Cross-device voiceprint registration method, electronic device and storage medium
By sharing the channel model between multiple terminal devices, converting the registered voice of one terminal device into the registered voice of other terminal devices, the problem of users requiring multiple voice inputs for voiceprint registration is solved, and multiple device registration is implemented at one time, improving user experience and device performance.
Patent Information
- Application Number
- CN202010650133.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-07
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-07-07
AI Technical Summary
In the prior art, users need to perform voice input in multiple terminal devices separately to register voiceprints, affecting the user experience.
By obtaining the registered voice of a terminal device and using the channel model between the device and other terminal devices for conversion processing, the registered voice corresponding to other terminal devices is generated, thereby realizing the voiceprint registration of multiple terminal devices at one voice input.
The number of voice input times for voiceprint registration of multiple terminal devices is reduced, the user experience is improved, and the computing volume of each terminal device is reduced, ensuring the performance of the device.
Smart Images

Figure CN114093368B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of terminals, and particularly relates to a cross-device voiceprint registration method, an electronic device, and a computer-readable storage medium. Background Art
[0002] Voiceprint recognition, that is, speaker recognition, is a technology for automatically identifying and verifying the identity of a speaker through voice, and is widely used in terminal devices such as mobile phones and smart speakers. Before voiceprint recognition, the user needs to first perform voiceprint registration in the terminal device, that is, the user needs to input registration voice in the terminal device to generate a voiceprint template based on the input registration voice, so as to identify the user's identity based on the voiceprint template. Currently, users often have multiple terminal devices, and to implement voiceprint recognition for each terminal device, the user needs to separately input voice in each terminal device to perform voiceprint registration for each terminal device, and the number of voice inputs is relatively large, affecting the user experience. Summary of the Invention
[0003] The embodiments of this application provide a cross-device voiceprint registration method, an electronic device, and a computer-readable storage medium, which can migrate the registration voice to achieve the purpose of performing voiceprint registration for multiple terminal devices with one voice input, reduce the number of voice inputs for cross-device voiceprint registration, and improve the user experience.
[0004] In a first aspect, the embodiments of this application provide a cross-device voiceprint registration method, which is applied to a second terminal device, and the method may include:
[0005] Obtain a first registration voice corresponding to a first terminal device;
[0006] Perform conversion processing on the first registration voice to obtain a second registration voice corresponding to the second terminal device;
[0007] Generate a voiceprint template corresponding to the second terminal device according to the second registration voice.
[0008] Through the above cross-device voiceprint registration method, the second terminal device can perform conversion processing on the first registration voice obtained by the first terminal device to generate a second registration voice corresponding to the second terminal device, and perform voiceprint registration on the second terminal device, so as to achieve the purpose of performing voiceprint registration for multiple terminal devices with one input of the registration voice, reduce the number of voice inputs for cross-device voiceprint registration, and improve the user experience.
[0009] In an example, the performing conversion processing on the first registration voice to obtain a second registration voice corresponding to the second terminal device may include:
[0010] Convert the first registered voice through the first channel model corresponding to the first terminal device to obtain the original voice corresponding to the first registered voice, where the first channel model is used to represent the mapping relationship between the voice corresponding to the first terminal device and the original voice;
[0011] Convert the original voice through the second channel model corresponding to the second terminal device to obtain the second registered voice corresponding to the second terminal device, where the second channel model is used to represent the mapping relationship between the voice corresponding to the second terminal device and the original voice.
[0012] In the embodiments of the present application, the channel model corresponding to the terminal device may be the channel model of the terminal device relative to the original voice signal, which is used to represent the mapping relationship between the voice signal obtained by the terminal device and the original voice signal, and the original voice signal is the voice signal without the addition of the channel information of the terminal device.
[0013] In another example, the converting the first registered voice to obtain the second registered voice corresponding to the second terminal device may include:
[0014] Convert the first registered voice through the third channel model to obtain the second registered voice corresponding to the second terminal device, where the third channel model is used to represent the mapping relationship between the voice corresponding to the first terminal device and the voice corresponding to the second terminal device.
[0015] In the embodiments of the present application, the channel model corresponding to the terminal device may also be the channel model between two terminal devices, which is used to represent the mapping relationship between the voice signals obtained by the two terminal devices.
[0016] In a possible implementation manner of the first aspect, the first channel model and the second channel model are channel models constructed based on the frequency response curve, or are channel models constructed based on the spectral characteristics.
[0017] Similarly, the third channel model is a channel model constructed based on the frequency response curve, or is a channel model constructed based on the spectral characteristics.
[0018] Exemplarily, when establishing a channel model of a terminal device relative to an original speech signal, an original sweep signal may be played to the terminal device, the frequency response curve St of the sound signal received by the terminal device may be measured, and the frequency response curve S of the original sweep signal may be measured. Then, each frequency response gain value may be calculated based on the frequency response curve St and the frequency response curve S, and a channel model of the terminal device relative to the original speech signal may be established based on each frequency response gain value. Among them, the original sweep signal may be an original sound signal output by a sweep signal generator, the frequency response gain value may be the ratio between the values corresponding to the same frequency in the frequency response curve St and the frequency response curve S, and the channel model of the terminal device relative to the original speech signal may be St / S.
[0019] Exemplarily, when establishing a channel model between two terminal devices, such as when establishing a channel model between a first terminal device and a second terminal device, an original sweep signal may be played to the first terminal device and the second terminal device respectively. Among them, the original sweep signals played to the first terminal device and the second terminal device are the same. The frequency response curve St1 of the sound signal received by the first terminal device and the frequency response curve St2 of the sound signal received by the second terminal device may be measured. Then, each frequency response gain value may be calculated based on the frequency response curves St1 and St2, and a channel model of the first terminal device relative to the second terminal device, and / or a channel model of the second terminal device relative to the first terminal device may be established based on the frequency response gain value. Among them, the channel model of the first terminal device relative to the second terminal device may be St1 / St2, and the channel model of the second terminal device relative to the first terminal device may be St2 / St1.
[0020] Exemplarily, when establishing a channel model between two terminal devices, such as when establishing a channel model between a first terminal device and a second terminal device, an original speech signal may be played to the first terminal device and the second terminal device respectively, the speech signals received by the first terminal device and the second terminal device may be obtained, and the original speech signal and the speech signals received by the first terminal device and the second terminal device may be sent to a preset neural network model. The neural network model may extract the spectral feature A corresponding to the original speech signal and the spectral feature B corresponding to the speech signal respectively, and learn the mapping relationship between the spectral feature A and the spectral feature B, so as to obtain a channel model of the first terminal device relative to the second terminal device and a channel model of the second terminal device relative to the first terminal device.
[0021] Exemplarily, when establishing a channel model between two terminal devices, such as when establishing a channel model between a first terminal device and a second terminal device, an original voice signal can be played to the first terminal device and the second terminal device respectively, the voice signal C received by the first terminal device and the voice signal D received by the second terminal device are obtained, and the voice signal C and the voice signal D can be sent to a preset neural network model. The neural network model can extract the spectral feature C corresponding to the voice signal C and the spectral feature D corresponding to the voice signal D respectively, and learn the mapping relationship between the spectral feature C and the spectral feature D, so as to obtain the channel model of the first terminal device relative to the second terminal device, and / or obtain the channel model of the second terminal device relative to the first terminal device.
[0022] It should be noted that generating the voiceprint template corresponding to the second terminal device according to the second registered voice may include:
[0023] Generating the voiceprint template corresponding to the second terminal device according to the voiceprint recognition model corresponding to the second terminal device and the second registered voice, where the voiceprint recognition model corresponding to the second terminal device is a voiceprint recognition model trained based on the training voice obtained by the second terminal device.
[0024] Optionally, after generating the voiceprint template corresponding to the second terminal device according to the second registered voice, it may further include:
[0025] Obtaining the authentication voice corresponding to the second terminal device;
[0026] Generating an authentication template corresponding to the authentication voice according to the voiceprint recognition model corresponding to the second terminal device and the authentication voice;
[0027] Determining the similarity between the authentication template and the voiceprint template;
[0028] When the similarity is greater than a preset similarity threshold, updating the voiceprint recognition model corresponding to the second terminal device according to the authentication voice.
[0029] It should be understood that when the similarity is greater than a preset similarity threshold, the method may further include: performing conversion processing on the authentication voice to obtain the training voice corresponding to the first terminal device, and sending the training voice to the first terminal device, where the training voice is used to update the voiceprint recognition model corresponding to the first terminal device.
[0030] Through the above optional methods, embodiments of the present application can obtain high-quality authentication voices during the daily use of users to update the voiceprint recognition models corresponding to each terminal device, improve the matching degree of the voiceprint recognition models corresponding to each terminal device with the actual usage scenarios, improve the robustness of voiceprint recognition in each terminal device, and thus improve the accuracy of voiceprint recognition in each terminal device.
[0031] In a second aspect, embodiments of the present application provide a cross-device voiceprint registration method, which is applied to a first terminal device or a server. The method may include:
[0032] Obtain a first registration voice corresponding to the first terminal device;
[0033] Perform conversion processing on the first registration voice to obtain a second registration voice corresponding to a second terminal device;
[0034] Send the second registration voice to the second terminal device, and the second registration voice is used to generate a voiceprint template corresponding to the second terminal device.
[0035] Through the above cross-device voiceprint registration method, the first terminal device or the server can perform conversion processing on the first registration voice obtained by the first terminal device, generate a second registration voice corresponding to the second terminal device, and send the second registration voice to the second terminal device. The second terminal device can directly perform voiceprint registration on the second terminal device based on the received second registration voice, achieving the purpose of voiceprint registration for multiple terminal devices with a single input of a registration voice, reducing the number of voice inputs for voiceprint registration of multiple terminal devices, and enhancing the user experience. At the same time, it can also reduce the computational load of the second terminal device and ensure the usage performance of the second terminal device.
[0036] In one example, the performing conversion processing on the first registration voice to obtain a second registration voice corresponding to a second terminal device may include:
[0037] Perform conversion processing on the first registration voice through a first channel model corresponding to the first terminal device to obtain an original voice corresponding to the first registration voice, where the first channel model is used to represent the mapping relationship between the voice corresponding to the first terminal device and the original voice;
[0038] Perform conversion processing on the original voice through a second channel model corresponding to the second terminal device to obtain a second registration voice corresponding to the second terminal device, where the second channel model is used to represent the mapping relationship between the voice corresponding to the second terminal device and the original voice.
[0039] In another example, the performing conversion processing on the first registration voice to obtain a second registration voice corresponding to a second terminal device may include:
[0040] The first registered voice is processed by a third channel model to obtain a second registered voice corresponding to the second terminal device, and the third channel model is used to characterize the mapping relationship between the voice corresponding to the first terminal device and the voice corresponding to the second terminal device.
[0041] It can be understood that when the method is applied to a server, after the server obtains the second registered voice corresponding to the second terminal device through conversion processing, it can also directly generate a voiceprint template corresponding to the second terminal device according to the voiceprint recognition model and the second registered voice corresponding to the second terminal device, and send the generated voiceprint template to the second terminal device, so as to directly generate a voiceprint template by the server and send it to the second terminal device, reducing the computing amount of the second terminal device and the performance requirements for the second terminal device.
[0042] Exemplarily, after sending the second registered voice to the second terminal device, it may further include:
[0043] Obtain an authentication voice corresponding to the second terminal device, where the similarity between the authentication voice and the voiceprint template corresponding to the second terminal device is greater than a preset similarity threshold;
[0044] Process the authentication voice to obtain a training voice corresponding to the first terminal device, and send the training voice to the first terminal device, where the training voice is used to update the voiceprint recognition model corresponding to the first terminal device.
[0045] Through the above optional method, when the method provided in the embodiments of the present application is applied to a server, the server can obtain high-quality authentication voices during the daily use of users, and process the authentication voices to obtain training voices corresponding to each terminal device, so as to update the voiceprint recognition models corresponding to each terminal device, improve the matching degree of the voiceprint recognition models corresponding to each terminal device with the actual use scenario, improve the robustness of voiceprint recognition in each terminal device, and thus improve the accuracy of voiceprint recognition in each terminal device.
[0046] In a third aspect, an embodiment of the present application provides a cross-device voiceprint registration method, which may include:
[0047] The first terminal device obtains a first registered voice corresponding to the first terminal device;
[0048] The first terminal device processes the first registered voice to obtain a first original voice corresponding to the first registered voice, and sends the first original voice to the second terminal device;
[0049] The second terminal device receives the first original voice from the first terminal device, and performs conversion processing on the first original voice to obtain a second registered voice corresponding to the second terminal device;
[0050] The second terminal device generates a voiceprint template corresponding to the second terminal device according to the second registered voice.
[0051] Through the above cross-device voiceprint registration method, the process of converting the first registered voice into the second registered voice can be decomposed into the process of converting the first registered voice into the original voice and the process of converting the original voice into the second registered voice. The process of converting the first registered voice into the original voice can be executed by the first terminal device, and the process of converting the original voice into the second registered voice can be executed by the second terminal device, so as to decompose the process of converting the first registered voice into the second registered voice to the first terminal device and the second terminal device for execution, which can reduce the computing amount of each terminal device and thus ensure the usage performance of each terminal device.
[0052] In a possible implementation manner of the third aspect, the first terminal device obtaining the first registered voice corresponding to the first terminal device may include:
[0053] The first terminal device obtains the interaction voice between the first terminal device and the user, and obtains the target voice in the interaction voice, where the target voice is the voice corresponding to the user;
[0054] The first terminal device obtains the first registered voice corresponding to the first terminal device from the target voice according to the signal-to-noise ratio and / or voice energy level corresponding to the target voice.
[0055] It should be noted that the first terminal device can also obtain the first registered voice from the daily voice interaction between the user and the first terminal device, so as to perform voiceprint registration of each terminal device in a self-learning and registration-free manner, thereby simplifying the operation process of voiceprint registration and improving the user experience.
[0056] In an example, after the second terminal device generates the voiceprint template corresponding to the second terminal device, it may further include:
[0057] The second terminal device obtains the authentication voice corresponding to the second terminal device, generates an authentication template corresponding to the authentication voice according to the voiceprint recognition model corresponding to the second terminal device and the authentication voice, and determines the similarity between the authentication template and the voiceprint template;
[0058] When the similarity is greater than a preset similarity threshold, the second terminal device updates the voiceprint recognition model corresponding to the second terminal device according to the authentication voice, and performs conversion processing on the authentication voice according to the second channel model corresponding to the second terminal device to obtain a second original voice corresponding to the authentication voice, and sends the second original voice to the first terminal device;
[0059] The first terminal device receives the second original voice from the second terminal device, performs conversion processing on the second original voice according to the first channel model corresponding to the first terminal device to obtain a training voice corresponding to the first terminal device, and updates the voiceprint recognition model corresponding to the first terminal device according to the training voice corresponding to the first terminal device.
[0060] Fourthly, an embodiment of the present application provides a cross-device voiceprint registration device, which is applied to a second terminal device. The device may include:
[0061] A registration voice acquisition module, configured to acquire a first registration voice corresponding to a first terminal device;
[0062] A conversion processing module, configured to perform conversion processing on the first registration voice to obtain a second registration voice corresponding to the second terminal device;
[0063] A voiceprint registration module, configured to generate a voiceprint template corresponding to the second terminal device according to the second registration voice.
[0064] In one example, the conversion processing module may include:
[0065] A first conversion processing unit, configured to perform conversion processing on the first registration voice through a first channel model corresponding to the first terminal device to obtain an original voice corresponding to the first registration voice, where the first channel model is used to represent a mapping relationship between the voice corresponding to the first terminal device and the original voice;
[0066] A second conversion processing unit, configured to perform conversion processing on the original voice through a second channel model corresponding to the second terminal device to obtain a second registration voice corresponding to the second terminal device, where the second channel model is used to represent a mapping relationship between the voice corresponding to the second terminal device and the original voice.
[0067] In another example, the conversion processing module may include:
[0068] A third conversion processing unit, configured to perform conversion processing on the first registered voice through a third channel model to obtain a second registered voice corresponding to the second terminal device, where the third channel model is used to characterize the mapping relationship between the voice corresponding to the first terminal device and the voice corresponding to the second terminal device.
[0069] In a possible implementation manner of the fourth aspect, the first channel model and the second channel model are channel models constructed based on a frequency response curve, or are channel models constructed based on spectral features.
[0070] In another possible implementation manner of the fourth aspect, the third channel model is a channel model constructed based on a frequency response curve, or is a channel model constructed based on spectral features.
[0071] Optionally, the voiceprint registration module is specifically configured to generate a voiceprint template corresponding to the second terminal device according to the voiceprint recognition model corresponding to the second terminal device and the second registered voice, where the voiceprint recognition model corresponding to the second terminal device is a voiceprint recognition model trained based on training voices obtained by the second terminal device.
[0072] In an example, the apparatus may further include:
[0073] An authentication voice acquisition module, configured to acquire an authentication voice corresponding to the second terminal device;
[0074] An authentication template generation module, configured to generate an authentication template corresponding to the authentication voice according to the voiceprint recognition model corresponding to the second terminal device and the authentication voice;
[0075] A similarity determination module, configured to determine the similarity between the authentication template and the voiceprint template;
[0076] A model update module, configured to update the voiceprint recognition model corresponding to the second terminal device according to the authentication voice when the similarity is greater than a preset similarity threshold.
[0077] It should be understood that when the similarity is greater than a preset similarity threshold, the apparatus may further include:
[0078] A training voice acquisition module, configured to perform conversion processing on the authentication voice to obtain a training voice corresponding to the first terminal device, and send the training voice to the first terminal device, where the training voice is used to update the voiceprint recognition model corresponding to the first terminal device.
[0079] In a fifth aspect, an embodiment of the present application provides a cross-device voiceprint registration apparatus, which is applied to a first terminal device or a server, and the apparatus may include:
[0080] A first registered voice acquisition module, configured to acquire a first registered voice corresponding to the first terminal device;
[0081] A conversion processing module, configured to perform conversion processing on the first registered voice to obtain a second registered voice corresponding to the second terminal device;
[0082] A second registered voice sending module, configured to send the second registered voice to the second terminal device, where the second registered voice is used to generate a voiceprint template corresponding to the second terminal device.
[0083] In a possible implementation manner of the fifth aspect, the conversion processing module may include:
[0084] A first conversion processing unit, configured to perform conversion processing on the first registered voice through a first channel model corresponding to the first terminal device to obtain an original voice corresponding to the first registered voice, where the first channel model is used to represent a mapping relationship between the voice corresponding to the first terminal device and the original voice;
[0085] A second conversion processing unit, configured to perform conversion processing on the original voice through a second channel model corresponding to the second terminal device to obtain a second registered voice corresponding to the second terminal device, where the second channel model is used to represent a mapping relationship between the voice corresponding to the second terminal device and the original voice.
[0086] In a possible implementation manner of the fifth aspect, the conversion processing module may further include:
[0087] A third conversion processing unit, configured to perform conversion processing on the first registered voice through a third channel model to obtain a second registered voice corresponding to the second terminal device, where the third channel model is used to represent a mapping relationship between the voice corresponding to the first terminal device and the voice corresponding to the second terminal device.
[0088] In a sixth aspect, an embodiment of the present application provides a cross-device voiceprint registration system, which may include a first terminal device and a second terminal device. The first terminal device includes a registered voice acquisition module and a first conversion processing module, and the second terminal device includes a second conversion processing module and a voiceprint registration module;
[0089] The registered voice acquisition module is configured to acquire a first registered voice corresponding to the first terminal device;
[0090] The first conversion processing module is configured to perform conversion processing on the first registered voice to obtain a first original voice corresponding to the first registered voice, and send the first original voice to the second terminal device;
[0091] The second conversion processing module is configured to receive the first original voice from the first terminal device, and perform conversion processing on the first original voice to obtain a second registered voice corresponding to the second terminal device;
[0092] The voiceprint registration module is configured to generate a voiceprint template corresponding to the second terminal device according to the second registered voice.
[0093] In one example, the registered voice acquisition module may include:
[0094] The target voice acquisition unit is configured to acquire the interaction voice between the first terminal device and the user, and acquire the target voice in the interaction voice, where the target voice is the voice corresponding to the user;
[0095] The registered voice acquisition unit is configured to acquire the first registered voice corresponding to the first terminal device from the target voice according to the signal-to-noise ratio and / or voice energy level corresponding to the target voice.
[0096] In a possible implementation manner of the sixth aspect, the second terminal device may further include an authentication voice acquisition module, an authentication template generation module, a similarity determination module, and a first model update module; the first terminal device may further include a training voice acquisition module and a second model update module:
[0097] The authentication voice acquisition module is configured to acquire the authentication voice corresponding to the second terminal device;
[0098] The authentication template generation module is configured to generate an authentication template corresponding to the authentication voice according to the voiceprint recognition model corresponding to the second terminal device and the authentication voice;
[0099] The similarity determination module is configured to determine the similarity between the authentication template and the voiceprint template;
[0100] The first model update module is configured to, when the similarity is greater than a preset similarity threshold, update the voiceprint recognition model corresponding to the second terminal device according to the authentication voice, and perform conversion processing on the authentication voice according to the second channel model corresponding to the second terminal device to obtain a second original voice corresponding to the authentication voice, and send the second original voice to the first terminal device;
[0101] The training voice acquisition module is configured to receive the second original voice from the second terminal device, and perform conversion processing on the second original voice according to the first channel model corresponding to the first terminal device to obtain a training voice corresponding to the first terminal device;
[0102] The second model update module is configured to update the voiceprint recognition model corresponding to the first terminal device according to the training voice corresponding to the first terminal device.
[0103] In a seventh aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the cross-device voiceprint registration method according to any one of the above first aspects or any one of the second aspects.
[0104] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a computer, the computer implements the cross-device voiceprint registration method according to any one of the above first aspects or any one of the second aspects.
[0105] In a ninth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device is caused to execute the cross-device voiceprint registration method according to any one of the above first aspects or any one of the second aspects. Description of the Drawings
[0106] Figure 1 is a schematic structural diagram of a terminal device provided by an embodiment of the present application;
[0107] Figure 2 is a schematic software architecture diagram of a terminal device provided by an embodiment of the present application;
[0108] Figure 3 is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0109] Figure 4 is a schematic diagram of an application interface provided by an embodiment of the present application Figure 1 ;
[0110] Figure 5 is a schematic diagram of an application interface provided by an embodiment of the present application Figure 2 ;
[0111] Figure 6 is a schematic diagram of an application interface provided by an embodiment of the present application Figure 3 ;
[0112] Figure 7 is a schematic diagram of an application interface provided by an embodiment of the present application Figure 4 ;
[0113] Figure 8 is a schematic flowchart of the cross-device voiceprint registration method provided by Embodiment 1 of the present application;
[0114] Figure 9 It is another schematic diagram of an application scenario provided by an embodiment of the present application;
[0115] Figure 10 It is a schematic flowchart of a cross-device voiceprint registration method provided by the second embodiment of the present application;
[0116] Figure 11 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0117] It should be understood that when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0118] It should also be understood that the term "and / or" used in the specification and the appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0119] As used in the specification and the appended claims of the present application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" according to the context.
[0120] In addition, in the description of the specification and the appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0121] Referring to "one embodiment" or "some embodiments" described in the specification of the present application means that a specific feature, structure, or characteristic described in conjunction with the embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0122] Voiceprint recognition is a technology for automatically identifying and verifying the identity of a speaker through voice, which can be applied to terminal devices such as mobile phones, smart watches, and smart speakers. Among them, voiceprint recognition includes two stages: voiceprint registration and voiceprint verification. In the voiceprint registration stage, the user needs to input registration voice to the terminal device, and the terminal device can generate a voiceprint template based on the obtained registration voice; in the voiceprint verification stage, the terminal device can score the similarity between the authentication voice input by the user and the voiceprint template generated in the registration stage to identify the user's identity.
[0123] When performing voiceprint registration, the registration voice obtained by the terminal device is affected by the channel corresponding to the terminal device, that is, the channel information corresponding to the terminal device is attached to the obtained registration voice. Different terminal devices are often composed of different hardware components, so different terminal devices have different channel information, resulting in different registration voices for voiceprint registration on different terminal devices. Due to the difference in channels, the registration voice obtained by a certain terminal device can only be used for voiceprint registration of that terminal device. When the user has multiple terminal devices and wants to use voiceprint recognition on multiple terminal devices, the user needs to input voice separately on each terminal device for voiceprint registration of each terminal device, and the number of voice inputs is relatively large, affecting the user experience.
[0124] To solve the above problems, the embodiments of the present application provide a cross-device voiceprint registration method, an electronic device, and a computer-readable storage medium, which can perform conversion processing on the registration voice obtained by a certain terminal device to generate the registration voice corresponding to other terminal devices, so as to perform voiceprint registration on other terminal devices, achieving the purpose of voiceprint registration of multiple terminal devices with a single input of registration voice, reducing the number of voice inputs for multi-terminal device voiceprint registration, and improving the user experience.
[0125] It should be noted that the terminal devices involved in the embodiments of the present application can be mobile phones, tablet computers, wearable devices (such as smart headphones, smart bracelets, etc.), smart speakers, smart homes, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), personal digital assistants (PDAs), desktop computers, etc.
[0126] First, the terminal devices involved in the embodiments of the present application will be introduced below. Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of the terminal device 100 provided by the embodiments of the present application.
[0127] The terminal device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0128] It can be understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on the terminal device 100. In other embodiments of this application, the terminal device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0129] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0130] The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0131] A memory can also be set in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0132] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0133] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple groups of I2C buses. The processor 110 can be respectively coupled to the touch sensor 180K, the charger, the flashlight, the camera 193, etc. through different I2C bus interfaces. For example: The processor 110 can be coupled to the touch sensor 180K through the I2C interface, enabling the processor 110 to communicate with the touch sensor 180K through the I2C bus interface to implement the touch function of the terminal device 100.
[0134] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple groups of I2S buses. The processor 110 can be coupled to the audio module 170 through the I2S bus to implement communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit an audio signal to the wireless communication module 160 through the I2S interface to implement the function of answering a call through a Bluetooth headset.
[0135] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled through a PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 through the PCM interface to implement the function of answering a call through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0136] The UART interface is a general-purpose serial data bus for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 through the UART interface to implement the Bluetooth function. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 through the UART interface to implement the function of playing music through a Bluetooth headset.
[0137] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 110 and the camera 193 communicate through the CSI interface to implement the shooting function of the terminal device 100. The processor 110 and the display screen 194 communicate through the DSI interface to implement the display function of the terminal device 100.
[0138] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or as a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0139] The USB interface 130 is an interface that conforms to the USB standard specification and can specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the terminal device 100, and can also be used for data transmission between the terminal device 100 and peripheral devices. It can also be used to connect a headset to play audio. This interface can also be used to connect other terminal devices, such as AR devices, etc.
[0140] It can be understood that the interface connection relationships among the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the terminal device 100. In other embodiments of the present application, the terminal device 100 may also adopt different interface connection methods or a combination of multiple interface connection methods in the above embodiments.
[0141] The charging management module 140 is used to receive a charging input from a charger. Among them, the charger may be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 may receive the charging input of the wired charger through the USB interface 130. In some embodiments of wireless charging, the charging management module 140 may receive the wireless charging input through the wireless charging coil of the terminal device 100. While charging the battery 142, the charging management module 140 may also supply power to the terminal device through the power management module 141.
[0142] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives the inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the display screen 194, the camera 193, the wireless communication module 160, etc. The power management module 141 may also be used to monitor parameters such as the battery capacity, the number of battery cycles, and the battery health status (leakage, impedance). In some other embodiments, the power management module 141 may also be disposed in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 may also be disposed in the same device.
[0143] The wireless communication function of the terminal device 100 may be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.
[0144] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the terminal device 100 may be used to cover a single or multiple communication frequency bands. Different antennas may also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 may be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna may be used in combination with a tuning switch.
[0145] The mobile communication module 150 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the terminal device 100. The mobile communication module 150 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves through the antenna 1, filter, amplify, etc. the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 150 can be disposed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 can be disposed in the same device.
[0146] The modulation and demodulation processor can include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, receiver 170B, etc.), or displays an image or video through the display screen 194. In some embodiments, the modulation and demodulation processor can be an independent device. In other embodiments, the modulation and demodulation processor can be independent of the processor 110 and be disposed in the same device as the mobile communication module 150 or other functional modules.
[0147] The wireless communication module 160 may provide solutions for wireless communications applied to the terminal device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 may also receive signals to be sent from the processor 110, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna 2 for radiation.
[0148] In some embodiments, antenna 1 of the terminal device 100 is coupled to the mobile communication module 150, and antenna 2 is coupled to the wireless communication module 160, enabling the terminal device 100 to communicate with the network and other devices through wireless communication technologies. The wireless communication technologies may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), BeiDou Navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).
[0149] The terminal device 100 implements the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0150] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-Oled, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the terminal device 100 may include one or N display screens 194, where N is a positive integer greater than 1.
[0151] The terminal device 100 can implement the shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, an application processor, etc.
[0152] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera photosensitive element. The optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the noise, brightness, and skin color of the image through algorithms. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0153] The camera 193 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In some embodiments, the terminal device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0154] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the terminal device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0155] The video codec is used to compress or decompress digital videos. The terminal device 100 can support one or more video codecs. In this way, the terminal device 100 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0156] The NPU is a neural-network (NN) computing processor. By learning from the biological neural network structure, such as learning from the transmission mode between human brain neurons, it can quickly process the input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the terminal device 100 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.
[0157] The external memory interface 120 can be used to connect to an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0158] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.). The data storage area can store data created during the use of the terminal device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the terminal device 100 by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.
[0159] The terminal device 100 can implement audio functions through the audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and the application processor, etc. For example, music playback, recording, etc.
[0160] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.
[0161] The speaker 170A, also known as a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The terminal device 100 can listen to music or hands-free calls through the speaker 170A.
[0162] The receiver 170B, also known as an "earpiece", is used to convert an audio electrical signal into a sound signal. When the terminal device 100 answers a call or a voice message, the voice can be listened to by placing the receiver 170B close to the human ear.
[0163] The microphone 170C, also known as a "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak by bringing the mouth close to the microphone 170C to input the sound signal into the microphone 170C. The terminal device 100 can be provided with at least one microphone 170C. In some other embodiments, the terminal device 100 can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the terminal device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.
[0164] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5 mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0165] The pressure sensor 180A is used to sense pressure signals and can convert the pressure signals into electrical signals. In some embodiments, the pressure sensor 180A may be disposed on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor may include at least two parallel plates having conductive materials. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The terminal device 100 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on the display screen 194, the terminal device 100 detects the intensity of the touch operation according to the pressure sensor 180A. The terminal device 100 can also calculate the position of the touch according to the detection signal of the pressure sensor 180A. In some embodiments, touch operations acting on the same touch position but with different touch operation intensities may correspond to different operation instructions. For example: when a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, the instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, the instruction to create a new short message is executed.
[0166] The gyroscope sensor 180B can be used to determine the motion posture of the terminal device 100. In some embodiments, the angular velocity of the terminal device 100 around three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for anti-shake during shooting. Exemplarily, when the shutter is pressed, the gyroscope sensor 180B detects the angle of jitter of the terminal device 100, calculates the distance that the lens module needs to compensate according to the angle, and enables the lens to offset the jitter of the terminal device 100 through reverse movement to achieve anti-shake. The gyroscope sensor 180B can also be used for navigation and somatosensory game scenarios.
[0167] The barometric pressure sensor 180C is used to measure the barometric pressure. In some embodiments, the terminal device 100 calculates the altitude according to the barometric pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.
[0168] The magnetic sensor 180D includes a Hall sensor. The terminal device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip leather case. In some embodiments, when the terminal device 100 is a flip phone, the terminal device 100 can detect the opening and closing of the flip according to the magnetic sensor 180D. Furthermore, according to the detected opening and closing state of the leather case or the opening and closing state of the flip, features such as automatic flip unlocking are set.
[0169] The acceleration sensor 180E can detect the magnitude of the acceleration of the terminal device 100 in various directions (generally three axes). When the terminal device 100 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the terminal device and is applied to applications such as horizontal and vertical screen switching and pedometers.
[0170] A distance sensor 180F is used to measure distance. The terminal device 100 can measure distance through infrared or laser. In some embodiments, when shooting a scene, the terminal device 100 can use the distance sensor 180F to measure distance to achieve fast focusing.
[0171] The proximity light sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The light-emitting diode may be an infrared light-emitting diode. The terminal device 100 emits infrared light outward through the light-emitting diode. The terminal device 100 uses the photodiode to detect the infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the terminal device 100. When insufficient reflected light is detected, the terminal device 100 can determine that there is no object near the terminal device 100. The terminal device 100 can use the proximity light sensor 180G to detect when the user holds the terminal device 100 close to the ear for a call, so as to automatically turn off the screen to achieve the purpose of power saving. The proximity light sensor 180G can also be used for automatic unlocking and locking of the holster mode and pocket mode.
[0172] The ambient light sensor 180L is used to sense the ambient light brightness. The terminal device 100 can adaptively adjust the brightness of the display screen 194 according to the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also cooperate with the proximity light sensor 180G to detect whether the terminal device 100 is in the pocket to prevent accidental touch.
[0173] The fingerprint sensor 180H is used to collect fingerprints. The terminal device 100 can use the collected fingerprint characteristics to achieve fingerprint unlocking, access application locks, fingerprint taking pictures, fingerprint answering calls, etc.
[0174] The temperature sensor 180J is used to detect temperature. In some embodiments, the terminal device 100 uses the temperature detected by the temperature sensor 180J to execute a temperature processing strategy. For example, when the temperature reported by the temperature sensor 180J exceeds a threshold, the terminal device 100 reduces the performance of the processor located near the temperature sensor 180J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the terminal device 100 heats the battery 142 to avoid abnormal shutdown of the terminal device 100 caused by low temperature. In other embodiments, when the temperature is lower than yet another threshold, the terminal device 100 boosts the output voltage of the battery 142 to avoid abnormal shutdown caused by low temperature.
[0175] The touch sensor 180K, also known as the "touch control device". The touch sensor 180K can be disposed on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also known as the "touch control screen". The touch sensor 180K is used to detect touch operations acting thereon or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In some other embodiments, the touch sensor 180K can also be disposed on the surface of the terminal device 100, at a different position from that of the display screen 194.
[0176] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire vibration signals of the vibrating bone mass of the human vocal part. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure pulsation signals. In some embodiments, the bone conduction sensor 180M can also be disposed in the earphone to form a bone conduction earphone. The audio module 170 can parse out voice signals based on the vibration signals of the vibrating bone mass of the human vocal part acquired by the bone conduction sensor 180M to implement the voice function. The application processor can parse out heart rate information based on the blood pressure pulsation signals acquired by the bone conduction sensor 180M to implement the heart rate detection function.
[0177] The button 190 includes a power-on button, a volume button, etc. The button 190 can be a mechanical button or a touch button. The terminal device 100 can receive button inputs to generate key signal inputs related to the user settings and function control of the terminal device 100.
[0178] The motor 191 can generate vibration prompts. The motor 191 can be used for incoming call vibration prompts and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, playing audio, etc.) can correspond to different vibration feedback effects. Touch operations acting on different regions of the display screen 194 can also correspond to different vibration feedback effects for the motor 191. Different application scenarios (such as time reminder, receiving information, alarm clock, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0179] The indicator 192 can be an indicator light and can be used to indicate the charging state, power change, and can also be used to indicate messages, missed calls, notifications, etc.
[0180] The SIM card interface 195 is used to connect to the SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation from the terminal device 100. The terminal device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The terminal device 100 interacts with the network through the SIM card to implement functions such as calls and data communication. In some embodiments, the terminal device 100 uses an eSIM, that is, an embedded SIM card. The eSIM card can be embedded in the terminal device 100 and cannot be separated from the terminal device 100.
[0181] The software system of the terminal device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of this application, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the terminal device 100.
[0182] Figure 2 It is a software structure block diagram of the terminal device 100 in the embodiments of this application.
[0183] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0184] The application layer can include a series of application packages.
[0185] As Figure 2 shown, the application packages can include applications such as the camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and short message.
[0186] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.
[0187] As Figure 2 shown, the application framework layer can include the window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0188] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0189] The content provider is used to store and obtain data, and make this data accessible to application programs. The data may include videos, images, audio, incoming and outgoing calls, browsing history and bookmarks, phone books, etc.
[0190] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build application programs. The display interface can be composed of one or more views. For example, a display interface including a text message notification icon may include a view for displaying text and a view for displaying pictures.
[0191] The phone manager is used to provide the communication function of the terminal device 100. For example, the management of call states (including connection, disconnection, etc.).
[0192] The resource manager provides various resources for application programs, such as localized strings, icons, pictures, layout files, video files, etc.
[0193] The notification manager enables application programs to display notification information in the status bar. It can be used to convey notification-type messages, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that the download is completed, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as the notification of a background-running application program, and can also be a notification that appears on the screen in the form of a dialogue window. For example, prompt text information in the status bar, emit a prompt sound, the terminal device vibrates, the indicator light flashes, etc.
[0194] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for the scheduling and management of the Android system.
[0195] The core libraries contain two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core libraries of Android.
[0196] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.
[0197] The system library may include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.
[0198] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.
[0199] The media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0200] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.
[0201] The 2D graphics engine is a drawing engine for 2D drawing.
[0202] The kernel layer is the layer between hardware and software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.
[0203] The following combines the capture and photo-taking scenario to exemplarily illustrate the working processes of the software and hardware of the terminal device 100.
[0204] When the touch sensor 180K receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including information such as touch coordinates and the timestamp of the touch operation). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking the touch operation as a touch click operation and the control corresponding to the click operation as the control of the camera application icon as an example, the camera application calls the interface of the application framework layer to start the camera application, and then starts the camera driver by calling the kernel layer, and captures a static image or video through the camera 193.
[0205] The following introduces the cross-device voiceprint registration method provided by the embodiments of the present application. Among them, the method has previously established corresponding channel models for each terminal device of the user. When the user performs voiceprint registration on a certain terminal device, the method can obtain the registration voice input by the user in this terminal device, and can perform conversion processing on the registration voice according to the channel models corresponding to each terminal device to obtain the registration voice corresponding to the user's other terminal devices, so as to perform voiceprint registration on the user's other terminal devices and obtain the voiceprint templates corresponding to each terminal device.
[0206] It should be noted that the channel model corresponding to the terminal device can be the channel model of the terminal device relative to the original voice signal, which is used to characterize the mapping relationship between the voice signal obtained by the terminal device and the original voice signal. The original voice signal is the voice signal without the addition of the channel information of the terminal device; alternatively, it can be the channel model between two terminal devices, which is used to characterize the mapping relationship between the voice signals obtained by the two terminal devices.
[0207] In the embodiments of the present application, the channel model corresponding to the terminal device can be constructed based on the frequency response curve of the voice signal, or the channel model corresponding to the terminal device can be constructed based on the spectral characteristics of the voice signal.
[0208] Exemplarily, when establishing the channel model of the terminal device relative to the original voice signal, the original swept-frequency signal can be played to the terminal device, the frequency response curve St of the sound signal received by the terminal device is measured, and the frequency response curve S of the original swept-frequency signal is measured. Then, the frequency response gain values can be calculated according to the frequency response curve St and the frequency response curve S, and the channel model of the terminal device relative to the original voice signal can be established based on the frequency response gain values. Among them, the original swept-frequency signal can be the original sound signal output by the swept-frequency signal generator, the frequency response gain value can be the ratio between the values corresponding to the same frequency in the frequency response curve St and the frequency response curve S, and the channel model of the terminal device relative to the original voice signal can be St / S.
[0209] Alternatively, the original voice signal can be played to the terminal device, the voice signal received by the terminal device is obtained, and the original voice signal and the voice signal received by the terminal device can be sent to a preset neural network model. The neural network model can extract the spectral feature A corresponding to the original voice signal and the spectral feature B corresponding to the voice signal respectively, and learn the mapping relationship between the spectral feature A and the spectral feature B, so as to obtain the channel model of the terminal device relative to the original voice signal. Among them, the original voice signal can be the voice signal collected based on a standard sound collection device (such as a microphone), and the preset neural network model can be a neural network model trained based on a large amount of voice signal data pairs. Each voice signal data pair includes the original voice signal and the voice signal received by the terminal device after passing through the terminal device.
[0210] Exemplarily, when establishing a channel model between two terminal devices, such as when establishing a channel model between a first terminal device and a second terminal device, an original swept-frequency signal can be played to the first terminal device and the second terminal device respectively. Among them, the original swept-frequency signals played to the first terminal device and the second terminal device are the same. Measure the frequency response curve St1 of the sound signal received by the first terminal device and the frequency response curve St2 of the sound signal received by the second terminal device. Then, the frequency response gain values can be calculated according to the frequency response curves St1 and St2, and a channel model of the first terminal device relative to the second terminal device, and / or a channel model of the second terminal device relative to the first terminal device can be established according to the frequency response gain values. Among them, the channel model of the first terminal device relative to the second terminal device can be St1 / St2, and the channel model of the second terminal device relative to the first terminal device can be St2 / St1.
[0211] Alternatively, an original voice signal can be played to the first terminal device and the second terminal device respectively, the voice signal C received by the first terminal device and the voice signal D received by the second terminal device are obtained, and the voice signal C and the voice signal D can be sent to a preset neural network model. The neural network model can extract the spectral feature C corresponding to the voice signal C and the spectral feature D corresponding to the voice signal D respectively, and learn the mapping relationship between the spectral feature C and the spectral feature D, so as to obtain a channel model of the first terminal device relative to the second terminal device, and / or obtain a channel model of the second terminal device relative to the first terminal device. Among them, the original voice signal can be a voice signal collected based on a standard sound collection device (such as a microphone), and the preset neural network model can be a neural network model trained based on a large amount of voice signal data pairs. Each voice signal data pair includes the voice signal corresponding to the first terminal device and the voice signal corresponding to the second terminal device.
[0212] It should be understood that the voiceprint templates corresponding to each terminal device can be generated according to the voiceprint recognition models corresponding to each terminal device. Therefore, in the embodiments of the present application, the voiceprint recognition models corresponding to each terminal device can be trained in advance to generate the voiceprint templates corresponding to each terminal device according to the voiceprint recognition models corresponding to each terminal device and the registered voices. Among them, the voiceprint template can be a feature vector output by the voiceprint recognition model, etc., that is, the voiceprint template can be a feature vector composed of the voiceprint features extracted by the voiceprint recognition model from the registered voice, etc.
[0213] Here, the voiceprint recognition model can be a voiceprint recognition model based on a Gaussian mixture model - universal background model (GMM - UBM), or can be a voiceprint recognition model based on a support vector machine (SVM), or can be a voiceprint recognition model based on joint factor analysis (JFA), or can be a voiceprint recognition model based on an identity vector (i - vector), or can be a voiceprint recognition model based on a time - delay neural network (TDNN). Among them, the voiceprint recognition models corresponding to each terminal device can be the same or different. For example, the voiceprint recognition models corresponding to each terminal device can all be voiceprint recognition models based on GMM - UBM, or can all be voiceprint recognition models based on TDNN. For example, the voiceprint recognition model corresponding to terminal device A can be a voiceprint recognition model based on GMM - UBM, the voiceprint recognition model corresponding to terminal device B can be a voiceprint recognition model based on SVM, and the voiceprint recognition model corresponding to terminal device C can be a voiceprint recognition model based on JFA, and so on.
[0214] It should be understood that the specific process of generating the voiceprint template corresponding to the terminal device according to the voiceprint recognition model corresponding to the terminal device and the registered voice can be: First, the voice features of the registered voice can be extracted. Among them, the extracted voice features can be mel - frequency cepstral coefficients (MFCC) or can be filter bank (FBank) features. Then, the voiceprint recognition model can be used to process the extracted voice features to obtain the voiceprint template corresponding to the registered voice. For example, the GMM - UBM model can be used to process the voice features to obtain the Gaussian mean supervector as the voiceprint template corresponding to the registered voice; or the i - vector model can be used to process the voice features to obtain the i - vector as the voiceprint template corresponding to the registered voice; or the deep neural network (DNN) can be used to process the voice features to obtain the d - vector as the voiceprint template corresponding to the registered voice; or the TDNN network can be used to process the voice features to obtain the x - vector as the voiceprint template corresponding to the registered voice, and so on.
[0215] It should be noted that the voiceprint recognition models corresponding to the respective terminal devices can be trained based on the training voice sets obtained by the respective terminal devices. For example, the voiceprint recognition model corresponding to terminal device A can be trained using the training voice set A obtained by terminal device A, the voiceprint recognition model corresponding to terminal device B can be trained using the training voice set B obtained by terminal device B, the voiceprint recognition model corresponding to terminal device C can be trained using the training voice set C obtained by terminal device C, and so on. Among them, the training voices obtained by the terminal devices are attached with the channel information corresponding to the respective terminal devices. Here, the training voice sets obtained by the respective terminal devices are used to train the respective voiceprint recognition models, so that the respective voiceprint recognition models can be better matched with the respective terminal devices, so as to improve the accuracy of the voiceprint templates corresponding to the respective terminal devices, thereby improving the recognition accuracy of the voiceprint recognition of the respective terminal devices.
[0216] It should be noted that the embodiments of the present application do not specifically limit the training process of the voiceprint recognition model. For example, existing training methods can be used to train the voiceprint recognition model.
[0217] The cross-device voiceprint registration method provided by the embodiments of the present application will be introduced below in combination with specific application scenarios.
[0218]
Embodiment 1
[0219] Figure 3 is a schematic diagram of the application scenario of the cross-device voiceprint registration method provided by Embodiment 1 of the present application. As Figure 3 shown, the application scenario may include multiple terminal devices 100, and the respective terminal devices 100 can be interconnected through short-range communication or a network. Among them, each terminal device 100 or the storage device communicatively connected to each terminal device 100 stores a channel model of the terminal device relative to the original voice signal, or stores a channel model between the terminal device and any other terminal device.
[0220] It should be understood that the terminal device 100 is a terminal device with a cross-device voiceprint registration function. After the cross-device voiceprint registration function of each terminal device 100 is enabled, when a user performs voiceprint registration on a certain terminal device 100 (hereinafter referred to as the first terminal device), the first terminal device can send the first registration voice input by the user to the first terminal device to other terminal devices 100 (collectively referred to as the second terminal devices hereinafter). Each second terminal device can use the channel model of the first terminal device relative to the original voice signal and the channel model of the second terminal device relative to the original voice signal to perform conversion processing on the first registration voice to obtain the second registration voice corresponding to each second terminal device; or each second terminal device can perform conversion processing on the first registration voice according to the channel model between the first terminal device and the second terminal device to obtain the second registration voice corresponding to each second terminal device, so as to generate the voiceprint template corresponding to each second terminal device according to each second registration voice, reducing the number of voice inputs for cross-device voiceprint registration of multiple terminal devices.
[0221] In this embodiment, the cross-device voiceprint registration function of any terminal device 100 can be manually enabled by the user.
[0222] In one example, the terminal device 100 can display an icon and / or a menu bar for cross-device voiceprint registration to allow the user to manually enable the cross-device voiceprint registration function of the terminal device 100. For example, as Figure 4 shown in (a) of, the terminal device 100 can display an icon for cross-device voiceprint registration in the quick control interface, and the quick control interface can also include icons for conventional functions such as Bluetooth, flight mode, mobile data, wireless local area network, flashlight, brightness, etc., to implement quick operations for related functions such as Bluetooth, flight mode, and mobile data. Or, as Figure 4 shown in (b) of, the terminal device 100 can display a menu bar for cross-device voiceprint registration in the settings interface, and the settings interface can also include setting menu bars for conventional functions such as Bluetooth, flight mode, mobile data, wireless local area network, brightness, etc., to implement setting operations for related functions such as Bluetooth, flight mode, and wireless local area network.
[0223] When the terminal device 100 detects a relevant operation of the user on the quick control interface or the settings interface (such as detecting a relevant touch operation or click operation), it can be determined that the cross-device voiceprint registration function of the terminal device 100 needs to be enabled. The terminal device 100 can directly enable the cross-device voiceprint registration function, or can pop up a window to ask the user whether to confirm to enable the cross-device voiceprint registration function. For example, the terminal device 100 detects that the icon for cross-device voiceprint registration shown in Figure 4 as (a) of is lit, or detects that as Figure 4When the power-on key of the menu bar for cross-device voiceprint registration shown in (b) in [Figure X] is turned on, the terminal device 100 can directly enable the cross-device voiceprint registration function; or a pop-up window as shown in Figure 4 (c) in [Figure X] can be displayed to show the inquiry message "Are you sure to enable cross-device voiceprint registration?" and the selection keys "Yes" and "No". When the terminal device 100 detects that the user clicks "Yes", the terminal device 100 can enable the cross-device voiceprint registration function.
[0224] In one example, when the terminal device 100 detects that the user inputs the registration voice in the terminal device 100, the terminal device 100 can ask the user whether to enable the cross-device voiceprint registration function through a pop-up window. For example, when the terminal device 100 detects that the user is inputting the registration voice in the terminal device 100 to perform voiceprint registration on the terminal device 100, the terminal device 100 can display a pop-up window as shown in Figure 5 to show "Voiceprint registration is in progress. Do you want to enable cross-device voiceprint registration?" and provide the selection keys "Yes" or "No". When the terminal device 100 detects that the user clicks "Yes", the terminal device 100 can enable the cross-device voiceprint registration function.
[0225] After the cross-device voiceprint registration function of the terminal device 100 is enabled, the terminal device 100 can provide a voiceprint registration management interface as shown in Figure 6 (a) in [Figure X] for the user to select the terminal device for cross-device voiceprint registration. Among them, the voiceprint registration management interface may include: a selected device column 60, a candidate device column 61, and an add control 62 for adding a new device. The selected device column 60 is used to display the terminal devices that the user has selected for cross-device voiceprint registration. If no terminal device is selected, a prompt message indicating that no terminal device is selected can be displayed. For example, a prompt message "No terminal device has been selected yet" as shown in Figure 6 (a) in [Figure X] can be displayed. The candidate device column 61 is used to display the terminal devices with the cross-device voiceprint registration function. Here, the terminal devices displayed in the candidate device column 61 can be associated with the user's account. When the user logs in to the account, all the terminal devices with the cross-device voiceprint registration function associated with the account can be displayed in the candidate device column 61. It should be understood that the terminal devices displayed in the selected device column 60 and the candidate device column 61 can be the device names and / or device identifiers of the terminal devices. For example, the device names or device identifiers "AAA", "BBB", "CCC", "DDD", etc. as shown in Figure 6 and Figure 7 in [Figure X] can be displayed.
[0226] When selecting a terminal device for cross-device voiceprint registration, the user can directly select the device name and / or device identifier in the candidate device list 61 to select the terminal device for cross-device voiceprint registration. After the user selects the device name and / or device identifier in the candidate device list 61, the terminal device 100 can directly add the device name and / or device identifier to the selected device list 60; or, a pop-up window can be displayed to ask the user if they are sure to select this terminal device. For example, as shown in (a) of Figure 6 , when the terminal device 100 detects that the user has performed a selection operation (such as a click operation or a touch operation) on the device name "AAA" in the candidate device list 61, the terminal device 100 can display a pop-up window as shown in (b) of Figure 6 to display the query information "Are you sure to select the terminal device AAA?", as well as the selection keys "Confirm" and "Cancel". When the terminal device 100 detects that the user clicks "Confirm", as shown in (c) of Figure 6 , the terminal device 100 can add the device name "AAA" to the selected device list 60 and can delete the device name "AAA" from the candidate device list 61.
[0227] If the user needs to delete a selected terminal device, the user can perform a selection operation (such as a click operation or a touch operation) on the device name and / or device identifier to be deleted in the selected device list 60. At this time, the terminal device 100 can delete the device name and / or device identifier from the selected device list 60; or, a pop-up window can be displayed to ask the user if they are sure to delete this terminal device. For example, as shown in (a) of Figure 7 , when the terminal device 100 detects that the user has performed a selection operation on the device name "BBB" in the selected device list 60, the terminal device 100 can display a pop-up window as shown in (b) of Figure 7 to display the query information "Are you sure to delete the terminal device BBB?", as well as the selection keys "Confirm" and "Cancel". When the terminal device 100 detects that the user clicks "Confirm", as shown in (c) of Figure 7 , the terminal device 100 can delete the device name "BBB" from the selected device list 60 and can add the device name "BBB" to the candidate device list 61 to facilitate the user to select again.
[0228] When, in addition to the terminal devices in the candidate device column 61, the user also wants to perform cross-device voiceprint registration on other terminal devices, the user can add devices by adding a control 62. Among them, the added terminal devices can be terminal devices that have enabled the cross-device voiceprint registration function, or can be terminal devices that have not enabled the cross-device voiceprint registration function. Specifically, when the added terminal device is a terminal device that has enabled the cross-device voiceprint registration function, the terminal device 100 can directly add the device name and / or device identifier of the added terminal device to the selected device column 60. When the added terminal device is a terminal device that has not enabled the cross-device voiceprint registration function, the terminal device 100 can send an activation request to the added terminal device, and a relevant pop-up window can pop up on the added terminal device to prompt the user to activate the cross-voiceprint registration function of the added terminal device. After the user activates the cross-device voiceprint registration function of the added terminal device, the terminal device 100 can add the device name and / or device identifier of the added terminal device to the selected device column 60.
[0229] Please refer to Figure 8 , Figure 8 which shows a schematic flow diagram of a cross-device voiceprint registration method provided in this embodiment. As Figure 8 shown, the method may include:
[0230] S801. The user inputs a first registration voice into a first terminal device.
[0231] S802. The first terminal device generates a voiceprint template corresponding to the first terminal device according to the first registration voice.
[0232] S803. The first terminal device sends the first registration voice to a second terminal device.
[0233] It can be understood that after determining the terminal devices for cross-device voiceprint registration, the user can input the first registration voice into the first terminal device. The first registration voice refers to the voice received by the first terminal device, that is, the first registration voice is attached with the channel information corresponding to the first terminal device. After receiving the first registration voice, the first terminal device can process the first registration voice according to the voiceprint recognition model corresponding to the first terminal device to obtain the voiceprint template corresponding to the first registration voice, so as to complete the voiceprint registration of the first terminal device. At the same time, the first terminal device can also send the first registration voice to each second terminal device, and each second terminal device can obtain the second registration voice corresponding to each second terminal device according to the first registration voice to perform voiceprint registration on each second terminal device.
[0234] It should be noted that the first terminal device can also obtain the first registered voice from the daily voice interaction between the user and the first terminal device, so as to perform voiceprint registration for each terminal device in a self-learning and registration-free manner, thereby simplifying the operation process of voiceprint registration and improving the user experience.
[0235] Specifically, when the user interacts with the first terminal device by voice, the first terminal device can screen out the user's voice through self-learning, for example, screen out the user's voice through clustering, and can select the voice with good quality from the screened-out voices as the first registered voice to perform voiceprint registration for each terminal device. Here, the voice with good quality can be selected by evaluating the signal-to-noise ratio and / or voice energy level of each voice.
[0236] S804. The second terminal device performs conversion processing on the first registered voice according to the first channel model of the first terminal device relative to the original voice signal and the second channel model of the second terminal device relative to the original voice signal; or performs conversion processing on the first registered voice according to the channel model between the first terminal device and the second terminal device to obtain a second registered voice.
[0237] S805. The second terminal device generates a voiceprint template corresponding to the second terminal device according to the second registered voice.
[0238] Exemplarily, when the channel model is the channel model of the terminal device relative to the original voice signal, after receiving the first registered voice sent by the first terminal device, the second terminal device can determine the first channel model of the first terminal device relative to the original voice signal according to the device name and / or device identifier of the first terminal device, and determine the second channel model of the second terminal device relative to the original voice signal according to the device name and / or device identifier of the second terminal device. Then, the channel information corresponding to the first terminal device in the first registered voice can be removed through the first channel model to obtain the original voice without channel information. Subsequently, the channel information corresponding to the second terminal device can be added to the original voice through the second channel model to obtain a second registered voice containing the channel information corresponding to the second terminal device, so that a voiceprint template corresponding to the second terminal device can be generated according to the second registered voice.
[0239] Specifically, after receiving the first registered voice, the second terminal device can perform frequency-domain conversion on the first registered voice to obtain the frequency-domain signal St1' corresponding to the first registered voice. For example, the first registered voice can be subjected to frequency-domain conversion through fast Fourier transform (FFT) to obtain the frequency-domain signal St1' corresponding to the first registered voice. Then, the frequency-domain signal St1' can be converted and processed through the first channel model of the first terminal device relative to the original voice signal, that is, the channel information corresponding to the first terminal device in the frequency-domain signal St1' can be removed according to the mapping relationship between the voice signal corresponding to the first terminal device and the original voice signal to obtain the original frequency-domain signal S'. Subsequently, the original frequency-domain signal S' can be converted and processed through the second channel model of the second terminal device relative to the original voice signal, that is, the channel information corresponding to the second terminal device can be added to the original frequency-domain signal S' according to the mapping relationship between the voice signal corresponding to the second channel model and the original voice signal to obtain the frequency-domain signal St2' corresponding to the second terminal device. Finally, the frequency-domain signal St2' can be subjected to inverse FFT to obtain the second registered voice corresponding to the second terminal device.
[0240] Exemplarily, when the channel model is the channel model between terminal devices, after the second terminal device receives the first registered voice sent by the first terminal device, it can determine the channel model between the first terminal device and the second terminal device according to the device name and / or device identifier of the first terminal device and the device name and / or device identifier of the second terminal device, and can directly convert the first registered voice into the second registered voice corresponding to the second terminal device according to this channel model, reducing the number of conversions of the registered voice to reduce information loss during the conversion process of the registered voice, thereby improving the accuracy of the voiceprint template generated based on the second registered voice.
[0241] Specifically, after receiving the first registered voice, the second terminal device can perform frequency-domain conversion on the first registered voice to obtain the frequency-domain signal St1' corresponding to the first registered voice. Then, the frequency-domain signal St1' can be directly converted into the frequency-domain signal St2' corresponding to the second terminal device according to the mapping relationship between the voice signal corresponding to the first terminal device and the voice signal corresponding to the second terminal device. Finally, the frequency-domain signal St2' can be subjected to inverse FFT to obtain the second registered voice corresponding to the second terminal device.
[0242] Exemplarily, the first terminal device can also directly send the voice features corresponding to the first registered voice to each second terminal device, and each second terminal device can obtain the second registered voice corresponding to each second terminal device based on the voice features corresponding to the first registered voice to perform voiceprint registration on each second terminal device. That is, the first terminal device can also directly send the voice features after frequency-domain conversion of the first registered voice to each second terminal device, which can enable each second terminal device to omit the frequency-domain conversion process and improve the processing performance of each second terminal device.
[0243] It should be understood that the conversion process of converting the first registered voice to the second registered voice in this embodiment can also be executed by the first terminal device.
[0244] Exemplarily, after the first terminal device obtains the first registered voice input by the user, it can remove the channel information of the first registered voice through the first channel model of the first terminal device relative to the original voice signal to obtain the original voice without channel information. Then, it can add the channel information corresponding to the second terminal device to the original voice through the second channel model of the second terminal device relative to the original voice signal to obtain the second registered voice containing the channel information corresponding to the second terminal device, and can send the second registered voice to the second terminal device.
[0245] Exemplarily, after the first terminal device obtains the first registered voice input by the user, it can directly convert the first registered voice into the second registered voice corresponding to the second terminal device through the channel model between the first terminal device and the second terminal device, and can send the second registered voice to the second terminal device.
[0246] In one example, when the channel model is the channel model of the terminal device relative to the original voice signal, to reduce the computational load of the first terminal device and / or the second terminal device, the conversion process of the registered voice can be decomposed to the first terminal device and the second terminal device. That is, after the first terminal device obtains the first registered voice input by the user, it can remove the channel information of the first registered voice through the first channel model of the first terminal device relative to the original voice signal to obtain the original voice, and can send the original voice to the second terminal device. After receiving the original voice, the second terminal device can add the channel information of the original voice through the second channel model of the second terminal device relative to the original voice signal to obtain the second registered voice containing the channel information corresponding to the second terminal device.
[0247] After generating the voiceprint template corresponding to the second terminal device, the user can directly use the voiceprint recognition function of the second terminal device without having to perform voiceprint registration in the second terminal device. Specifically, the user can directly input the authentication voice to the second terminal device. After receiving the authentication voice, the second terminal device can obtain the feature vector (i.e., the authentication template) corresponding to the authentication voice through the voiceprint recognition model, and calculate the similarity between the obtained feature vector and the voiceprint template in the second terminal device, so as to identify the user's identity according to the similarity and a preset first similarity threshold. When the similarity is greater than or equal to the first similarity threshold, it can be determined that the authentication voice comes from the registered user; when the similarity is less than the first similarity threshold, it can be determined that the authentication voice may not come from the registered user. Among them, the first similarity threshold can be specifically set according to the actual situation. For example, the first similarity threshold can be set to 70%.
[0248] This embodiment does not limit the algorithm for calculating the similarity between the feature vector corresponding to the authentication voice and the voiceprint template. Exemplarily, any one of algorithms such as cosine distance (CDS), linear discriminant analysis (LDA), and probabilistic linear discriminant analysis (PLDA) can be used to calculate the similarity between the feature vector corresponding to the authentication voice and the voiceprint template.
[0249] It should be noted that during the process of voiceprint recognition by the terminal device (including the above-mentioned first terminal device and second terminal device), when the terminal device determines that the similarity between the authentication voice and the voiceprint template in the terminal device is greater than or equal to a preset second similarity threshold, the terminal device can determine the authentication voice as a high-quality voice sample, and can use the high-quality voice sample to perform incremental learning on the voiceprint recognition model corresponding to the terminal device to update the voiceprint recognition model corresponding to the terminal device. At the same time, the terminal device can also generate high-quality voice samples corresponding to other terminal devices according to the authentication voice, so that other terminal devices can perform incremental learning on the voiceprint recognition models corresponding to other terminal devices according to the high-quality voice samples to update the voiceprint recognition models corresponding to other terminal devices. That is, this embodiment can obtain high-quality authentication voices during the user's daily use to update the voiceprint recognition models corresponding to each terminal device, improve the matching degree between the voiceprint recognition models corresponding to each terminal device and the actual use scenario, improve the robustness of voiceprint recognition in each terminal device, and thus improve the accuracy of voiceprint recognition in each terminal device.
[0250] Among them, the second similarity threshold can be specifically set according to the actual situation, and the second similarity threshold can be greater than or equal to the first similarity threshold. For example, the second similarity threshold can be set to 90%.
[0251] This embodiment does not limit the algorithm for incremental learning using high-quality voice samples. Exemplarily, the high-quality voice samples can be jointly trained with the original training data corresponding to each terminal device in a weighted manner to update the voiceprint recognition model corresponding to each terminal device.
[0252] This embodiment can perform conversion processing on the registration voice obtained by the first terminal device to generate the registration voice corresponding to each second terminal device, so as to perform voiceprint registration on each second terminal device, achieving the purpose of voiceprint registration for multiple terminal devices with a single input of the registration voice, reducing the number of voice inputs for voiceprint registration of multiple terminal devices, and improving the user experience.
[0253]
Embodiment 2
[0254] The method provided in the above Embodiment 1 needs to perform conversion processing on the registration voice through the first terminal device and / or the second terminal device, greatly increasing the computing amount of the first terminal device and / or the second terminal device, and affecting the usage performance of the first terminal device and / or the second terminal device.
[0255] Please refer to Figure 9 , Figure 9 which shows a schematic diagram of the application scenario of the cross-device voiceprint registration method provided in the second embodiment of the present application. This application scenario can include multiple terminal devices 100 and a server 90. Among them, the server 90 can be a cloud server or a control center, etc., to perform conversion processing on the registration voice through the server, reducing the computing amount of the first terminal device and / or the second terminal device, and ensuring the usage performance of the first terminal device and / or the second terminal device.
[0256] Among them, the server 90 can communicate with each terminal device 100 respectively. The server 90 or the storage device communicatively connected to the server 90 can store the channel model of each terminal device 100 relative to the original voice signal, or store the channel model between any two terminal devices 100.
[0257] It should be understood that when a user performs voiceprint registration on a first terminal device, the first terminal device may send the first registration voice input by the user to the server 90. The server 90 may perform conversion processing on the first registration voice according to the first channel model of the first terminal device relative to the original voice signal and the second channel models of the second terminal devices relative to the original voice signal to obtain the second registration voices corresponding to the second terminal devices; alternatively, the server 90 may perform conversion processing on the first registration voice according to the channel model between the first terminal device and the second terminal devices to obtain the second registration voices corresponding to the second terminal devices. Thus, voiceprint templates corresponding to the second terminal devices may be generated according to the second registration voices, so as to reduce the number of voice inputs for voiceprint registration of multiple terminal devices.
[0258] Please refer to Figure 10 , Figure 10 which shows a schematic flowchart of a cross-device voiceprint registration method provided in this embodiment. As Figure 10 shown, the method may include:
[0259] S1001. The user inputs a first registration voice to the first terminal device.
[0260] S1002. The first terminal device generates a voiceprint template corresponding to the first terminal device according to the first registration voice.
[0261] S1003. The first terminal device sends the first registration voice to the server.
[0262] It can be understood that after determining the terminal devices for cross-device voiceprint registration, the user may input a first registration voice to the first terminal device. The first registration voice refers to the voice received by the first terminal device, that is, the channel information corresponding to the first terminal device is attached to the first registration voice. After receiving the first registration voice, the first terminal device may use the voiceprint recognition model corresponding to the first terminal device to process the first registration voice to obtain the voiceprint template corresponding to the first registration voice, so as to complete the voiceprint registration of the first terminal device. At the same time, the first terminal device may also send the first registration voice to the server 90, so that the server 90 can obtain the registration voices corresponding to other terminal devices according to the first registration voice to perform voiceprint registration on other terminal devices.
[0263] S1004. The server performs conversion processing on the first registration voice according to the first channel model of the first terminal device relative to the original voice signal and the second channel models of the second terminal devices relative to the original voice signal; or performs conversion processing on the first registration voice according to the channel model between the first terminal device and the second terminal devices to obtain the second registration voice.
[0264] S1005. The server sends the second registered voice to the second terminal device.
[0265] Exemplarily, when the channel model is the channel model of the terminal device relative to the original voice signal, after receiving the first registered voice sent by the first terminal device, the server can determine the first channel model of the first terminal device relative to the original voice signal according to the device name and / or device identifier of the first terminal device, and determine the second channel model of the second terminal device relative to the original voice signal according to the device name and / or device identifier of the second terminal device. Then, the channel information corresponding to the first terminal device in the first registered voice can be removed according to the first channel model to obtain the original voice without channel information. Subsequently, the channel information corresponding to the second terminal device can be added to the original voice according to the second channel model to obtain the second registered voice containing the channel information corresponding to each second terminal device, and each second registered voice can be sent to the corresponding second terminal device.
[0266] Exemplarily, when the channel model is the channel model between terminal devices, after receiving the first registered voice sent by the first terminal device, the server can determine the channel model between the first terminal device and the second terminal device according to the device name and / or device identifier of the first terminal device and the device name and / or device identifier of the second terminal device, and directly convert the first registered voice into the second registered voice corresponding to the second terminal device according to this channel model, so as to reduce the conversion times of the registered voice, reduce the information loss in the conversion process of the registered voice, and thus improve the accuracy of the voiceprint template generated based on the second registered voice.
[0267] S1006. The second terminal device generates a voiceprint template corresponding to the second terminal device according to the second registered voice.
[0268] Here, after each second terminal device receives the corresponding second registered voice, it can process each second registered voice according to the voiceprint recognition model corresponding to each second terminal device to obtain the voiceprint template corresponding to the second registered voice, that is, extract the voiceprint features from each second registered voice to obtain the voiceprint template corresponding to each second terminal device. For example, after the second terminal device A receives the second registered voice A, it can extract the voiceprint features of the second registered voice A according to the voiceprint recognition model A corresponding to the second terminal device A to obtain the voiceprint template A corresponding to the second terminal device A; after the second terminal device B receives the second registered voice B, it can extract the voiceprint features of the second registered voice B according to the voiceprint recognition model B corresponding to the second terminal device B to obtain the voiceprint template B corresponding to the second terminal device B, and so on. Among them, the second registered voice A is the voice added with the channel information corresponding to the second terminal device A, and the second registered voice B is the voice added with the channel information corresponding to the second terminal device B.
[0269] In one example, to further reduce the computational load of each second terminal device, the server 90 may directly generate the voiceprint templates corresponding to each second terminal device. The voiceprint recognition models corresponding to each terminal device may also be stored in the server 90 or a storage device communicatively connected to the server 90. When the server 90 obtains the second registration voices corresponding to each second terminal device, the server 90 may obtain the voiceprint recognition models corresponding to each second terminal device according to the device names and / or device identifiers corresponding to each second terminal device, and may process the second registration voices according to the voiceprint recognition models corresponding to each second terminal device to obtain the voiceprint templates corresponding to each second terminal device, and send the voiceprint templates corresponding to each second terminal device to the corresponding second terminal devices respectively.
[0270] It should be noted that during the process of voiceprint recognition by the terminal device (including the above-mentioned first terminal device and second terminal device), when the terminal device determines that the similarity between the authentication voice and the voiceprint template in the terminal device is greater than or equal to a preset second similarity threshold, the terminal device may send the authentication voice to the server 90. The server 90 may perform incremental learning on the voiceprint recognition model corresponding to the terminal device by using the authentication voice to update the voiceprint recognition model corresponding to the terminal device, and may send the updated voiceprint recognition model to the terminal device. At the same time, the server 90 may also generate training voices of other terminal devices according to the authentication voice to perform incremental learning on the voiceprint recognition models corresponding to other terminal devices to update the voiceprint recognition models corresponding to other terminal devices, and may send the updated voiceprint recognition models to the corresponding terminal devices respectively. That is, in this embodiment, the server may obtain high-quality authentication voices during the daily use of the user to update the voiceprint recognition models corresponding to each terminal device, so as to improve the matching degree of the voiceprint recognition models corresponding to each terminal device with the actual use scenario, improve the robustness of voiceprint recognition in each terminal device, and thus improve the accuracy of voiceprint recognition in each terminal device.
[0271] In this embodiment, by performing conversion processing on the registration voice by the server to generate the registration voices corresponding to other terminal devices, it can not only achieve the purpose of voiceprint registration for multiple terminal devices with one input of the registration voice, reduce the number of voice inputs for voiceprint registration of multiple terminal devices, and improve the user experience. At the same time, it can also reduce the computational load of each terminal device and ensure the performance of each terminal device.
[0272] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0273] Figure 11The structural schematic diagram of an electronic device provided by an embodiment of the present application. As Figure 11 shown, the electronic device 11 of this embodiment includes: at least one processor 1100 ( Figure 11 only one is shown in the figure), a memory 1101, and a computer program 1102 stored in the memory 1101 and executable on the at least one processor 1100. When the processor 1100 executes the computer program 1102, the electronic device 11 implements the steps in any of the above cross-device voiceprint registration method embodiments.
[0274] The electronic device 11 may be a terminal device or a server. The electronic device 11 may include, but is not limited to, a processor 1100 and a memory 1101. Those skilled in the art can understand that Figure 11 this is only an example of the electronic device 11 and does not constitute a limitation on the electronic device 11. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0275] The processor 1100 may be a central processing unit (CPU), and the processor 1100 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0276] In some embodiments, the memory 1101 may be an internal storage unit of the electronic device 11, such as a hard disk or memory of the electronic device 11. In other embodiments, the memory 1101 may also be an external storage device of the electronic device 11, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 11. Further, the memory 1101 may also include both the internal storage unit and the external storage device of the electronic device 11. The memory 1101 is used to store an operating system, application programs, a boot loader, data, and other programs, such as program codes of the computer program. The memory 1101 may also be used to temporarily store data that has been output or will be output.
[0277] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, which when executed by a computer, causes the computer to implement the steps in the above method embodiments.
[0278] An embodiment of the present application provides a computer program product, which when running on an electronic device, causes the electronic device to implement the steps in the above method embodiments.
[0279] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable storage medium may at least include: any entity or device capable of carrying the computer program code to the device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable storage medium may not be an electrical carrier signal and a telecommunication signal.
[0280] In the above embodiments, the descriptions of the various embodiments have their respective emphases. For parts not described or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0281] Those of ordinary skill in the art will appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0282] In the embodiments provided in this application, it should be understood that the disclosed device / electronic device and method can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0283] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0284] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.
Claims
1. A cross-device voiceprint registration method, characterized in that, Applied to a second terminal device, the method includes: Obtain a first registered voice corresponding to a first terminal device; the first registered voice is a voice input by a user in the first terminal device; Perform conversion processing on the first registered voice to obtain a second registered voice corresponding to the second terminal device; the second registered voice is a voice after removing channel information corresponding to the first terminal device in the first registered voice and adding channel information corresponding to the second terminal device; Generate a voiceprint template corresponding to the second terminal device according to the second registered voice.
2. The method according to claim 1, characterized in that, The performing conversion processing on the first registered voice to obtain a second registered voice corresponding to the second terminal device includes: Perform conversion processing on the first registered voice through a first channel model corresponding to the first terminal device to obtain an original voice corresponding to the first registered voice, where the first channel model is used to represent the mapping relationship between the voice corresponding to the first terminal device and the original voice; Perform conversion processing on the original voice through a second channel model corresponding to the second terminal device to obtain a second registered voice corresponding to the second terminal device, where the second channel model is used to represent the mapping relationship between the voice corresponding to the second terminal device and the original voice.
3. The method according to claim 1, characterized in that, The performing conversion processing on the first registered voice to obtain a second registered voice corresponding to the second terminal device includes: Perform conversion processing on the first registered voice through a third channel model to obtain a second registered voice corresponding to the second terminal device, where the third channel model is used to represent the mapping relationship between the voice corresponding to the first terminal device and the voice corresponding to the second terminal device.
4. The method according to claim 2, wherein The first channel model and the second channel model are channel models constructed based on frequency response curves or are channel models constructed based on spectral features.
5. The method according to claim 3, characterized in that The third channel model is a channel model constructed based on a frequency response curve or is a channel model constructed based on spectral features.
6. The method according to any one of claims 1-5, characterized in that, The generating a voiceprint template corresponding to the second terminal device according to the second registered voice includes: Generate a voiceprint template corresponding to the second terminal device according to a voiceprint recognition model corresponding to the second terminal device and the second registered voice, where the voiceprint recognition model corresponding to the second terminal device is a voiceprint recognition model trained based on training voices obtained by the second terminal device.
7. The method according to any one of claims 1 to 6, characterized in that, After the generating a voiceprint template corresponding to the second terminal device according to the second registered voice, it further includes: Obtain an authentication voice corresponding to the second terminal device; Generate an authentication template corresponding to the authentication voice according to the voiceprint recognition model corresponding to the second terminal device and the authentication voice; Determine the similarity between the authentication template and the voiceprint template; When the similarity is greater than a preset similarity threshold, update the voiceprint recognition model corresponding to the second terminal device according to the authentication voice.
8. The method according to claim 7, wherein When the similarity is greater than a preset similarity threshold, the method further includes: Perform conversion processing on the authenticated voice to obtain the training voice corresponding to the first terminal device, and send the training voice to the first terminal device, where the training voice is used to update the voiceprint recognition model corresponding to the first terminal device.
9. A cross-device voiceprint registration method, characterized in that Applied to a first terminal device or a server, the method includes: Obtain a first registration voice corresponding to the first terminal device; the first registration voice is a voice input by a user in the first terminal device; Perform conversion processing on the first registration voice to obtain a second registration voice corresponding to the second terminal device; the second registration voice is a voice after removing the channel information corresponding to the first terminal device in the first registration voice and adding the channel information corresponding to the second terminal device; Send the second registration voice to the second terminal device, where the second registration voice is used to generate a voiceprint template corresponding to the second terminal device.
10. The method according to claim 9, characterized in that The performing conversion processing on the first registration voice to obtain a second registration voice corresponding to the second terminal device includes: Perform conversion processing on the first registration voice through a first channel model corresponding to the first terminal device to obtain the original voice corresponding to the first registration voice, where the first channel model is used to represent the mapping relationship between the voice corresponding to the first terminal device and the original voice; Perform conversion processing on the original voice through a second channel model corresponding to the second terminal device to obtain the second registration voice corresponding to the second terminal device, where the second channel model is used to represent the mapping relationship between the voice corresponding to the second terminal device and the original voice.
11. The method according to claim 9, characterized in that, The performing conversion processing on the first registration voice to obtain a second registration voice corresponding to the second terminal device includes: Perform conversion processing on the first registration voice through a third channel model to obtain the second registration voice corresponding to the second terminal device, where the third channel model is used to represent the mapping relationship between the voice corresponding to the first terminal device and the voice corresponding to the second terminal device.
12. A cross-device voiceprint registration method, characterized in that, Includes: The first terminal device obtains a first registration voice corresponding to the first terminal device; The first registration voice is a voice input by a user in the first terminal device; The first terminal device performs conversion processing on the first registration voice to obtain the first original voice corresponding to the first registration voice, and sends the first original voice to the second terminal device; The second terminal device receives the first original voice from the first terminal device, and performs conversion processing on the first original voice to obtain the second registration voice corresponding to the second terminal device; the second registration voice is a voice after removing the channel information corresponding to the first terminal device in the first registration voice and adding the channel information corresponding to the second terminal device; The second terminal device generates a voiceprint template corresponding to the second terminal device according to the second registration voice.
13. The method according to claim 12, wherein The first terminal device obtaining the first registration voice corresponding to the first terminal device includes: The first terminal device acquires the interactive voice between the first terminal device and the user, and acquires the target voice in the interactive voice, where the target voice is the voice corresponding to the user; The first terminal device acquires the first registration voice corresponding to the first terminal device from the target voice according to the signal-to-noise ratio and / or voice energy level corresponding to the target voice.
14. The method according to claim 12 or 13, characterized in that After the second terminal device generates the voiceprint template corresponding to the second terminal device according to the second registration voice, it includes: The second terminal device acquires the authentication voice corresponding to the second terminal device, generates an authentication template corresponding to the authentication voice according to the voiceprint recognition model corresponding to the second terminal device and the authentication voice, and determines the similarity between the authentication template and the voiceprint template; When the similarity is greater than a preset similarity threshold, the second terminal device updates the voiceprint recognition model corresponding to the second terminal device according to the authentication voice, and performs conversion processing on the authentication voice according to the second channel model corresponding to the second terminal device to obtain the second original voice corresponding to the authentication voice, and sends the second original voice to the first terminal device; The first terminal device receives the second original voice from the second terminal device, performs conversion processing on the second original voice according to the first channel model corresponding to the first terminal device to obtain the training voice corresponding to the first terminal device, and updates the voiceprint recognition model corresponding to the first terminal device according to the training voice corresponding to the first terminal device.
15. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the electronic device implements the cross-device voiceprint registration method according to any one of claims 1 to 8, or any one of claims 9 to 11.
16. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a computer, the computer implements the cross-device voiceprint registration method according to any one of claims 1 to 8, or any one of claims 9 to 11.
Citation Information
Patent Citations
Cross-device voiceprint recognition method and system
CN109378006A