A voiceprint recognition and registration device and a cross-device voiceprint recognition method
By recording biometric voiceprints after noise removal on the registration device and matching them after noise removal on the recognition device, the problem of cumbersome registration required for voiceprint recognition devices in existing technologies is solved, achieving high-accuracy voiceprint recognition across devices and improving user experience.
Patent Information
- Application Number
- CN202080001170.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-19
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2040-05-19
AI Technical Summary
In existing technologies, voiceprint recognition devices require a complicated registration process on each device, and the voiceprint mapping model is device-dependent, making it difficult to guarantee recognition accuracy, especially when the device channel changes, resulting in a high error rate.
By recording the user's biometric voiceprint information after noise removal from the registered device, and matching it after noise removal from the identified device, cross-device voiceprint recognition is achieved by using machine learning and filter technology to extract and share biometric voiceprint information across devices.
Users only need to register their voiceprint on one device to be recognized on multiple devices, which improves the user experience, has a high recognition accuracy, and eliminates the influence of device noise.
Smart Images

Figure CN114026637B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of voiceprint recognition, in particular to a voiceprint recognition and registration device and a cross-device voiceprint recognition method. BACKGROUND
[0002] Voiceprint recognition is a kind of biometric technology, which converts a voice signal into a digital signal and uses a computer to identify the identity of a speaker. Currently, when a device uses a voiceprint recognition function, the voiceprint of a user needs to be registered in advance. With the continuous development of smart device technology and the increasing number of voiceprint devices, the registration of voiceprints has become a complicated process.
[0003] In the current technology, a voiceprint mapping model can be established between different devices, a voiceprint can be extracted from a voice collected by a first device, voiceprint registration can be performed, a voiceprint feature can be extracted from a voice instruction collected by a second device, and the voiceprint feature can be mapped to the voiceprint registered by the first device based on the established voiceprint mapping model, so that the voiceprint does not need to be registered on the second device.
[0004] In the above scheme, the saved voiceprint mapping model is related to the device, and two-to-two mapping of all devices needs to be performed. The number of mapping relationships is too large. For example, when there are Q devices in an environment, the number of two-to-two voiceprint mapping relationship models is Qx(Q-1). If there is a change in the channel of an individual device, the corresponding voiceprint mapping model group is completely incorrect, and the recognition accuracy is difficult to guarantee. SUMMARY
[0005] Embodiments of the present application provide a voiceprint recognition device in a recognition device, a voiceprint registration device in a registration device, and a cross-device voiceprint recognition method. The method is applied to a cross-device voiceprint recognition system, which includes a recognition device and a registration device. The registration device registers the voice of a user. The registration device eliminates noise related to the registration device in the collected voice to obtain a biological voiceprint of the user, records the biological voiceprint under the voiceprint identifier of the user to obtain registration voiceprint information. It can be understood that the biological voiceprint of the user has been "recorded in the register". The biological voiceprint of the user that has been recorded is referred to as the registration voiceprint information of the user. The registration voiceprint information provides a basis for voiceprint recognition. The recognition device is a device that performs voiceprint recognition on the voice of a user. After the recognition device receives the voice input by the user, it eliminates noise related to the recognition device in the voice to obtain a biological voiceprint of the user (referred to as target voiceprint information). The voiceprint recognition function is realized through the registration voiceprint information shared by the registration device, that is, the target voiceprint information is matched with the registration voiceprint information, and the user identity is indicated through the matching result.
[0006] In a first aspect, an embodiment of the present application provides a voiceprint recognition apparatus in an identification device, the identification device can be a terminal, the voiceprint recognition apparatus can be a processor, a chip or a chip system in the terminal; or the voiceprint recognition apparatus is a terminal, and the terminal is a component of the identification device; the apparatus comprises a processor, a voice input and output device connected to the processor, and a transceiver; the voice input and output device is configured to receive a voice input by a user; the processor is configured to eliminate noise related to the identification device in the received voice to obtain target voiceprint information, that is, the target voiceprint information is a biological voiceprint after the influence of the identification device on the voice is eliminated, and the biological voiceprint of the user is equivalent to the voiceprint of a speaker heard by a listener face to face; then, the transceiver is configured to obtain registered voiceprint information of the user, and the registered voiceprint information comprises at least one biological voiceprint template obtained by eliminating noise related to a registered device; the processor is further configured to match the target voiceprint information with the registered voiceprint information to obtain a matching result, and the matching result is used to indicate the identity of the user; in the embodiment of the present application, the registered voiceprint information is biological voiceprint information of the user extracted from the voice after eliminating noise related to the registered device itself, and since the influence of the registered device on the biological voiceprint is eliminated, the registered voiceprint information can be shared by other identification devices, and the user only needs to register the voiceprint on one device (the registered device) to realize the voiceprint recognition function on multiple devices, the identification device does not need to register the voiceprint, and the complicated voiceprint registration process is saved, cross-device voiceprint recognition is realized, the user experience can be greatly improved, and the registered voiceprint information and the target voiceprint information are both biological voiceprints after eliminating the noise related to the device, so that the influence of the device itself on the voiceprint is eliminated, and the identification device has a high accuracy regardless of the type of terminal.
[0007] In a possible implementation, the apparatus further comprises a transceiver connected to the processor; the transceiver is configured to receive the registered voiceprint information and a voiceprint identifier of the registered voiceprint information sent by a storage device or a registered device, and the voiceprint identifier is used to indicate the user; in the embodiment, the registered voiceprint information can be stored in the storage device, thereby providing a basis for sharing the registered voiceprint information by other devices, and when the user needs to use the identification device for voiceprint recognition, the user does not need to register the voiceprint on the identification device, but can share the registered voiceprint information in the storage device or the registered device.
[0008] In a possible implementation, the transceiver is further configured to send a voiceprint identifier of the user to the storage device or the registered device to request the registered voiceprint information of the user; the biological voiceprint identifier of the user has a corresponding relationship with the registered voiceprint information, the voiceprint identifier of the user is sent to the storage device or the registered device, and the storage device or the registered device sends the registered voiceprint information corresponding to the voiceprint identifier to the voiceprint recognition apparatus.
[0009] In a possible implementation, the processor is specifically configured to extract a voice input voiceprint through a voiceprint extraction model, and eliminate noise related to the identification device in the voice to obtain target voiceprint information; in the embodiment of the application, the voiceprint extraction model is pre-trained through machine learning, and the biological voiceprint of the user in the voice is directly extracted through the pre-trained voiceprint extraction model, which has good robustness.
[0010] In a possible implementation, the voiceprint extraction model is obtained by learning corpus collected by multiple devices; the multiple devices can be different types of terminals, including but not limited to terminals in a smart home, lighting systems, smart speakers, robots, vehicle terminals, user devices, mobile phones, tablet computers, personal computers, virtual reality terminal devices, augmented reality terminal devices, and the like; in the embodiment of the application, the voiceprint extraction model is obtained by learning corpus collected by multiple different devices, and the voiceprint information output by the voiceprint extraction model can eliminate the influence of different devices on the voice to obtain the biological voiceprint of the user.
[0011] In a possible implementation, the processor is further configured to perform frequency response compensation on the voice through a filter, and the frequency response compensation is used to eliminate noise related to the identification device in the voice to obtain target voiceprint information; in the embodiment of the application, the voice signal is processed through the filter, the highlighted signal is attenuated, and the weak signal is enhanced, so as to compensate the frequency response, so as to achieve the purpose of removing the device-related noise. The filter filtering mode is to directly filter the noise signal, so that the output signal is the biological voiceprint signal of the user, which is simple and fast.
[0012] In a possible implementation, the processor is further configured to obtain user data associated with the user identity, and perform an operation corresponding to the user data; in the embodiment of the application, the processor performs an operation corresponding to the historical information, without the need for the user to perform habitual operations every time, thereby improving the user experience.
[0013] In a second aspect, an embodiment of the present application provides a voiceprint registration device in a registration device, which can be a terminal, and the voiceprint registration device can be a processor, a chip or a chip system in the terminal; or the voiceprint registration device is a terminal, and the terminal is a component of an identification device; the device includes a processor and a voice input and output device connected to the processor; the voice input and output device is configured to receive a voice input by a user; the processor is configured to eliminate noise in the voice related to the registration device to obtain registration voiceprint information of the user, and the registration voiceprint information includes a biological voiceprint of the user; the registration voiceprint information is used to provide a basis for voiceprint identification for other devices, and the registration voiceprint information can be registration voiceprint information shared by multiple identification devices; in the embodiment of the present application, the registration device extracts the voice of the user received, eliminates the noise in the voice related to the device, and obtains a “clean” biological voiceprint of the user (the biological voiceprint is most close to a voiceprint when the user speaks face to face), and because the influence of the device on the voice is eliminated, the extracted biological voiceprint information can be registration voiceprint information shared by multiple devices, cross-device voiceprint identification is realized, and user experience can be greatly improved.
[0014] In a possible implementation, the device further includes a transceiver connected to the processor; the transceiver is configured to send the registration voiceprint information and a voiceprint identifier corresponding to the registration voiceprint information to a storage device or an identification device, and the voiceprint identifier is used to indicate the user; in the embodiment of the present application, the transceiver sends the registration voiceprint information and the voiceprint identifier corresponding to the registration voiceprint information to the storage device or the identification device, takes the storage device or the identification device as a storage center of the registration voiceprint information and the voiceprint identifier, and provides a basis for sharing of the registration voiceprint information.
[0015] In a possible implementation, the processor is specifically configured to input the voice into a voiceprint extraction model, and eliminate the noise in the voice related to the registration device by the voiceprint extraction model to obtain the registration voiceprint information; in the embodiment of the present application, the voiceprint extraction model is pre-trained in a manner of machine learning, the biological voiceprint of the user in the voice is directly extracted by the voiceprint extraction model that has been trained, and the voiceprint extraction model has good robustness.
[0016] In a possible implementation, the voiceprint extraction model is obtained by learning a plurality of devices to collect a corpus; in the embodiment of the present application, the voiceprint information output by the voiceprint extraction model can eliminate the influence of different devices on the voice to obtain the biological voiceprint of the user.
[0017] In a possible implementation, the processor is further configured to perform frequency response compensation on the voice by using a filter, and the frequency response compensation is used to eliminate noise related to the identification device in the voice to obtain the registered voiceprint information; in the embodiment of the application, the voice signal is processed by using the filter, the highlighted signal is attenuated, and the weaker signal is enhanced, so that the frequency response is compensated, and the purpose of removing the device-related noise is achieved. The filter filtering mode is to directly filter the noise signal, so that the output signal is the biological voiceprint signal of the user, which is simple and fast.
[0018] In a third aspect, the embodiment of the application provides a method for cross-device voiceprint recognition, which is applied to an identification device and can include the following steps: the identification device receives voice input by a user; then, noise related to the identification device in the voice is eliminated to obtain biological voiceprint of the user, which is referred to as target voiceprint information; registered voiceprint information of the user is obtained, and the registered voiceprint information includes at least one biological voiceprint template obtained by eliminating noise related to a registration device; the target voiceprint information is matched with the registered voiceprint information to obtain a matching result, and the matching result is used to indicate the identity of the user; in the embodiment of the application, the registered voiceprint information is biological voiceprint information of the user himself extracted from the voice after eliminating noise related to the registration device itself, and since the influence of the registration device on the biological voiceprint is eliminated, the registered voiceprint information can be shared by other identification devices. The user only needs to register the voiceprint on the registration device, and can realize the voiceprint recognition function on multiple devices, without the need to register the voiceprint on each identification device, thereby saving the complicated voiceprint registration process, realizing cross-device voiceprint recognition, greatly improving the user experience, and the registered voiceprint information and the target voiceprint information are both biological voiceprints separated from the device-related noise, so that the influence of the device itself on the voiceprint is eliminated, and the identification device has a high accuracy regardless of the type of terminal.
[0019] In an optional implementation, obtaining the registered voiceprint information of the user can include: receiving the registered voiceprint information and a voiceprint identifier corresponding to the registered voiceprint information sent by a storage device or a registration device, and the voiceprint identifier is used to indicate the user.
[0020] In an optional implementation, before receiving the registered voiceprint information and the voiceprint identifier corresponding to the registered voiceprint information sent by the storage device or the registration device, the method further includes: sending the voiceprint identifier to the storage device or the registration device to request the registered voiceprint information of the user.
[0021] In an optional implementation, eliminating the noise related to the recognition device in the voice to obtain the target voiceprint information can include inputting the voice into a voiceprint extraction model, and eliminating the noise related to the recognition device in the voice by the voiceprint extraction model to obtain the target voiceprint information. In the embodiment of the application, the voiceprint extraction model is pre-trained by machine learning, and the biological voiceprint of the user in the voice is directly extracted by the pre-trained voiceprint extraction model, which has good robustness.
[0022] In an optional implementation, the voiceprint extraction model is obtained by learning a plurality of devices collected corpus; in the embodiment of the application, the voiceprint information output by the voiceprint extraction model can eliminate the influence of different devices on the voice to obtain the biological voiceprint of the user.
[0023] In an optional implementation, after matching the target voiceprint information with the registered voiceprint information to obtain the matching result, the method further includes: obtaining user data associated with the user identity, the user data can be historical data of the user operation, and performing an operation corresponding to the user data; in the embodiment of the application, the recognition device obtains the user data of the user and performs an operation corresponding to the historical information, without the need for the user to perform habitual operations every time, thereby improving the user experience.
[0024] In a fourth aspect, the embodiment of the application provides a cross-device voiceprint recognition method applied to a registration device, which can include: receiving voice input by a user; eliminating noise related to the registration device in the voice to obtain registered voiceprint information; the registered voiceprint information includes a biological voiceprint of the user; in the embodiment of the application, the registration device extracts the voice input by the user, eliminates the noise related to the device in the voice, and obtains the "clean" biological voiceprint of the user (the biological voiceprint is closest to the voiceprint when the user talks face to face), because the influence of the device on the voice is reduced, the extracted biological voiceprint information can be shared as registered voiceprint information of a plurality of devices, cross-device voiceprint recognition is realized, and the user experience can be greatly improved.
[0025] In an optional implementation, the registration device sends the registered voiceprint information and a voiceprint identifier corresponding to the registered voiceprint information to a storage device or a recognition device; the registration device sends the registered voiceprint information and the voiceprint identifier corresponding to the registered voiceprint information to the storage device or the recognition device, and takes the storage device or the recognition device as a storage center of the registered voiceprint information and the voiceprint identifier, thereby providing a basis for sharing of the registered voiceprint information.
[0026] In an optional implementation, eliminating the noise in the voice related to the registration device to obtain the biometric voiceprint information can include: inputting the voice into a voiceprint extraction model, and eliminating the noise in the voice related to the registration device by the voiceprint extraction model to obtain the registration voiceprint information; in the embodiment of the application, the voiceprint extraction model is pre-trained in a manner of machine learning, and the biometric voiceprint of the user in the voice is directly extracted by the voiceprint extraction model that has been trained, which has good robustness.
[0027] In an optional implementation, the voiceprint extraction model is obtained by learning a plurality of devices collected corpus; in the embodiment of the application, the voiceprint information output by the voiceprint extraction model can eliminate the influence of different devices on the voice to obtain the biometric voiceprint of the user.
[0028] In an optional implementation, eliminating the noise in the voice related to the registration device to obtain the biometric voiceprint information can include: performing frequency response compensation on the voice by a filter, and the frequency response compensation is used to eliminate the noise in the voice related to the identification device to obtain the registration voiceprint information; in the embodiment of the application, the voice signal is processed by the filter, the highlighted signal is attenuated, and the weak signal is enhanced, so that the frequency response is compensated, thereby achieving the purpose of removing the device-related noise. The filter filtering mode is to directly filter the noise signal, so that the output signal is the biometric voiceprint signal of the user, which is simple and fast.
[0029] In a fifth aspect, the embodiment of the application provides a voiceprint recognition device, which has the functions of the identification device in the third aspect; the functions can be implemented by hardware, or the corresponding software is executed by hardware; the hardware or software includes one or more modules corresponding to the functions, and the device includes: a voice input and output module, configured to receive the voice input by the user; a processing module, configured to eliminate the noise in the voice received by the voice input and output module and related to the identification device to obtain target voiceprint information, and the target voiceprint information is the biometric voiceprint of the user; a transceiver module, configured to obtain the registration voiceprint information of the user, and the registration voiceprint information includes at least one biometric voiceprint template obtained by eliminating the noise related to the registration device; and the processing module is further configured to match the target voiceprint information obtained by the processing module with the registration voiceprint information obtained by the transceiver module 1303 to obtain a matching result, and the matching result is used to indicate the identity of the user.
[0030] In a sixth aspect, an embodiment of the present application provides a voiceprint registration device, which has the functions of the registration device in the fourth aspect above; the functions can be implemented by hardware, or by corresponding software executed by hardware; the hardware or software includes one or more modules corresponding to the functions, and the device includes: a voice input and output module, configured to receive voice input by a user; and a processing module, configured to eliminate noise in the voice received by the voice input and output module and related to the registration device to obtain registration voiceprint information of the user.
[0031] In a seventh aspect, an embodiment of the present application provides a cross-device voiceprint identification system, which includes a registration device and an identification device; the registration device receives first voice input by a user, and eliminates noise in the first voice and related to the identification device to obtain registration voiceprint information, which includes biological voiceprint information of the user and provides a basis for voiceprint identification by the identification device; the identification device receives second voice input by the user, and eliminates noise in the second voice and related to the identification device to obtain target voiceprint information, which is biological voiceprint information of the user; and the identification device matches the target voiceprint information with the registration voiceprint information to obtain a matching result, which is used to indicate the identity of the user. In the embodiment of the present application, the registration voiceprint information is biological voiceprint information of the user extracted from the voice after eliminating noise in the voice and related to the registration device itself. Since the influence of the registration device on the biological voiceprint information is eliminated, the registration voiceprint information can be shared by other voiceprint identification devices. The user only needs to register voiceprint on one device (the registration device), and can implement voiceprint identification on multiple devices, without the need to register voiceprint on each device, thereby saving the complicated voiceprint registration process, achieving cross-device voiceprint identification, greatly improving user experience, and eliminating the influence of the device on the voiceprint information, so that the user has a high accuracy rate when using any voiceprint device for voiceprint identification.
[0032] In an optional implementation, the system further includes a storage device; the storage device receives and stores the registration voiceprint information and voiceprint identification information corresponding to the registration voiceprint information sent by the registration device, and the voiceprint identification information is used to indicate the user; and the identification device receives the registration voiceprint information and the voiceprint information corresponding to the registration voiceprint information sent by the storage device. In the embodiment of the present application, the storage device stores the registration voiceprint information, thereby providing a basis for sharing the registration voiceprint information by other devices. When the user needs to use other devices (also referred to as voiceprint identification devices) for voiceprint identification, the user does not need to register voiceprint on the voiceprint identification devices, but can share the registration voiceprint information registered on the registration device.
[0033] In an eighth aspect, the embodiments provide a chip, comprising a processor and a memory, the memory being configured to store a program or instructions, which, when executed by the processor, cause the identification device to perform any one of the methods of the third aspect, or cause the registration device to perform any one of the methods of the fourth aspect.
[0034] In a ninth aspect, the embodiments provide a computer readable medium configured to store a computer program or instructions, which, when executed by a computer, cause the computer to perform any one of the methods of the third aspect, or cause the computer to perform any one of the methods of the fourth aspect. BRIEF DESCRIPTION OF DRAWINGS
[0035] Figure 1 A schematic diagram of an embodiment of a communication system in the embodiments of the present application;
[0036] Figure 2 A flowchart of an embodiment of cross-device voiceprint recognition in the embodiments of the present application;
[0037] Figure 3 A flowchart of an embodiment of voiceprint registration of a registration device in the embodiments of the present application;
[0038] Figure 4 A schematic diagram of an embodiment of storing registered voiceprint information in the embodiments of the present application;
[0039] Figure 5 A schematic diagram of another embodiment of storing registered voiceprint information in the embodiments of the present application;
[0040] Figure 6 A schematic diagram of another embodiment of storing registered voiceprint information in the embodiments of the present application;
[0041] Figure 7 A flowchart of an embodiment of voiceprint recognition in the embodiments of the present application;
[0042] Figure 8 A schematic diagram of an application scenario of cross-device voiceprint recognition in the embodiments of the present application;
[0043] Figure 9 A schematic diagram of another application scenario of cross-device voiceprint recognition in the embodiments of the present application;
[0044] Figure 10 A schematic diagram of training a voiceprint extraction model in the embodiments of the present application;
[0045] Figure 11A A schematic diagram of a curve without frequency response compensation in the embodiments of the present application;
[0046] Figure 11BFig. 2 is a schematic diagram of a curve compensated by frequency response in an embodiment of the present application;
[0047] Figure 12 Fig. 4 is a schematic diagram of generating at least one voiceprint information template in an embodiment of the present application;
[0048] Figure 13 Fig. 5 is a schematic diagram of one embodiment of an apparatus in an embodiment of the present application;
[0049] Figure 14 Fig. 6 is a schematic diagram of another embodiment of an apparatus in an embodiment of the present application. DETAILED DESCRIPTION
[0050] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. The terms "first", "second", "third", "fourth" and the like (if any) in the description of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.
[0051] The embodiments of the present application provide a cross-device voiceprint recognition method, which refers to completing voiceprint registration on one device, and using the registered voiceprint information on other unregistered devices to complete corresponding tasks. The multiple devices can include the same type of devices or different types of devices.
[0052] The method is applied to a communication system, and an example of the architecture of the communication system is shown in Fig. 1. Figure 1As shown, the communication system includes a server 101 and a plurality of terminals 102. Among them, the server 101 can be a server, a server cluster, or a cloud server; the terminal 102 can be a terminal in a smart home, including but not limited to smart home appliances (such as smart screens, televisions, smart washing machines, smart air conditioners, etc.), lighting systems, smart speakers, robots, etc.; the terminal 102 can also be a vehicle terminal, a user equipment (UE), a mobile phone, a tablet computer, a personal computer, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device; as an example but not limitation, in this application, the terminal can also be a wearable device. Wearable devices can also be called wearable smart devices, which are smart design of daily wear using wearable technology. It is a general term for devices that can be worn, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices that can be worn directly on the body or integrated into the user's clothes or accessories. Wearable devices are not just hardware devices, but also have strong functions through software support and data interaction, cloud interaction. Broadly speaking, wearable smart devices include full-featured, large-sized devices that can achieve complete or partial functions without relying on smartphones, such as smartwatches or smartglasses, and devices that focus on a specific application and need to be used with other devices such as smartphones, such as various smart wristbands and smart jewelry for monitoring vital signs.
[0053] In this application, the terminal 102 can be a terminal in an internet of things (IoT) system. IoT is an important part of future information technology development, and its main technical feature is to connect objects through communication technology and network to realize the interconnection of man-machine and the interconnection of things. In order to better illustrate the embodiments of the present application, the words involved in the present application are first explained.
[0054] Device-related noise: the device adds device-related noise in the process from receiving the voice to processing the voice signal, and the device-related noise includes but is not limited to channel noise, encoding noise, microphone physical characteristics and quantity, gain, distance, environment, pre-processing algorithm, etc. It can be understood that by registering the device to enter the user's voice, the voiceprint information in the voice will be affected by the device and thus change, that is, it is no longer a "clean" voiceprint, but a voiceprint after adding the device-related noise. For example, different devices have different channels, and the noise or change exists. For example, the signal will attenuate or delay during transmission, so the sound signal will change during transmission. The sound signal is composed of signals of different frequencies, so if the attenuation degree or delay degree of each frequency signal is not uniform during the transmission of the sound, the received sound signal will be distorted, which will cause channel distortion. Or when transmitting the signal, converting the analog signal into a digital signal for transmission may also introduce distortion. During transmission, due to the limited bandwidth of the transmission channel, the more the bit rate, the more high-quality signals can be transmitted. However, the signal cannot be completely lossless during encoding and decoding. Therefore, the voice signal after encoding and decoding will have some loss. Different devices may have different effects on the sound signal.
[0055] Voiceprint identifier: the voiceprint identifier is used to distinguish the registered voiceprint information, each registered voiceprint information has a unique voiceprint identifier, and the voiceprint identifier in the present application can be a general account independently owned by each user. For example, the voiceprint identifier can be a user ID (identity), which can be used as a general account to perform related operations of the device. For example, if the device is a mobile phone, the general account can be used to use complete services such as downloading software, data synchronization, mobile phone positioning, etc. If the device is a television, it can perform user-related data recommendation (favorite programs), etc. Or the voiceprint identifier can also be a special account or identifier for voiceprint information identification function. One voiceprint identifier can correspond to at least one biological voiceprint template of the same user, and the voiceprint identifier is used to indicate the user. The voiceprint identifier in the embodiments of the present application can be described by taking the user ID as an example.
[0056] Registered voiceprint information: the related information of the user's biological voiceprint after eliminating the noise related to the registered device, and the registered voiceprint information includes at least one biological voiceprint template (i.e., a preset biological voiceprint for matching or identification).
[0057] User's biological voiceprint: the voiceprint of the user himself, which is independent of the device recording the voice, and is equivalent to the voiceprint of the speaker heard by the listener face to face.
[0058] In the present application, the Figure 1The devices included in the corresponding communication system are divided according to functions, including a registration device, a storage device, and a voiceprint recognition device. The registration device is used to register the voice of a user. The registration device extracts registration voiceprint information in the voice, which includes the biological voiceprint of the user himself after separating the noise related to the registration device (i.e., registration voiceprint information); the storage device (which can also be referred to as a voiceprint sharing device) is used to store the registration voiceprint information of the user and provides a registration voiceprint information sharing service. The voiceprint recognition device is a device for voiceprint recognition of the voice of a user (which can be a device that has not registered the voiceprint of the user), but can realize the voiceprint recognition function through the shared registration voiceprint information.
[0059] The registration device (or voiceprint registration apparatus) can be any terminal in a plurality of terminals in the communication system, for example, the terminal can be a mobile phone, a tablet computer, etc.
[0060] The storage device can be a server, a cloud server, or can be any terminal in a plurality of terminals, such as at least one terminal in a plurality of terminals, or can be each terminal in a plurality of terminals.
[0061] The voiceprint recognition device (or recognition device, or voiceprint recognition apparatus) can be a terminal in a plurality of terminals that has not registered a voiceprint. For example, if the registration terminal is a mobile phone and a tablet computer, the voiceprint recognition device can be a wearable device, a vehicle terminal, a smart screen, a headset, a personal computer, etc.
[0062] In this application, the voiceprint recognition device is also referred to as a "first terminal", and the registration device is also referred to as a "second terminal".
[0063] It should be noted that the above-mentioned registration device, storage device, and voiceprint recognition device are only for the convenience of explaining the division of the devices from the functional level, and are not limited to each device only performing one function. For example, a terminal can be a registration device, a storage device, and a voiceprint recognition device, such as a mobile phone that can be used to register the voiceprint of a user, can also be used to store registration voiceprint information, and can also be used for voiceprint recognition of a user; a terminal can be a storage device and a voiceprint recognition device, for example, a personal computer that can be used to store registration voiceprint information and can also be a voiceprint recognition device.
[0064] In the embodiment of the present application, the registered device receives the voice input by the user, separates the noise in the voice related to the registered device itself, extracts the biological voiceprint information of the user in the voice, and registers the voiceprint information to obtain registered voiceprint information. Since the registered voiceprint information is the voiceprint information after eliminating the noise related to the device, the influence of the device on the biological voiceprint of the user is eliminated, and the registered voiceprint information can be shared by other devices. The storage device stores the registered voiceprint information, thereby providing a basis for sharing the registered voiceprint information by other devices. When the user needs to use other devices (also referred to as voiceprint recognition devices) for voiceprint recognition, voiceprint registration on the voiceprint recognition device is not needed, and the registered voiceprint information registered on the registered device can be shared. The voiceprint recognition device receives the voice input by the user, separates the noise in the voice related to the voiceprint recognition device itself, extracts the biological voiceprint information (also referred to as target voiceprint information) of the user in the voice, obtains the registered voiceprint information from the storage device, matches the registered voiceprint information with the target voiceprint information, and performs voiceprint recognition to identify the user identity. In the embodiment of the present application, the registered voiceprint information is the biological voiceprint information of the user extracted after eliminating the noise in the voice related to the registered device itself. Since the influence of the registered device on the biological voiceprint is eliminated, the registered voiceprint information can be shared by other voiceprint recognition devices. The user only needs to register the voiceprint on one device (registered device), and can realize the voiceprint recognition function on multiple devices. The user does not need to register the voiceprint on each device, thereby saving the complicated voiceprint registration process, realizing cross-device voiceprint recognition, greatly improving the user experience, and eliminating the influence of the device on the voiceprint. The registered voiceprint information and the target voiceprint information are both voiceprint information after eliminating the noise related to the device, and the influence of the device on the voiceprint is eliminated. The user has a high accuracy rate when using any voiceprint device for voiceprint recognition.
[0065] The present application mainly includes three stages: 1) collection stage of registered voiceprint information; 2) storage stage of registered voiceprint information; and 3) speaker voiceprint recognition stage. Please refer to Figure 2 In the collection stage of registered voiceprint information, the noise in the voice related to the device is eliminated to obtain the biological voiceprint of the user, and the biological voiceprint is recorded under the existing user ID, that is, the biological voiceprint is registered to obtain registered voiceprint information. In the storage stage of registered voiceprint information, the registered voiceprint information is stored in association with the user ID, thereby providing a basis for sharing the registered voiceprint information. In the speaker recognition stage, the registered voiceprint information is matched with target voiceprint information to obtain a matching result, and the matching result is used to indicate the user identity.
[0066] First, for the collection stage of registered voiceprint information, the execution subject of the stage is a registered device, or the execution subject of the stage is a processor, a chip or a chip system in the registered device. Please refer toFigure 3 As shown, the execution subject of this stage takes the registration device as an example, which can perform the following steps: step 301, the registration device receives the voice input by the user. The voice does not focus on the text content of the voice itself, and the text content of the voice is not limited, which is mainly used to extract the voiceprint information. The voice can be collected by the voice input and output device, and the specific implementation can be referred to the specific description of the voice input and output device. Figure 14 Step 302, the registration device eliminates the noise in the voice related to the registration device to obtain the biological voiceprint of the user. The specific processing scheme can be referred to the specific description of the subsequent noise elimination scheme, which is not described here.
[0067] Step 303, the registration device registers the biological voiceprint of the user to obtain the registration voiceprint information of the user. "Registration" refers to the operation of recording the biological voiceprint of the user under the existing voiceprint identifier, or can be understood as the operation of establishing the correspondence between the biological voiceprint of the user and the voiceprint identifier, or can be understood as the operation of configuring the voiceprint identifier of the user for the biological voiceprint of the user. The voiceprint identifier can be a general account for performing related operations of the device. For example, the registration device takes a mobile phone as an example, after the user purchases the mobile phone, a user ID will be registered, through which the user can use the services such as downloading software and mobile phone positioning. The voiceprint identifier can also be a special account for voiceprint recognition function, for example, the user's mobile phone number can be used as the special account. In order to distinguish between the voiceprint information that has been "recorded" and the voiceprint information that has not been "recorded", the voiceprint information that has been "recorded" or the biological voiceprint that has been configured with the voiceprint identifier is called "registration voiceprint information". The registration device establishes the correspondence between the biological voiceprint of the user and the user ID, and completes the registration process of the biological voiceprint of the user. The registration voiceprint information includes the biological voiceprint of the user in step 302.
[0068] The registration device can be used as a storage device to store the registration voiceprint information. Alternatively, the registration device can also send the registration voiceprint information and the corresponding user ID to other storage devices, and the storage devices can store the registration voiceprint information. The storage device can be a cloud server, a server or other terminal.
[0069] In the collection stage of the registration voiceprint information in the present application, the received voice of the user is extracted, the noise in the voice related to the device is eliminated, and the "clean" biological voiceprint of the user (which is closest to the voiceprint when the user talks face to face) is obtained. Because the influence of the device on the voice is reduced, the extracted biological voiceprint information can be used as the registration voiceprint information shared by multiple devices.
[0070] Then, for the 2) registration voiceprint information storage (or registration voiceprint information sharing) stage, the execution subject of storing the registration voiceprint information can be a storage device, or can also be a memory in the storage device. For example, if the storage device is a server, the execution subject of storing the registration voiceprint information can be the server, or can also be a memory in the server. The storage device receives the registration voiceprint information and the corresponding user ID sent by the registration device, and the storage device stores the registration voiceprint information and the corresponding user ID. Optionally, the registration device can also send the registration time of the registration voiceprint information, the identification of the registration device, the voice intensity received by the registration device, and the like to the storage device, and the storage device stores the above information in association with the user ID. The registration time is used to record the time of registration of the voiceprint information, and can be used to prompt the user to update regularly according to the registration time. The identification of the registration device is used for the storage device to identify the registration device, so that when the storage device receives the user ID, it can be identified whether it is sent by the registration device or a non-registration device (voiceprint recognition device). The voice intensity information can represent the distance between the speaker and the registration device when the registration voiceprint information is collected, and the user ID can correspond to multiple biological voiceprint templates of the same user, and the multiple biological voiceprint templates can correspond to the same or different voice intensity information. The registration device can obtain one biological voiceprint template by performing the collection, noise elimination and registration once, i.e., through a process of one execution step 301 to step 303; and the registration device can obtain multiple biological voiceprint templates by performing the above process multiple times. When the multiple biological voiceprint templates correspond to different voice intensity information, the multiple biological voiceprint templates of the same user can cover multiple situations (different distance voiceprint collection situations), and the accuracy of voiceprint recognition can be improved.
[0071] Optionally, in one example, referring to Figure 4 As shown, the registration device sends the registration voiceprint information and the user ID to the cloud (or server), and the cloud (or server) stores the registration voiceprint information and the user ID in association. In this example, the registration voiceprint information of the user is stored in the cloud (or server), and the registration voiceprint information is distinguished by the user ID, which can save the storage space of the terminal, and can be applied to a wider application scenario, not only including indoor home application scenarios, but also including outdoor scenarios, such as application scenarios of cellular vehicle networks, as long as terminals that can connect to the network can share the registration voiceprint information.
[0072] Optionally, in a second example, referring to Figure 5As shown, the registration device sends the registered voiceprint information and the user ID to the cloud (or server), and any terminal (such as terminal 2, also referred to as a third terminal) in the plurality of terminals can download the registered voiceprint information to the local according to the storage resource condition of the terminal, and the registered voiceprint information can be stored in the third terminal (such as a tablet computer), which can be any terminal in a local area network, such as in a home scenario, the terminals included in the home scenario include a mobile phone, a smart screen, a smart lamp, and a tablet computer, and the plurality of terminals can be connected through wireless fidelity (WiFi), Bluetooth, infrared, network card, etc. When the voiceprint recognition device needs to perform voiceprint recognition, the voiceprint recognition device can quickly obtain the registered voiceprint information from the third terminal.
[0073] Optionally, in the third example, referring to Figure 6 As shown, the registered voiceprint information can be distributed and stored in each terminal (such as a smart screen, a personal computer, a vehicle-mounted terminal, and a tablet computer, etc.) in the plurality of terminals. For example, each terminal sends a user ID to the cloud, and each terminal can receive the registered voiceprint information corresponding to the user ID from the cloud (or server). Alternatively, the voiceprint data can be shared between terminals through Bluetooth, WiFi, infrared, network card, etc. The voiceprints of the voiceprint devices can be shared between terminals in a periodic or non-periodic manner. The user can also configure not to update the voiceprint and not to share the voiceprint on each device. If the registered voiceprint data cannot be shared between terminals through the above communication methods, the registered voiceprint information can also be shared when the terminals and other devices are connected.
[0074] Each terminal has stored the registered voiceprint information corresponding to the user ID. At this time, if voiceprint recognition is needed, the voiceprint recognition device is also a storage device, and the voiceprint recognition device can obtain the registered voiceprint information from the local, without the need to obtain the registered voiceprint information from the cloud or from other devices, that is, the registered voiceprint information is shared, and voice recognition can be quickly performed.
[0075] Finally, for 3) the speaker voiceprint recognition stage, the execution subject of this stage is the voiceprint recognition device (also referred to as the first terminal), or it can also be a processor, a chip, or a chip system in the first terminal. Taking the first terminal as an example, the first terminal receives the voice input by the user, separates the noise related to the first terminal in the voice, extracts the biological voiceprint information (also referred to as the target voiceprint information) of the user, the first terminal obtains the registered voiceprint information from the storage device, matches the registered voiceprint information and the target voiceprint information, and thus performs cross-device voiceprint recognition. If the registered voiceprint information and the target voiceprint information match, the voiceprint recognition is successful, and if the registered voiceprint information and the target voiceprint information do not match, the voiceprint recognition fails.
[0076] Please refer toFigure 7 As shown, in the speaker voiceprint recognition stage, the execution subject is taken as an example to illustrate a first terminal (such as a smart speaker) which can perform the following steps: step 701, the first terminal receives a voice input by a user, which can be achieved by voice collection of a voice input / output device, and will not be described here. Step 702, the first terminal eliminates noise in the voice related to the recognition device to obtain target voiceprint information, and the target voiceprint information is the biological voiceprint of the user, which will be described later. Step 703, the first terminal obtains the registered voiceprint information of the user. The registered voiceprint information can be obtained from a registration device or other devices, or pre-stored in the memory of the first terminal and read from the memory when used. The specific process of obtaining the information from the registration device or other devices can be referred to the description below.
[0077] In a possible implementation, the registered voiceprint information and the ID corresponding to the registered voiceprint information are stored by a storage device or a registration device. The storage device can be a cloud server, a server, or the storage device can be a terminal, or the storage device can be multiple terminals.
[0078] The voiceprint recognition device sends a user ID to the storage device or the registration device to request the registered voiceprint information of the user, and the voiceprint recognition device receives the registered voiceprint information and the corresponding user ID sent by the storage device or the registration device.
[0079] In another possible implementation, if the storage device is a terminal or multiple terminals in a communication system, the storage device can share (or synchronize) the registered voiceprint information and the corresponding user ID to other terminals. For example, the storage device can synchronize periodically, or multiple terminals can be connected to the same local area network, and thus can be synchronized. For example, if the storage device is a tablet computer, the communication system includes three terminals, the tablet computer, a television and a mobile phone, when the three terminals are connected (such as through WiFi), the tablet computer transmits the stored registered voiceprint information and the corresponding ID to the mobile phone and the television, that is, the voiceprint recognition device receives the registered voiceprint information and the corresponding ID sent by the storage device. In this implementation, the voiceprint recognition device does not need to send a user ID to the storage device to request the registered voiceprint information.
[0080] In another possible implementation, the first terminal sends a user ID and an identifier of the first terminal to the cloud server, the cloud server determines a storage location (i.e., a storage device) of registered voiceprint information corresponding to the user ID according to the user ID, the cloud server determines that the registered voiceprint information is stored in the first device according to the user ID, the cloud server sends the user ID and the identifier of the first terminal to the first device, the first device is in communication connection with the first terminal, the first device sends the registered voiceprint information to the first terminal according to the identifier of the first terminal, and the first terminal receives the registered voiceprint information sent by the first device.
[0081] In step 704, the first terminal matches the target voiceprint information with the registered voiceprint information to obtain a matching result, and the matching result is used to indicate a user identity. If the target voiceprint information matches the registered voiceprint information, it is determined that the user is the same user as the user corresponding to the registered voiceprint information, i.e., the user identity is true (or a preset user). If the target voiceprint information does not match the registered voiceprint information, it is determined that the user is a different user from the user corresponding to the registered voiceprint information, i.e., the user identity is false. Optionally, if the user identity is the preset user, the first terminal executes a control instruction of the voice. If the user identity is false (or a non-pre-set user), it is determined that the user is a different user from the user corresponding to the registered voiceprint information, and the first terminal does not need to execute the control instruction of the voice.
[0082] For the above three stages, an application scenario is taken as an example, please refer to Figure 8As shown, the terminals used by the user S in daily life include a mobile phone, a tablet computer, a computer, a smart screen and a vehicle terminal. The user S can register the voiceprint of himself by registering a device (such as a mobile phone). The mobile phone receives the voice input by the user S, such as "Hello, Xia Yi". The mobile phone does not focus on the text itself, extracts the biological voiceprint information in the voice, and associates the biological voiceprint information with the user ID of the user S. The mobile phone sends the user ID and the biological voiceprint information to the cloud, which serves as a storage center for registered voiceprint information. When the user S wants to control the vehicle terminal by voice and log in to the user ID on the vehicle terminal (which can be pre-logged in and does not need to be logged in every time), the vehicle terminal sends the user ID to the cloud through a cellular network. The cloud sends the registered voiceprint information corresponding to the user ID to the vehicle terminal. The vehicle terminal receives the voice input by the user S, such as "play music". The vehicle terminal receives the voice input by the user S, separates the noise related to the vehicle terminal in the voice, extracts the target voiceprint information in the voice, and matches the registered voiceprint information received from the cloud with the target voiceprint information. When the registered voiceprint information matches the target voiceprint information, the vehicle terminal confirms that the user identity of the user S is a preset user and executes the voice instruction "play music".
[0083] It can be understood that according to the purpose of the voiceprint recognition technology, it can be classified into "speaker verification" and "speaker identification". "Speaker verification" refers to judging whether the testee is a specified person, and "speaker identification" refers to identifying which one of the recorded speakers the testee is.
[0084] The above Figure 8 The corresponding application scenario is the application scenario of "speaker verification", that is, the vehicle terminal receives the voice and extracts the target voiceprint information in the voice, determines whether the target voiceprint information and the registered voiceprint information corresponding to the user ID are the voiceprints of the same person, and determines the identity of the user S. After that, the vehicle terminal can execute the voice instruction of the user S. If the target voiceprint information and the registered voiceprint information do not match, the vehicle terminal does not need to execute the voice instruction of the user S. Figure 8 The corresponding application scenario is only an example. The application scenarios of the present application include but are not limited to account login (such as bank account login), identity verification (such as voice recognition of anti-theft doors and identity recognition in financial securities transactions).
[0085] This application can also be applied to "speaker identification" scenarios. Optionally, each registered voiceprint information corresponds to a user ID. For example, in a home scenario, each user ID corresponds to user data. This user data can contain different information for different application scenarios, such as the user's favorite TV programs or the air conditioner temperature. Taking the air conditioner temperature as an example, the correspondence between user IDs and user data can be shown in Table 1 below:
[0086] Table 1
[0087] User ID User Air Conditioner Temperature 1A f 25 degrees Celsius 2D g 20 degrees Celsius 3C c 27 degrees Celsius
[0088] The target voiceprint information is matched with multiple registered voiceprint information corresponding to the voiceprint identifier to determine the target registered voiceprint information that matches the target voiceprint information; user data associated with the user ID corresponding to the target registered voiceprint information is obtained; and then the operation corresponding to the user data is executed.
[0089] For example, please see Figure 9 As shown, in a "speaker recognition" application scenario, a family includes three family members: user f (e.g., father), user g (e.g., mother), and user c (e.g., child). Using a tablet computer as the registration device, all three family members can register their voiceprints through the tablet. The tablet can register a preset number of voiceprints, which can be set by the user or by the tablet's system settings. In this application scenario, taking the example of a terminal registering the voiceprints of three users, user f logs in with their user ID (or pre-logs in). The tablet receives the voice input by user f for registration. This voice can be any text content (e.g., "Hello, Xiaoyi"), and the text content is not limited. The tablet separates the noise related to the tablet from the voice, extracts user f's biometric voiceprint information from the voice, associates this biometric voiceprint information with the user ID (e.g., "1A"), and registers it, obtaining user f's first registered voiceprint information. Similarly, the tablet receives the voice input for registration from user g and obtains the second registration voiceprint information of user g, the user ID corresponding to the second registration voiceprint information is "e.g., 2D". The tablet receives the voice input for registration from user c and obtains the third registration voiceprint information of user c, the user ID corresponding to the third registration voiceprint information is "e.g., 3C". The tablet sends the user ID corresponding to each registration voiceprint information to the cloud for storage.
[0090] If the user needs to use the smart speaker and wants to control the air conditioner through the smart speaker (the smart speaker has been associated with the user ID on the tablet computer smart home application), the smart speaker receives the voice "Xiao Yi, turn on the air conditioner" input by the user f, separates the noise related to the smart speaker in the voice, extracts the user's biological voiceprint information (target voiceprint information) in the voice, and sends the user ID to the cloud. The cloud sends the three user IDs and the corresponding three registered voiceprint information to the smart speaker, or the smart speaker has pre-stored the three registered voiceprint information and the corresponding user ID. The smart speaker matches the target voiceprint information with the three registered voiceprint information received. If the target voiceprint information matches the first registered voiceprint information (such as the registered voiceprint information of user f) in the three registered voiceprint information, the voice instruction is executed. Optionally, the smart speaker determines the user ID (such as "1A") corresponding to the first registered voiceprint information. The smart speaker can determine the relevant user data corresponding to the voiceprint identifier, such as the relevant user data being "temperature 25℃". It can be understood that the smart speaker identifies the user ID (such as "1A") corresponding to the first registered voiceprint information, indicating that the voice is the voice input by user f. The historical user data recorded by the smart speaker can be: the user ID (1A) corresponds to "temperature 25℃", indicating that user f often adjusts the temperature of the air conditioner to 25℃. The smart speaker sends a control instruction to the air conditioner according to the relevant information to adjust the temperature of the air conditioner to 25℃.
[0091] The voiceprint recognition technology is used to identify a specific user, and selectively execute a command. At the same time, different users are identified, and the smart device can also provide personalized services for different users, expanding the application of the smart device.
[0092] The above application scenario is a home application scenario. The present application can also be applied to a work scenario. The registered voiceprint information of multiple members of a work group corresponds to multiple IDs. The voiceprint recognition technology in the present application can be used to identify which user in the work group the speaker is (i.e., to identify the speaker). Different users have different permissions. If it is identified that the speaker is the user d corresponding to the user ID, the voiceprint recognition device directly executes the user data of the permission of user d.
[0093] In this example, the speaker recognition can be performed by the voiceprint recognition, the speaker is recognized, and the user data corresponding to the speaker is determined, which includes but is not limited to historical information associated with the speaker or user data corresponding to the authority of the speaker. The device can be intelligently customized or intelligently recommended according to the user data. For example, in an application scenario, the smart speaker can perform an operation corresponding to the historical information (such as a temperature of 25°C) associated with the speaker (the user corresponding to the user ID "1A") according to the historical information, without the need for the user to perform the habitual operation each time, thereby improving the user experience.
[0094] Optionally, the above Figure 7 In step 702 of the corresponding implementation, the first terminal extracts the target voiceprint information in the voice that is irrelevant to the first terminal. The specific manner of extracting the target voiceprint information in the voice that is irrelevant to the first terminal can include the following manners. Figure 3 In step 302 of the corresponding implementation, the registration device extracts the biological voiceprint information in the voice that is irrelevant to the device. The specific manner of extracting the biological voiceprint information in the voice that is irrelevant to the device can include the following manners.
[0095] The machine learning manner includes the following manners.
[0096] The first terminal (or the voiceprint recognition device) inputs the voice into a voiceprint extraction model, and outputs the target voiceprint information through the voiceprint extraction model. The voiceprint extraction model is obtained by learning a plurality of devices.
[0097] The voiceprint extraction model includes but is not limited to a gauss mixture model (GMM), a GMM universal background model (GMM-UBM), an i-vector, an x-vector, a dnn-ivector, a deep neural network (DNN), speech analysis, speech factor decomposition, clustering, transformation, and the like.
[0098] The machine learning includes a training phase of the voiceprint extraction model and an application phase of the voiceprint extraction model.
[0099] Please refer to Figure 10As shown, in the training phase of the voiceprint extraction model, a large amount of corpus for learning and reference data for reference are obtained, and the large amount of corpus includes corpus collected by different types of devices. For example, the different types of devices include but are not limited to smart home appliances (such as smart screens, televisions, smart washing machines, smart air conditioners, etc.), lighting systems, smart speakers, robots, etc.; the terminal can also be a vehicle terminal, a user equipment, a mobile phone, a tablet computer, a personal computer, a virtual reality terminal device, an augmented reality terminal device, a wearable device, etc. The reference data is voiceprint information independent of (or weakly related to) the device.
[0100] A large amount of corpus input by the same user through different types of devices is input to a voiceprint model (such as a GMM-UBM model), and the voiceprint data output by the GMM-UBM model is compared with reference data, which is the user's biological voiceprint data. If the difference between the output voiceprint data and the reference data is greater than or equal to a threshold value, the output voiceprint data is re-input to the GMM-UBM model, and the voiceprint data output by the GMM-UBM model is then compared with the reference data. Through continuous iterative training, if the difference between the voiceprint data output by the GMM-UBM model and the reference data is less than the threshold value, it indicates that the voiceprint data output by the model can approach the reference data, and a voiceprint extraction model is obtained, which is used to separate device-related noise and extract the user's biological voiceprint. In this example, the voiceprint extraction model is pre-trained through machine learning, and the user's biological voiceprint in the voice is directly extracted through the already trained voiceprint extraction model, which has good robustness.
[0101] II. Signal processing method
[0102] In this application, the voice signal can be processed by a filter for frequency response compensation, which is used to eliminate the noise related to the identification device in the voice to obtain the user's biological voiceprint information. The filter is the most basic signal processing device, which extracts the required signal from a plurality of mixed signals. In this application, the main function of the filter is to eliminate various noises affecting signal processing. The filter produces different gains according to different frequencies, so that specific signals are highlighted, the highlighted signals are attenuated, and weaker signals are enhanced, thereby achieving the purpose of eliminating device noise.
[0103] As shown in the following expression:
[0104]
[0105] Equation (1) above is a finite impulse response filter, where n is the time point, N is the unit impulse response length of the digital filter, the coefficient a is convolved with the coefficient x to produce the filter output y, and k starts from 0 and continues to N.
[0106] Alternatively, the filter can be expressed as follows:
[0107]
[0108] Equation (2) above is an infinite impulse response filter, where n refers to the time point, N and P are the unit impulse response lengths of the digital filter, k starts from 0 and continues to N, and the coefficients a and x are convolved; j starts from 0 and continues to P, and the sum of the convolutions of the coefficients b and y produces the filtered output y.
[0109] If the frequency response of different devices tends to be level or consistent by adjusting coefficients a and x in equation (1) above, and if the frequency response of different devices tends to be level or consistent by adjusting coefficients a and b in equation (2) above, then device-related noise in speech can be filtered out.
[0110] Please see Figure 11A As shown, in Figure 11A This includes three curves: the upper limit curve 1101, the lower limit curve 1102, and the uncompensated frequency response curve 1103. Between 600Hz and 1.2kHz, the frequency response curve 1103 shows a peak exceeding that of the upper limit curve 1101. Between 300Hz and 500Hz, the frequency response curve 1103 exhibits an upward trend. Please refer to [link / reference]. Figure 11B As shown, Figure 11B The formula includes three curves: upper limit curve 1101, lower limit curve 1102, and frequency response compensation curve 1104. In the above formula (1), the frequency response of different devices is made to tend to be horizontal by adjusting coefficients a and x. Alternatively, in the above formula (2), the frequency response of different devices is made to tend to be horizontal by adjusting coefficients a and b. In the frequency range of 300 Hz to 2.5 kHz, the frequency response tends to be consistent (e.g., -2 dBr), thereby filtering out device-related noise in the speech and obtaining the user's bio-voiceprint.
[0111] In this example, the speech signal is processed by a filter to attenuate prominent signals and enhance weaker signals, thereby compensating for the frequency response and removing device-related noise. The filter method directly filters out noise signals, and the output signal is the user's biometric voiceprint signal, which is simple and fast to implement.
[0112] In one alternative implementation, each person's biometric voiceprint is not static but changes, such as at different times of the day, different health conditions (e.g., healthy and sick), or different ages. These factors can all cause the same person's biometric voiceprint to change. To improve the robustness of the system, multiple voiceprints can be registered for the same user, that is, the same user corresponds to multiple biometric voiceprint templates, and one user ID corresponds to one registered voiceprint information, which includes multiple biometric voiceprint templates.
[0113] The storage device can generate voiceprint models (i.e., voiceprint templates) based on multiple biometric voiceprints of the same user. Please refer to [link / reference]. Figure 12 As shown, taking the x-vector system as an example, the x-vector system is a speaker recognition system built on DNN. By training the DNN, the speaker's speech is mapped to fixed-dimensional embeddings, called x-vectors.
[0114] The x-vector network receives voiceprint information from the same user, which is a user voiceprint with device-related noise removed. The x-vector network can capture user voiceprint information using shorter speech, exhibiting stronger robustness on short speech. One input corresponds to one x-vector, which is the user's voiceprint information (since device-related information has been removed, this voiceprint information is device-independent), or it becomes the user's voiceprint model. If multiple device-independent voiceprints are input from multiple devices, or multiple device-independent voiceprints are input from one device, multiple x-vectors will be generated. Dimensionality reduction is performed using linear discriminant analysis (LDA). These multiple vectors can cover more user pronunciation scenarios (such as voiceprint information from different time periods, healthy or unhealthy states, etc.), meaning one user corresponds to multiple biometric voiceprint models (or templates). These multiple biometric voiceprint models form a voiceprint information template library, which can further enhance the voiceprint recognition effect. The number of templates corresponding to the same user can be 10-30, and they can be updated regularly or irregularly, replacing old templates with new ones to improve the robustness of biometric voiceprint templates.
[0115] It can be understood that the registration device can register one bio-acoustic print template at a time, or register multiple bio-acoustic print templates in multiple times. Or each of the multiple registration devices can register one bio-acoustic print template respectively. Alternatively, the registration device can send one or more bio-acoustic print templates to the storage device or the identification device through one message at a time, or send multiple bio-acoustic print templates to the storage device or the identification device through multiple messages respectively. The channel for sending includes but is not limited to wireless and wired ways, and can be implemented through a transceiver, which can be referred to the corresponding description of Figure 14 The identification device can obtain the one or more bio-acoustic print templates from the registration device or the storage device. For the identification device, the one or more bio-acoustic print templates are used for matching in acoustic print recognition, and belong to the registration acoustic print information that the identification device can use, although the registration acoustic print information is not generated and registered by the identification device itself, but comes from other devices. Through this scheme, device-independent acoustic print information sharing is realized, flexible cross-device acoustic print recognition is realized, and user experience is improved.
[0116] Corresponding to the method given in the above method embodiment, the embodiment of the present application also provides a corresponding device, which includes modules for executing the corresponding modules of the above embodiments. The module can be software, hardware, or a combination of software and hardware. As Figure 13 The embodiment of the present application also provides a device 1300, which can be a terminal or a component of a terminal (for example, an integrated circuit, a chip, etc.), and the acoustic print recognition device includes a voice input and output module 1301 (or a voice input and output unit), a processing module 1302 (or a processing unit), and a transceiver module 1303 (or a transceiver unit).
[0117] In a possible design, the device 1300 can perform the functions of the identification device in the above method embodiments: the voice input and output module 1301 is configured to receive a voice input by a user; the processing module 1302 is configured to eliminate noise related to the identification device from the voice received by the voice input and output module 1301 to obtain target acoustic print information, the target acoustic print information being a bio-acoustic print of the user; and the transceiver module 1303 is configured to obtain registration acoustic print information of the user, the registration acoustic print information including at least one bio-acoustic print template obtained by eliminating noise related to a registration device; and the processing module 1302 is further configured to match the target acoustic print information obtained by the processing module 1302 with the registration acoustic print information obtained by the transceiver module 1303 to obtain a matching result, the matching result being used to indicate a user identity.
[0118] Further, the voice input and output module 1301 is configured to perform the above Figure 7The step 701 in the corresponding embodiment is specifically implemented by referring to the detailed description in the step 701, which is not repeated here. The processing module 1302 is configured to perform the processing described above. Figure 7 The steps 702 and 704 in the corresponding embodiment are specifically implemented by referring to the detailed description of the steps 702 and 704, which is not repeated here. The transceiver module 1303 is configured to perform the processing described above. Figure 7 The step 703 in the corresponding embodiment is specifically implemented by referring to the detailed description of the step 703, which is not repeated here.
[0119] In another possible design, the apparatus 1300 can perform the function of registering a device in the method embodiments described above. The voice input and output module 1301 is configured to receive a voice input by a user. The processing module 1302 is configured to eliminate noise related to the registered device from the voice received by the voice input and output module 1301 to obtain registered voiceprint information of the user. Optionally, the transceiver module 1303 is configured to send the registered voiceprint information obtained by the processing module 1302 and the corresponding voiceprint identifier to another storage device.
[0120] Further, the voice input and output module 1301 is configured to perform the processing described above. Figure 3 The step 301 in the corresponding embodiment is specifically implemented by referring to the detailed description in the step 301, which is not repeated here. The processing module 1302 is configured to perform the processing described above. Figure 3 The steps 302 and 303 in the corresponding embodiment are not repeated here.
[0121] In another implementation manner, the apparatus can be a chip or an integrated circuit. At this time, the transceiver module 1303 can be a communication interface, the processing module 1302 can be a logic circuit, and the voice input and output module 1301 can be an audio circuit. Optionally, the communication interface can be an input and output interface or a transceiver circuit. The input and output interface can include an input interface and an output interface. The transceiver circuit can include an input interface circuit and an output interface circuit.
[0122] In an implementation manner, the processing module 1302 can be a processing apparatus, and the functions of the processing apparatus can be partially or entirely implemented by software. Optionally, the functions of the processing apparatus can be partially or entirely implemented by software. At this time, the processing apparatus can include a memory and a processor, where the memory is used to store a computer program, and the processor reads and executes the computer program stored in the memory to perform the corresponding processing and / or steps in any one of the method embodiments. Optionally, the processing apparatus can include only the processor. The memory used to store the computer program is located outside the processing apparatus, and the processor is connected with the memory through a circuit / wire to read and execute the computer program stored in the memory.
[0123] It can be understood that, Figure 13Each of the functional components can be implemented by software, hardware, or a combination of both, and is not limited in particular.
[0124] In addition, Figure 14 A schematic structural diagram of an apparatus 1400 is provided for an embodiment of the present application. The apparatus can be a terminal, or can also be an integrated circuit, a chip or a chip system in the terminal. The apparatus takes the terminal as an example, which can include but is not limited to a terminal in a smart home, a lighting system, a smart speaker, a robot, etc.; the terminal can also be a vehicle terminal, a user equipment, a mobile phone, a tablet computer, a personal computer, etc.
[0125] As Figure 14 shown, the apparatus 1400 includes a processor 1401, a transceiver 1402, a memory 1403, and a voice input / output device 1404. Among them, the processor 1401, the transceiver 1402, the memory 1403, and the voice input / output device 1404 can communicate with each other through internal connection paths to transfer control signals and / or data signals. Among them, the memory 1403 is used to store computer programs, and the processor 1401 is used to call and run the computer programs from the memory 1403 to control the transceiver 1402 to transceive signals. The processor 1401 is the control center of the apparatus, connects various parts of the entire mobile phone through various interfaces and lines, and performs various functions and processes data of the mobile phone by running or executing software programs and / or modules stored in the memory 1403, and calling data stored in the memory 1403.
[0126] The voice input / output device 1404 is used for the audio interface between the user and the mobile phone. The voice input unit can be an audio circuit, or can be a voice recognizer. The audio circuit can include a loudspeaker 14041 and a microphone 14042, the microphone 14042 converts the collected sound signal into an electrical signal, which is received by the audio circuit and converted into audio data, and then the audio data is output to the processor 1401 for processing, and the processor 1401 eliminates the noise related to the apparatus to obtain the user's biological voiceprint.
[0127] Optionally, the apparatus can further include an antenna. The transceiver 1402 transmits or receives wireless signals through the antenna. The transceiver 1402 can be configured to send or receive the registered voiceprint information and the corresponding voiceprint identifier to other devices. Optionally, the processor 1401 and the memory 1403 can be combined into one processing apparatus, and the processor 1401 is configured to execute the program codes stored in the memory 1403 to implement the above functions. Optionally, the memory 1403 can also be integrated in the processor 1401. Alternatively, the memory 1403 is independent of the processor 1401, i.e., located outside the processor 1401. Optionally, the transceiver 1402 includes but is not limited to radio frequency (RF) circuit, communication interface, WiFi module, Bluetooth module, etc.
[0128] Optionally, the apparatus can further include a display unit 1405 configured to display information input by the user or information provided to the user, as well as various images. The display unit can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.
[0129] In one possible design, the apparatus 1400 can be configured to perform the functions of identifying devices in the method embodiments: a voice input / output device 1404 configured to receive voice input by the user; a processor 1401 configured to eliminate noise related to the identifying device from the voice to obtain target voiceprint information, the target voiceprint information being a biological voiceprint of the user; a transceiver 1402 configured to obtain registered voiceprint information of the user, the registered voiceprint information including at least one biological voiceprint template obtained by eliminating noise related to the registering device; and the processor 1401 is further configured to match the target voiceprint information with the registered voiceprint information to obtain a matching result, the matching result being used to indicate the identity of the user.
[0130] Optionally, the transceiver 1402 is configured to receive the registered voiceprint information and the voiceprint identifier of the registered voiceprint information sent by the storage device or the registering device, the voiceprint identifier being used to indicate the user. Optionally, the transceiver 1402 is further configured to send the voiceprint identifier of the user to the storage device or the registering device to request the registered voiceprint information of the user.
[0131] Optionally, the processor 1401 is specifically configured to extract a voice input voiceprint through a voiceprint extraction model, and eliminate noise in the voice related to the identification device through the voiceprint extraction model to obtain target voiceprint information. Optionally, the voiceprint extraction model is obtained by learning corpus collected by multiple devices. Optionally, the processor 1401 is further configured to perform frequency response compensation on the voice through a filter, and the frequency response compensation is used to eliminate noise in the voice related to the identification device to obtain the target voiceprint information. Optionally, the processor 1401 is further configured to obtain user data associated with the user identity, and perform an operation corresponding to the user data.
[0132] In a possible design, the apparatus 1400 is configured to perform functions performed by the registered device in the above method embodiments: the voice input output device 1404 is configured to receive voice input by a user; and the processor 1401 is configured to: eliminate noise in the voice related to the registered device to obtain registered voiceprint information of the user, and the registered voiceprint information specifically includes biological voiceprint of the user.
[0133] Optionally, the transceiver 1402 is configured to send the registered voiceprint information and a voiceprint identifier corresponding to the registered voiceprint information to a storage device or an identification device, and the voiceprint identifier is used to indicate the user. Optionally, the processor 1401 is specifically configured to extract a voice input voiceprint through a voiceprint extraction model, and eliminate noise in the voice related to the registered device through the voiceprint extraction model to obtain the registered voiceprint information. Optionally, the voiceprint extraction model is obtained by learning corpus collected by multiple devices. Optionally, the processor 1401 is further configured to perform frequency response compensation on the voice through a filter, and the frequency response compensation is used to eliminate noise in the voice related to the identification device to obtain the registered voiceprint information.
[0134] It can be understood that the processor in the embodiments of the present application can be an integrated circuit chip having a signal processing capability. In the implementation process, each step of the above method embodiments can be completed by an integrated logic circuit or an instruction in the form of software in the processor. The processor can be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0135] The schemes described in this application can be implemented in various ways. For example, these techniques can be implemented using hardware, software, or a combination thereof. For a hardware implementation, the processing units used to perform the techniques at a communication device (e.g., a base station, a terminal, a network entity, or a chip) can be implemented within one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, micro-controllers, microprocessors, other electronic units designed to perform the functions described herein, a combination thereof, or such any other hardware equivalents.
[0136] It is to be understood that the memory in the embodiments of the present application can be a volatile memory or a nonvolatile memory, or can include both volatile and nonvolatile memory. The nonvolatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache memory. By way of example, and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct Rambus RAM (DR RAM). Note that the system and method described herein are intended to include all such memory types and any other suitable type of memory.
[0137] The application further provides a computer readable medium, which stores a computer program, and the computer program realizes the functions of any of the method embodiments when executed by a computer. The application further provides a computer program product, which realizes the functions of any of the method embodiments when executed by a computer.
[0138] It can be understood that the "embodiments" mentioned throughout the specification mean that the specific features, structures or characteristics related to the embodiments are included in at least one embodiment of the application. Therefore, the various embodiments throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It can be understood that the size of the sequence number of each process described above in various embodiments of the application does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the application.
[0139] It can be understood that in the present application, "when", "if" and "if" all refer to the corresponding processing of the device under certain objective circumstances, not limited by time, and do not require the device to have a judgment action when it is implemented, nor does it mean that there are other limitations. In the present application, the use of singular elements is intended to represent "one or more", not "one and only one", unless otherwise specified. In the present application, "at least one" is intended to mean "one or more", and "multiple" is intended to mean "two or more". In addition, the terms "system" and "network" are often used interchangeably in this document.
[0140] The term "at least one of" or "at least one of" in this document means all or any combination of the listed items, for example, "at least one of A, B and C" can mean: A exists alone, B exists alone, C exists alone, A and B exist together, B and C exist together, A, B and C exist together, of which A can be singular or plural, B can be singular or plural, and C can be singular or plural.
[0141] It can be understood that in various embodiments of the present application, "B corresponding to A" means that B is associated with A and can be determined according to A. However, it should also be understood that determining B according to A does not mean that B is determined only according to A, but B can also be determined according to A and / or other information. The term "and / or" in this document is a description of the association relationship between the associated objects, which means that there can be three kinds of relationships, for example, A and / or B can mean: A exists alone, A and B exist together, B exists alone, of which A can be singular or plural, and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it.
[0142] Those skilled in the art can understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0143] It can be understood that the system, device and method described in the present application can also be implemented in other ways. For example, the above-described device embodiments are only schematic, and the division of the units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or in other forms.
[0144] If the functions described in the embodiments are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0145] The same or similar parts in the various embodiments in the present application can be mutually referred to. In the various embodiments in the present application, and the various implementation manners / implementation methods / realization methods in each embodiment, if there is no special description and no logical conflict, the terms and / or descriptions in different embodiments, and the various implementation manners / implementation methods / realization methods in each embodiment are consistent and can be mutually referred to. The technical features in different embodiments, and the various implementation manners / implementation methods / realization methods in each embodiment can be combined to form new embodiments, implementation manners, implementation methods, or realization methods according to their inherent logical relationship. The above-described implementation manners of the present application do not constitute a limitation on the protection scope of the present application.
[0146] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application.
Claims
1. A voiceprint recognition apparatus in a recognition device, characterized by, The device comprises: a processor, a voice input and output device connected to the processor, and a transceiver; the voice input and output device is configured to receive voice input by a user; the processor is configured to eliminate noise related to the identification device in the voice to obtain target voiceprint information, the target voiceprint information being a biological voiceprint of the user; the transceiver is configured to obtain registered voiceprint information of the user, the registered voiceprint information comprising at least one biological voiceprint template obtained by eliminating noise related to a registration device; the processor is further configured to match the target voiceprint information with the registered voiceprint information to obtain a matching result, the matching result being used to indicate the identity of the user; and the processor is specifically configured to use a voiceprint extraction model to eliminate noise related to the identification device in the voice to obtain the target voiceprint information, the voiceprint extraction model being obtained by using biological voiceprint data of the user as reference data to learn a corpus input by the same user using a plurality of devices, the plurality of devices comprising the registration device and the identification device, or the processor is configured to use a filter to perform frequency response compensation on the voice, the frequency response compensation being used to eliminate noise related to the identification device in the voice to obtain the target voiceprint information, the filter being used to make the frequency responses of different devices tend to be level or consistent, the different devices comprising the registration device and the identification device.
2. The device of claim 1, wherein: the transceiver is configured to receive the registered voiceprint information and a voiceprint identifier of the registered voiceprint information sent by a storage device or the registration device, the voiceprint identifier being used to indicate the user.
3. The device of claim 2, wherein: the transceiver is further configured to send a voiceprint identifier of the user to the storage device or the registration device to request registered voiceprint information of the user.
4. The apparatus of any one of claims 1-3, wherein, the processor is further configured to: obtain user data associated with the identity of the user; and perform an operation corresponding to the user data.
5. A voiceprint registration apparatus in a registration device, characterized by, The device comprises: a processor and a voice input and output device connected to the processor; the voice input and output device is configured to receive voice input by a user; the processor is configured to eliminate noise related to the registration device in the voice to obtain registered voiceprint information of the user, the registered voiceprint information comprising a biological voiceprint of the user; and the processor is further configured to match the target voiceprint information with the registered voiceprint information to obtain a matching result, the matching result being used to indicate the identity of the user; and The processor is specifically configured to extract the voice input voiceprint model, eliminate noise related to the registration device in the voice through the voiceprint extraction model to obtain the registration voiceprint information, the voiceprint extraction model is obtained by learning corpus input by the same user collected by multiple devices with the biological voiceprint data of the user as reference data, the multiple devices include the registration device and the identification device, or, for frequency response compensation of the voice through a filter, the frequency response compensation is used to eliminate noise related to the registration device in the voice to obtain the registration voiceprint information, the filter is used to make the frequency responses of different devices tend to be horizontal or consistent, the different devices include the registration device and the identification device.
6. The apparatus of claim 5, wherein, The device further includes a transceiver connected with the processor; The transceiver is configured to send the registration voiceprint information and a voiceprint identifier corresponding to the registration voiceprint information to a storage device or an identification device, and the voiceprint identifier is used to indicate the user. 7.A method of cross-device voiceprint recognition, applied to a recognition device, and having the steps of: Comprise: Receiving voice input by a user; Eliminate noise related to the identification device in the voice to obtain target voiceprint information, the target voiceprint information being the biological voiceprint of the user; Obtain the registration voiceprint information of the user, the registration voiceprint information including at least one biological voiceprint template obtained by eliminating noise related to the registration device; Match the target voiceprint information with the registration voiceprint information to obtain a matching result, the matching result being used to indicate the user identity; wherein The elimination of noise related to the identification device in the voice to obtain the target voiceprint information comprises: Extract the voice input voiceprint model, eliminate noise related to the identification device in the voice through the voiceprint extraction model to obtain the target voiceprint information, the voiceprint extraction model being obtained by learning corpus input by the same user collected by multiple devices with the biological voiceprint data of the user as reference data, the multiple devices including the registration device and the identification device, or, for frequency response compensation of the voice through a filter, the frequency response compensation is used to eliminate noise related to the identification device in the voice to obtain the target voiceprint information, the filter is used to make the frequency responses of different devices tend to be horizontal or consistent, the different devices including the registration device and the identification device.
8. The method of claim 7, wherein, The obtaining of the registration voiceprint information of the user comprises: Receiving the registration voiceprint information and a voiceprint identifier corresponding to the registration voiceprint information sent by a storage device or the registration device, and the voiceprint identifier is used to indicate the user.
9. The method of claim 8, wherein, Before the receiving of the registration voiceprint information and the voiceprint identifier corresponding to the registration voiceprint information sent by the storage device or the registration device, the method further comprises: Sending the voiceprint identifier to the storage device or the registration device to request the registration voiceprint information of the user.
10. The method according to any one of claims 7-9, characterized in that, After the matching of the target voiceprint information with the registration voiceprint information to obtain the matching result, the method further comprises: Obtain user data associated with the user identity; Performing an operation corresponding to the user data. 11.A method of cross-device voiceprint recognition, applied to a registration device, the method comprising: Comprising: Receiving a voice input by a user; Eliminating noise related to the registered device in the voice to obtain registered voiceprint information; The registered voiceprint information comprises a biological voiceprint of the user; wherein The eliminating noise related to the registered device in the voice to obtain biological voiceprint information comprises: Inputting the voice into a voiceprint extraction model, eliminating noise related to the registered device in the voice through the voiceprint extraction model to obtain the registered voiceprint information, the voiceprint extraction model being obtained by learning corpus input by the same user through a plurality of devices, the plurality of devices comprising the registered device and an identification device, or performing frequency response compensation on the voice through a filter, the frequency response compensation being used for eliminating noise related to the registered device in the voice to obtain the registered voiceprint information, the filter being used for making frequency responses of different devices tend to be horizontal or consistent, the different devices comprising the registered device and the identification device.
12. The method of claim 11, wherein, The method further comprises: Sending the registered voiceprint information and a voiceprint identifier corresponding to the registered voiceprint information to a storage device or an identification device.
Citation Information
Patent Citations
A voiceprint recognition method based on 3D convolution neural network
CN109215665A
Cross-device voiceprint recognition method and system
CN109378006A
Voiceprint registration method and system
CN111161746A