Voiceprint recognition method and device, wearable device and computer readable medium
By integrating audio and location acquisition devices into wearable devices, and using voiceprint feature profiles to estimate the voiceprint features corresponding to posture coordinates, the problem of decreased voiceprint recognition rate caused by changes in user posture is solved, and the accuracy of voiceprint recognition is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VOICEAI TECH CO LTD
- Filing Date
- 2022-12-01
- Publication Date
- 2026-05-08
AI Technical Summary
When performing voiceprint recognition on wearable devices, changes in user posture lead to changes in voice, resulting in a decrease in the voiceprint recognition pass rate. Furthermore, adjusting the threshold strategy is difficult to effectively address the changes under different users and postures.
By integrating audio acquisition devices and location acquisition devices into wearable devices, the user's voice information and the device's posture information are obtained. The voiceprint features corresponding to the posture coordinates are estimated using the pre-acquired voiceprint feature profile, and voiceprint recognition is performed.
It improves the success rate of voiceprint recognition in different postures and enhances the accuracy and reliability of voiceprint recognition.
Smart Images

Figure CN115910072B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voiceprint recognition technology, and more specifically, to a voiceprint recognition method, apparatus, wearable device, and computer-readable medium. Background Technology
[0002] Currently, in order to meet the needs of different usage scenarios and improve the user interaction experience, wearable devices are often equipped with voiceprint recognition modules. However, if the wearer changes their posture during use, such as squatting, leaning back, or bending over, the wearer's vocal organs will be compressed, which will cause the voice produced to change significantly, thus reducing the success rate of voiceprint recognition. Summary of the Invention
[0003] This application proposes a voiceprint recognition method, apparatus, wearable device, and computer-readable medium to improve the above-mentioned deficiencies.
[0004] In a first aspect, embodiments of this application provide a voiceprint recognition method applied to a wearable device. The wearable device includes an audio acquisition device and a location acquisition device. The method includes: when the wearable device is worn by a user and used in a voiceprint recognition scenario, obtaining a first voiceprint feature based on the user's voice information acquired by the audio acquisition device at a specified time; performing a voiceprint recognition operation on the first voiceprint feature; if the first voiceprint feature cannot be recognized by voiceprint recognition, obtaining a first posture coordinate based on the posture information of the wearable device acquired by the location acquisition device at a specified time; estimating a second voiceprint feature based on the first posture coordinate and a pre-acquired voiceprint feature profile, wherein the voiceprint feature profile is used to characterize the correspondence between posture coordinates and voiceprint features; and performing a voiceprint recognition operation on the second voiceprint feature.
[0005] Secondly, embodiments of this application also provide a voiceprint recognition device applied to a wearable device. The wearable device includes an audio acquisition device and a location acquisition device. The device includes: a voiceprint feature acquisition unit, a voiceprint recognition unit, a posture coordinate acquisition unit, and a voiceprint feature analysis unit. The voiceprint feature acquisition unit is used to acquire voiceprint features based on the user's voice information acquired by the audio acquisition device; the voiceprint recognition unit is used to perform voiceprint recognition on the voiceprint features; the posture coordinate acquisition unit is used to acquire posture coordinates based on the posture information of the wearable device acquired by the location acquisition device when the voiceprint features acquired by the voiceprint feature acquisition unit do not pass voiceprint recognition; and the voiceprint feature analysis unit is used to analyze and acquire voiceprint features based on the posture coordinates.
[0006] Thirdly, embodiments of this application also provide a wearable device, including: one or more processors; a memory; an audio acquisition device and a location acquisition device; one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to perform the methods described above.
[0007] Fourthly, embodiments of this application also provide a computer-readable medium storing processor-executable program code, which, when executed by the processor, causes the processor to perform the above-described method.
[0008] The voiceprint recognition method, apparatus, wearable device, and computer-readable medium provided in this application, when the wearable device is worn by a user and used in a voiceprint recognition scenario, obtain a first voiceprint feature based on the user's voice information collected by the audio acquisition device at a specified time, and perform a voiceprint recognition operation on the first voiceprint feature. Then, if the first voiceprint feature cannot be recognized by voiceprint recognition, obtain a first posture coordinate based on the posture information of the wearable device collected by the position acquisition device at a specified time, estimate a second voiceprint feature from a pre-acquired voiceprint feature profile, wherein the voiceprint feature profile is used to represent the correspondence between posture coordinates and voiceprint features; then, perform a voiceprint recognition operation on the second voiceprint feature. Therefore, when the first voiceprint feature cannot be recognized by voiceprint identification, that is, when the user's voice is distorted due to changes in posture and cannot be recognized by voiceprint identification, by collecting the posture information of the wearable device, that is, obtaining the first posture coordinates, and determining the new voiceprint feature, that is, the second voiceprint feature corresponding to the first posture coordinates, based on the pre-acquired voiceprint feature profile, the voiceprint recognition of the second voiceprint feature can be performed, which can improve the voiceprint recognition pass rate when the user wears the wearable device.
[0009] Other features and advantages of the embodiments of this application will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the embodiments of this application. The objects and other advantages of the embodiments of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1A flowchart of a voiceprint recognition method provided in an embodiment of this application is shown.
[0012] Figure 2 A flowchart of a voiceprint recognition method provided in another embodiment of this application is shown.
[0013] Figure 3 A flowchart illustrating the method for training and obtaining voiceprint feature profiles provided in this application is shown.
[0014] Figure 4 A flowchart of the method for estimating second voiceprint features from a voiceprint feature profile provided in this application is shown.
[0015] Figure 5 A flowchart of a voiceprint recognition method provided in another embodiment of this application is shown.
[0016] Figure 6 A flowchart of the method for updating voiceprint feature profiles provided in this application is shown.
[0017] Figure 7 A block diagram of a voiceprint recognition device provided in one embodiment of this application is shown.
[0018] Figure 8 A schematic diagram of a wearable device provided in one embodiment of this application is shown.
[0019] Figure 9 A schematic diagram of a storage unit according to an embodiment of this application is shown. Detailed Implementation
[0020] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. The components of the embodiments of the present application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without inventive effort are within the scope of protection of the present application.
[0021] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0022] Voiceprint recognition, a type of biometric technology also known as speaker identification, includes speaker recognition and speaker verification. Different tasks and applications utilize different voiceprint recognition technologies. For example, identification techniques may be needed to narrow down the scope of criminal investigations, while verification techniques are required for banking transactions. The essence of voiceprint recognition is to find the voiceprint features that describe a specific object. Most voiceprint recognition systems use acoustic features, specifically a set of acoustic descriptive parameters extracted from sound signals by computer algorithms / mathematical methods, represented in vector form.
[0023] A typical voiceprint recognition system framework consists of two main modules: training and recognition. The training module acquires a large number of training voices from different people, extracts corresponding training voiceprint features from the training voices, trains a voiceprint feature model, and collects the voiceprint feature models of multiple people to obtain a model library. The recognition module acquires recognition voices, extracts corresponding voiceprint features to be recognized from the recognition voices, scores and judges the voiceprint features to be recognized in the model library, and finds the result with the highest or closest judgment score.
[0024] Currently, with the advent of smart glasses, smart bracelets and other smart wearable devices, wearable devices have entered the lives of modern people. Similar to other smart terminal devices, wearable devices are also equipped with voiceprint recognition modules to meet the needs of different scenarios, such as making voice payments and providing personalized services to users, thereby improving the user interaction experience.
[0025] However, the inventors discovered in their research that if a user is not in a normal posture when wearing a wearable device for voiceprint recognition, such as squatting, leaning back, or bending over, or if the head is turned, such as looking down or turning the head, the user's vocal organs will be compressed to varying degrees, and the sound produced will also change significantly, thus leading to a decrease in the success rate of voiceprint recognition.
[0026] In the scoring process of voiceprint recognition, a threshold is set. If the score is below the threshold, the voiceprint recognition fails. Only when the score exceeds the threshold is an acceptance decision made. In some implementations, the pass rate can be improved by adjusting the threshold. However, the inventors further discovered that different users exhibit different changes in voiceprint characteristics under different postures, making it very difficult to formulate a strategy for adjusting the threshold.
[0027] Therefore, in order to overcome the above-mentioned defects, this application provides a voiceprint recognition method, device, wearable device, and computer-readable medium. When the first voiceprint feature cannot be recognized by voiceprint recognition, that is, when the user's voice is distorted due to changes in posture and cannot be recognized by voiceprint recognition, the posture information of the wearable device is collected, that is, the first posture coordinates are obtained. Based on the pre-acquired voiceprint feature profile, a new voiceprint feature is determined, that is, the second voiceprint feature corresponding to the first posture coordinates. Voiceprint recognition is performed on the second voiceprint feature, which can improve the voiceprint recognition pass rate when the user wears the wearable device.
[0028] Please see Figure 1 , Figure 1 This application illustrates a voiceprint recognition method provided in an embodiment. The method is applied to a wearable device, which may be a Bluetooth headset, a head-mounted virtual reality (VR) device, or an augmented reality (AR) device capable of collecting user voice information. The wearable device may include an audio acquisition device and a location acquisition device. As one implementation, the wearable device may include a processor connected to both the audio acquisition device and the location acquisition device, and the processor may be the execution entity of the method. Specifically, the method includes steps S101 to S105.
[0029] S101: When the wearable device is worn by a user and used in a voiceprint recognition scenario, a first voiceprint feature is obtained based on the user's voice information collected by the audio acquisition device at a specified time.
[0030] In one implementation, the specified time is any moment within a usage cycle when the wearable device is used in a voiceprint recognition scenario. Specifically, the usage cycle can be the entire process from when the wearable device is put on by the user to when it is removed. Furthermore, the wearable device may include a capacitive sensing sensor to detect the proximity / distance between the wearable device and the human body by detecting changes in capacitance, thereby determining the state of being put on / taken off and obtaining the usage cycle. Further, the wearable device may also include a user verification unit capable of confirming the user's identity when the device is detected to be worn. Specifically, the user verification unit may send a specific instruction to the wearable device to confirm the user's identity when a specific operation is detected. Thus, the voice information of the user collected by the audio acquisition device at the specified time can be determined to all originate from the same user.
[0031] As one implementation, the audio acquisition device can be a microphone. Specifically, the audio acquisition device can be a bone conduction microphone. Bone conduction is a sound transmission method that converts sound into mechanical vibrations of different frequencies, which are then transmitted through the skull. Compared to classic sound transmission methods such as air conduction, bone conduction eliminates many steps in sound wave transmission, enabling clear sound reproduction even in noisy environments. Alternatively, the audio acquisition device can be a dual-channel microphone. A dual-channel microphone refers to using a pair of matched microphones at different locations to receive sound signals. Combined with Environmental Noise Cancellation (ENC) technology, it can accurately calculate the speaker's location and remove various interfering noises from the environment. In other words, the audio acquisition device can ignore the sounds of the environment or other people around, ensuring that the acquired voice information comes entirely from the user.
[0032] As one implementation, the voiceprint features can be a set of acoustic description parameters extracted from the sound signal based on computer algorithms / mathematical methods, represented in vector form. Specifically, the acoustic description parameters may include short-time spectrum, pitch period, short-time energy, short-time zero-crossing rate, linear predictive coding (LPC) parameters, and Mel-scale frequency cepstral coefficients (MFCC), etc. Further, if the voiceprint features are represented as a four-dimensional vector composed of four description parameters, for example, the first voiceprint feature can be [1.2, 4.2, 2.1, 3.3].
[0033] S102: Perform a voiceprint recognition operation on the first voiceprint feature.
[0034] In one implementation, the first voiceprint feature can be input into a pre-obtained voiceprint feature model in the device for scoring and judgment to obtain a judgment score. Then, based on a pre-set threshold, it is determined whether the voiceprint recognition passes. If the judgment score is equal to or higher than the threshold, the voiceprint recognition is judged to pass; if the judgment score is lower than the threshold, the voiceprint recognition is judged to fail. The voiceprint feature model can be a random model, which uses a probability density function to simulate the user. The training process involves inputting multiple voice segments provided by the user into the probability density function to predict the parameters of the function, thereby obtaining the user's personalized voiceprint feature model. Further, the random model can be a Gaussian Mixture Model (GMM) or a Hidden Markov Model (HMM), etc. Further, the scoring and judgment operation specifically involves calculating the similarity of the first voiceprint feature of the corresponding model to obtain a judgment score.
[0035] S103: If the voiceprint recognition fails, obtain the first posture coordinates based on the posture information of the wearable device collected by the position acquisition device at a specified time;
[0036] S104: Based on the first posture coordinates, determine the second voiceprint feature corresponding to the first posture coordinates from the pre-acquired voiceprint feature profile, wherein the voiceprint feature profile is used to characterize the correspondence between posture coordinates and voiceprint features.
[0037] As can be seen from the foregoing, if a user is not in a normal posture when wearing a wearable device for voiceprint recognition, such as squatting, leaning back, or bending over, or if their head is turned, such as looking down or turning their head, their vocal organs will be subjected to varying degrees of pressure, resulting in significant changes in the emitted sound. This will cause the first voiceprint feature to fail the voiceprint recognition scoring, leading to a failed voiceprint recognition. Therefore, to improve the success rate of voiceprint recognition when users use wearable devices in different postures, when the first voiceprint feature fails, the first posture coordinates of the wearable device can be obtained. Based on these coordinates, the second voiceprint feature can be estimated, facilitating subsequent voiceprint recognition based on the second voiceprint feature.
[0038] In one implementation, the posture coordinates can be a six-dimensional coordinate system composed of posture coordinates and angle coordinates. Specifically, the posture coordinates can be three-dimensional spatial coordinates, used to characterize the specific position of the wearable device in three-dimensional space at a specified moment. The angle coordinates can be rotation angles including front-back, left-right, and horizontal directions, used to characterize the specific angle of the wearable device relative to the initial angle at a specified moment. Furthermore, the initial angle can be the angle of the wearable device when the user wears it and maintains a natural posture. For example, when the wearable device is a Bluetooth headset, the initial angle can be the angle of the wearable device when the user's head is straight.
[0039] In one implementation, the position acquisition device can be a sensor. Specifically, the position acquisition device may include an infrared sensor and a gyroscope, wherein the infrared sensor can measure the specific position of the wearable device at a specified time, and the gyroscope can measure the angle of the wearable device at a specified time.
[0040] Specifically, the first attitude coordinate can be a six-dimensional coordinate of the wearable device at a specified time, calculated based on the position and angle data of the wearable device collected by the position acquisition device at a specified time, and denoted as (x, y, z, α, β, γ).
[0041] In one implementation, the voiceprint feature profile can be an estimation model trained based on the user's voice information pre-collected by the audio acquisition device and the wearable device's posture information pre-collected by the location acquisition device. This model represents the correspondence between posture coordinates and voiceprint features. Specifically, posture coordinates can be obtained based on posture information, voiceprint features can be obtained based on voice information, and then a random model can be used to train and obtain an estimation model of voiceprint features under posture coordinates. Further, the random model can be a Gaussian Mixture Model (GMM) or a Hidden Markov Model (HMM), etc. In other words, each posture coordinate can be estimated to have a unique corresponding voiceprint feature based on the voiceprint feature profile model. Therefore, based on the first posture coordinate, the corresponding second voiceprint feature can be estimated from the voiceprint feature profile.
[0042] S105: Perform a voiceprint recognition operation on the second voiceprint feature.
[0043] The operation of performing voiceprint recognition on the second voiceprint feature can refer to the operation of performing voiceprint recognition on the first voiceprint feature described above, and will not be repeated here.
[0044] Therefore, the voiceprint recognition method provided in this application, when the wearable device is worn by a user and used in a voiceprint recognition scenario, obtains a first voiceprint feature based on the user's voice information collected by the audio acquisition device at a specified time, and performs a voiceprint recognition operation on the first voiceprint feature. Then, if the wearable device is used in a voiceprint recognition scenario and the user is in an unconventional posture, the emitted sound will change due to the posture change, causing the first voiceprint feature to fail voiceprint recognition. Then, based on the posture information of the wearable device collected by the position acquisition device at a specified time, a first posture coordinate is obtained, and a second voiceprint feature is estimated from a pre-acquired voiceprint feature profile, and a voiceprint recognition operation is performed on the second voiceprint feature. Therefore, when a user wears a wearable device for voiceprint recognition, if the voice distortion caused by the user's unconventional posture prevents the first voiceprint feature obtained directly from the voice information from passing voiceprint recognition, a second voiceprint feature can be estimated based on the voiceprint feature profile and posture coordinates. This results in a second voiceprint feature estimated from the voiceprint feature profile based on the first posture coordinates, which is different from the first voiceprint feature obtained when the user is in an unconventional posture. Voiceprint recognition is then performed on the second voiceprint feature to improve the scoring in the voiceprint recognition process, thereby increasing the voiceprint recognition pass rate when the user wears the wearable device.
[0045] Please see Figure 2 , Figure 2 This application illustrates a voiceprint recognition method provided by an embodiment of the present application. This method is applied to the aforementioned wearable device. Before the wearable device is worn by a user and used in a voiceprint recognition scenario, a voiceprint feature profile is trained and acquired. Then, after a first voiceprint feature fails voiceprint recognition, a second voiceprint feature is estimated from the voiceprint feature profile based on a first pose coordinate, and then voiceprint recognition is performed on the second voiceprint feature. Specifically, the method includes steps S201 to S206.
[0046] S201: Based on the user's voice information pre-collected by the audio acquisition device and the posture information of the wearable device pre-collected by the location acquisition device, train and obtain a voiceprint feature profile.
[0047] S202: When the wearable device is worn by a user and used in a voiceprint recognition scenario, a first voiceprint feature is obtained based on the user's voice information collected by the audio acquisition device at a specified time.
[0048] S203: Perform a voiceprint recognition operation on the first voiceprint feature;
[0049] S204: If the first voiceprint feature cannot be identified by voiceprint, obtain the first posture coordinates based on the posture information of the wearable device collected by the position acquisition device at a specified time;
[0050] S205: Based on the first posture coordinates, estimate the second voiceprint feature from the pre-acquired voiceprint feature profile, wherein the voiceprint feature profile is used to characterize the correspondence between posture coordinates and voiceprint features;
[0051] S206: Perform a voiceprint recognition operation on the second voiceprint feature.
[0052] The implementation methods for steps S202 to S206 can be referred to the aforementioned embodiments, and will not be repeated here.
[0053] As one implementation method, please refer to Figure 3 , Figure 3 This application provides a method for training and obtaining a voiceprint feature profile in step S201, which may include steps S301 to S306.
[0054] S301: When the user is in the first posture position, the user's voice information is collected based on the audio acquisition device to obtain the reference voiceprint features, wherein the first posture position is the position where the user is standing and the wearable device is in the initial angle.
[0055] S302: When the user is in the second posture position, the user's voice information is collected based on the audio acquisition device to obtain posture voiceprint features, wherein the second posture position is any position that does not coincide with the first posture position;
[0056] S303: Based on the reference voiceprint features and the posture voiceprint features, obtain the first voiceprint deviation.
[0057] As one implementation, the initial angle can be the angle of the wearable device when the user wears it and maintains a natural posture. Specifically, when the wearable device is a Bluetooth headset, the initial angle can be the angle of the wearable device when the user's head is upright. That is, the first posture position can be the position where the user is standing with their head upright. Further, the second posture position is different from the first posture position. Specifically, it can mean that at least one coordinate parameter in the second posture coordinates has a different value than the corresponding coordinate parameter value of the reference posture coordinates. The second posture coordinates are the posture coordinates obtained by the position acquisition device based on the user's posture information in the second posture position, and the reference posture coordinates are the posture coordinates obtained by the position acquisition device based on the user's posture information in the first posture position.
[0058] As one implementation method, the user's current posture can be confirmed by receiving instructions from the user side. Specifically, the instructions from the user side can be sent by the user performing a confirmation operation. Further, the confirmation operation can be the user pressing a button on the wearable device or the user pressing a control on the display screen of the wearable device.
[0059] As one implementation, the first voiceprint deviation can be a set of acoustic description parameters obtained by subtracting the remaining voiceprint features from the reference voiceprint features, and represented in vector form. Further, if the voiceprint features are represented by a four-dimensional vector composed of four description parameters, for example, the reference voiceprint features are [1.1, 4.3, 2.2, 3.1] and the posture voiceprint features are [1.2, 4.2, 2.1, 3.3], then the first voiceprint deviation can be obtained as [0.1, -0.1, -0.1, 0.2].
[0060] S304: When the user is in the first posture position, obtain the reference posture coordinates based on the posture information of the wearable device collected by the position acquisition device;
[0061] S305: When the user is in the second posture position, based on the posture information of the wearable device collected by the position acquisition device, the second posture coordinates are obtained, wherein the second posture coordinates are in a reference coordinate system determined by the reference posture coordinates.
[0062] The posture information of the wearable device collected by the location acquisition device is the posture coordinates of the wearable device in the world physical coordinate system. After obtaining the reference posture coordinates, a reference coordinate system is established based on the reference posture coordinates. This reference coordinate system is also a coordinate system in physical space, but its reference point is different from that of the world physical coordinate system. Then, when the user is in the second posture position, the posture coordinates obtained based on the posture information of the wearable device collected by the location acquisition device are also in the world physical coordinate system. Then, these posture coordinates are mapped to the reference coordinate system to obtain the posture coordinates in the reference coordinate system, which are denoted as the second posture coordinates. The process of obtaining the second posture coordinates based on the posture information of the wearable device collected by the location acquisition device can be referred to the previous embodiment, and will not be repeated here.
[0063] S306: Based on the first voiceprint deviation and the second posture coordinates, train to obtain the voiceprint feature profile, which is used to characterize the correspondence between the posture coordinates of the reference coordinate system and the voiceprint features.
[0064] As one implementation method, the training data used to train and acquire the voiceprint feature profile consists of a large amount of sample data obtained when the user is in a second pose position. This sample data includes multiple first voiceprint deviations and the second pose coordinates corresponding to each first voiceprint deviation. It should be noted that the training and acquisition of the voiceprint feature profile based on the first voiceprint deviation and the second pose coordinates can be referred to the foregoing embodiments, and will not be repeated here.
[0065] In one implementation, the value of each coordinate parameter in the second posture coordinates is different from the value of the coordinate parameter corresponding to the reference posture coordinates. Specifically, the inventors found in their research that when training to obtain the voiceprint feature profile based on the first voiceprint deviation and the second posture coordinates, the more dispersed the distribution of the second posture coordinates in the reference coordinate system determined with the reference posture coordinates as the reference, the more accurately the obtained voiceprint feature profile reflects the relationship between the user's posture coordinates and voiceprint features. Furthermore, when determining the second posture position of the user, it is also considered to select a second posture position in which the values of the coordinate parameters of the second posture coordinates are significantly different from those of the coordinate parameters of the reference posture coordinates.
[0066] Please see Figure 4 , Figure 4 This application illustrates a voiceprint recognition method provided in an embodiment of the present application. This method, applied to the aforementioned wearable device, can estimate a second voiceprint feature based on a pre-acquired voiceprint feature profile using a first pose coordinate. Specifically, the method includes steps S401 to S404.
[0067] S401: When the user is in the first posture position, obtain the reference posture coordinates based on the posture information of the wearable device collected by the position acquisition device, wherein the first posture position is the position where the user is standing and the wearable device is in the initial angle.
[0068] S402: When the user is in the first posture position, the user's voice information is collected based on the audio acquisition device to obtain the reference voiceprint features.
[0069] The implementation methods of steps S401 to S402 can be referred to the foregoing embodiments, and will not be repeated here.
[0070] S403: Based on the reference attitude coordinates and the first attitude coordinates, interpolate to obtain the third attitude coordinates.
[0071] Interpolation is an important method for approximating discrete functions. It can be used to estimate approximate values at other points by considering the values at a finite number of points.
[0072] As one implementation method, when interpolating to obtain the third pose coordinates, the interpolation control parameters are related to the application scenario in which the wearable device is used for voiceprint recognition or the object the user interacts with. Specifically, the interpolation control parameters can be the interpolation ratio between the reference pose coordinates and the first pose coordinates. For example, when the wearable device is in a scenario with high security requirements, such as a payment scenario, the interpolation will be more biased towards bringing the third pose coordinates closer to the first pose coordinates. That is, based on the third pose coordinates that are closer to the actual position, the voiceprint feature profile can obtain voiceprint features that are closer to the user's actual voiceprint, which can improve device security. In scenarios with high user experience requirements, such as personalized service scenarios, the interpolation will be more biased towards bringing the third pose coordinates closer to the reference pose coordinates. When the user interacts with an unfamiliar object, the interpolation will be more biased towards bringing the third pose coordinates closer to the first pose coordinates. When the user interacts with a familiar object, the interpolation will be more biased towards bringing the third pose coordinates closer to the reference pose coordinates.
[0073] S404: Based on the third posture coordinates, obtain the second voiceprint feature from the voiceprint feature distribution profile.
[0074] The implementation method of step S404 can be referred to the foregoing embodiments, and will not be repeated here.
[0075] Furthermore, considering that the aforementioned voiceprint feature profile is pre-collected with voice and posture information and trained and initialized in the device before the user uses the wearable device, the amount of collected data is limited, and the accuracy of the resulting voiceprint feature profile is not high. Therefore, after the user wears and uses the wearable device in a voiceprint recognition scenario, updating the voiceprint feature profile based on the second voiceprint feature can improve the accuracy of the user's individual voiceprint feature profile. For details, please refer to [link to relevant documentation]. Figure 5 , Figure 5 An embodiment of this application provides a voiceprint recognition method, which is applied to the aforementioned wearable device. Specifically, the method includes steps S501 to S503.
[0076] S501: Based on the user's voice information pre-collected by the audio acquisition device and the wearable device's posture information pre-collected by the location acquisition device, train and obtain a voiceprint feature profile.
[0077] S502: When the wearable device is worn by a user and used in a voiceprint recognition scenario, a first voiceprint feature is obtained based on the user's voice information collected by the audio acquisition device at a specified time.
[0078] S503: Perform a voiceprint recognition operation on the first voiceprint feature;
[0079] S504: If the first voiceprint feature cannot be identified by voiceprint, obtain the first posture coordinates based on the posture information of the wearable device collected by the position acquisition device at a specified time.
[0080] S505: Based on the first posture coordinates, estimate the second voiceprint feature from the pre-acquired voiceprint feature profile, wherein the voiceprint feature profile is used to characterize the correspondence between posture coordinates and voiceprint features.
[0081] S506: Perform a voiceprint recognition operation on the second voiceprint feature;
[0082] S507: Update the voiceprint feature profile based on the second voiceprint feature.
[0083] The implementation methods of steps S501 to S506 can be referred to the foregoing embodiments, and will not be repeated here.
[0084] As one implementation method, please refer to Figure 6 , Figure 6 This application provides a method for updating the voiceprint feature profile based on the second voiceprint feature in step S507, as provided in an embodiment of the present application. Specifically, the method may include steps S601 to S603.
[0085] S601: Based on the second voiceprint feature and the first voiceprint feature, obtain the second voiceprint deviation.
[0086] In one implementation, the second voiceprint deviation can be a set of acoustic description parameters obtained by subtracting the second voiceprint feature from the first voiceprint feature, based on the first voiceprint feature, and represented in vector form.
[0087] S602: Compare the standard deviation of the Gaussian distribution of the second voiceprint deviation and the first voiceprint deviation. When the second voiceprint deviation is greater than the standard deviation, update the mean of the Gaussian distribution of the first voiceprint deviation.
[0088] As one implementation method, the voiceprint feature profile can be obtained by training a Gaussian Mixture Model (GMM). Specifically, each bit in the parameter vector of the first voiceprint deviation follows a Gaussian distribution, that is, each bit corresponds to the mean μ of a Gaussian distribution, and the standard deviation of the Gaussian distribution is a preset value σ.
[0089] As one implementation, the method for comparing the standard deviation of the Gaussian distribution of the second voiceprint deviation and the first voiceprint deviation can be to compare each bit of the vector in the second voiceprint deviation with the standard deviation σ of the Gaussian distribution of each bit of the vector in the first voiceprint deviation. Furthermore, if the second voiceprint deviation being compared is less than the standard deviation σ, then the mean μ of that bit is not updated.
[0090] As one implementation, the method for updating the mean of the Gaussian distribution of the first voiceprint deviation when the second voiceprint deviation is greater than the standard deviation can be as follows: when the second voiceprint deviation being compared is greater than the standard deviation σ but less than 2σ, update μ to the average of the first voiceprint feature and the second voiceprint feature; when the second voiceprint deviation being compared is greater than 2σ, update μ to the first voiceprint feature.
[0091] As another implementation, the method of updating the mean of the Gaussian distribution of the first voiceprint deviation when the second voiceprint deviation is greater than the standard deviation, which determines whether to update the mean based on the comparison result of the second voiceprint deviation and the standard deviation, and the method of updating the mean based on the first voiceprint feature and the second voiceprint feature, can be selected and adjusted according to the requirements, and will not be elaborated here.
[0092] S603: Update the voiceprint feature profile based on the updated first voiceprint deviation and the second attitude coordinates.
[0093] The implementation method of step S603 can be referred to the foregoing embodiments, and will not be repeated here.
[0094] Please see Figure 7 The diagram shows the structural frame of a voiceprint recognition device 700 provided in an embodiment of this application. The device may include a voiceprint feature acquisition unit 701, a voiceprint recognition unit 702, a posture coordinate acquisition unit 703, and a voiceprint feature analysis unit 704.
[0095] The voiceprint feature acquisition unit 701 is used to acquire voiceprint features based on the user's voice information acquired by the audio acquisition device.
[0096] The voiceprint recognition unit 702 is used to perform voiceprint recognition on the voiceprint features.
[0097] The posture coordinate acquisition unit 703 is used to acquire posture coordinates based on the posture information of the wearable device acquired by the position acquisition device when the voiceprint feature acquired by the voiceprint feature acquisition unit does not pass voiceprint recognition.
[0098] The voiceprint feature analysis unit 704 is used to estimate voiceprint features based on the posture coordinates and a pre-acquired voiceprint feature profile.
[0099] Furthermore, the voiceprint feature analysis unit 704 is also used to train and obtain a voiceprint feature profile based on the user's voice information pre-collected by the audio acquisition device and the posture information of the wearable device pre-collected by the location acquisition device.
[0100] Furthermore, the voiceprint feature analysis unit 704 is also used to update the voiceprint feature profile based on the voiceprint features.
[0101] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0102] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0103] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0104] Please refer to Figure 8 This diagram illustrates a structural block diagram of a wearable device according to an embodiment of this application. The wearable device 800 can be a Bluetooth headset, a head-mounted VR (Virtual Reality) device, or other wearable device capable of collecting user voice information. The wearable device 800 in this application may include one or more of the following components: a processor 810, a memory 820, an audio acquisition device 830, a location acquisition device 840, and one or more applications. The one or more applications may be stored in the memory 820 and configured to be executed by one or more processors 810. The one or more applications are configured to perform the methods described in the foregoing method embodiments. The audio acquisition device 830 may be a microphone, for example, a bone conduction microphone or a dual-channel microphone array. The location acquisition device 840 may be a gyroscope, an infrared sensor, or other similar device.
[0105] Processor 810 may include one or more processing cores. Processor 810 connects to various parts within the wearable device 800 using various interfaces and lines, and performs various functions and processes data of the wearable device 800 by running or executing instructions, programs, code sets, or instruction sets stored in memory 820, and by calling data stored in memory 820. Optionally, processor 810 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 110 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 810 and may be implemented separately through a communication chip. The memory 820 may include random access memory (RAM) or read-only memory (ROM). The memory 820 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 820 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data created by the terminal during use (such as phonebook data, audio and video data, chat log data, etc.).
[0106] Please refer to Figure 9 This diagram illustrates a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. The computer-readable medium 900 stores program code that can be called by a processor to execute the methods described in the above method embodiments.
[0107] The computer-readable storage medium 900 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 900 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 900 has storage space for program code 910 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 910 may be compressed, for example, in a suitable form. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit it; although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A voiceprint recognition method, characterized in that, Applied to wearable devices, the wearable devices including an audio acquisition device and a location acquisition device, the method includes: When the wearable device is worn by a user and used in a voiceprint recognition scenario, a first voiceprint feature is obtained based on the user's voice information collected by the audio acquisition device at a specified time. Perform a voiceprint recognition operation on the first voiceprint feature; If the first voiceprint feature cannot be identified by voiceprint, the first posture coordinates are obtained based on the posture information of the wearable device collected by the position acquisition device at the specified time. Based on the first posture coordinates, the second voiceprint feature is estimated from the pre-acquired voiceprint feature profile, wherein the voiceprint feature profile is used to characterize the correspondence between posture coordinates and voiceprint features. Perform a voiceprint recognition operation on the second voiceprint feature.
2. The method according to claim 1, characterized in that, Before obtaining the first voiceprint feature based on the user's voice information collected by the audio acquisition device at a specified time, the method further includes: Based on the user's voice information pre-collected by the audio acquisition device and the wearable device's posture information pre-collected by the location acquisition device, the voiceprint feature profile is trained and obtained.
3. The method according to claim 2, characterized in that, The process of training and obtaining a voiceprint feature profile based on the user's voice information pre-collected by the audio acquisition device and the wearable device's posture information pre-collected by the location acquisition device includes: When the user is in a first posture position, the audio acquisition device collects the user's voice information to obtain a reference voiceprint feature, wherein the first posture position is the position where the user is standing and the wearable device is in an initial angle. When the user is in the second posture position, the user's voice information is collected based on the audio acquisition device to obtain posture voiceprint features, wherein the second posture position is any position different from the first posture position; Based on the reference voiceprint features and the posture voiceprint features, the first voiceprint deviation is obtained; When the user is in the first posture position, the reference posture coordinates are obtained based on the posture information of the wearable device collected by the position acquisition device; When the user is in the second posture position, the second posture coordinates are obtained based on the posture information of the wearable device collected by the position acquisition device, wherein the second posture coordinates are in a reference coordinate system determined by the reference posture coordinates; Based on the first voiceprint deviation and the second posture coordinates, the voiceprint feature profile is trained and obtained. The voiceprint feature profile is used to characterize the correspondence between the posture coordinates of the reference coordinate system and the voiceprint features.
4. The method according to claim 3, characterized in that, The value of each coordinate parameter in the second attitude coordinate system is different from the value of the coordinate parameter corresponding to the reference attitude coordinate system.
5. A method according to claim 3, characterized in that, After performing voiceprint recognition on the second voiceprint feature, the method further includes: The voiceprint feature profile is updated based on the second voiceprint feature.
6. A method according to claim 5, characterized in that, The step of updating the voiceprint feature profile based on the second voiceprint feature includes: Based on the second voiceprint feature and the first voiceprint feature, the second voiceprint deviation is obtained; The voiceprint feature profile is updated based on the second voiceprint deviation and the first voiceprint deviation.
7. A method according to claim 6, characterized in that, The first voiceprint deviation follows a Gaussian distribution, and the update of the voiceprint feature profile based on the second voiceprint deviation and the first voiceprint deviation further includes: Compare the standard deviation of the Gaussian distribution of the second voiceprint deviation with that of the first voiceprint deviation. If the second voiceprint deviation is greater than the standard deviation, update the mean of the Gaussian distribution of the first voiceprint deviation. The voiceprint feature profile is updated based on the updated first voiceprint deviation and the second posture coordinates.
8. A method according to claim 1, characterized in that, The step of estimating the second voiceprint feature based on the first pose coordinates and the pre-acquired voiceprint feature profile includes: When the user is in the first posture position, a reference posture coordinate is obtained based on the posture information of the wearable device collected by the position acquisition device, wherein the first posture position is the position where the user is standing and the wearable device is in the initial angle. When the user is in the first posture position, the audio acquisition device collects the user's voice information to obtain the reference voiceprint features. Based on the reference attitude coordinates and the first attitude coordinates, the third attitude coordinates are obtained by interpolation; Based on the third pose coordinates, the second voiceprint feature is obtained from the voiceprint feature distribution profile.
9. A voiceprint recognition device, characterized in that, Applied to wearable devices, the wearable device includes an audio acquisition device and a location acquisition device, the device comprising: The voiceprint feature acquisition unit is used to acquire voiceprint features based on the user's voice information collected by the audio acquisition device. A voiceprint recognition unit is used to perform voiceprint recognition on the voiceprint features; The posture coordinate acquisition unit is used to acquire posture coordinates based on the posture information of the wearable device acquired by the position acquisition device when the voiceprint features acquired by the voiceprint feature acquisition unit are not recognized by voiceprint recognition. The voiceprint feature analysis unit is used to estimate voiceprint features based on the posture coordinates and a pre-acquired voiceprint feature profile.
10. A wearable device, characterized in that, include: One or more processors; Memory; Audio acquisition device and location acquisition device; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to perform the method as described in any one of claims 1-8.
11. A computer-readable medium, characterized in that, The computer-readable medium stores processor-executable program code, which, when executed by the processor, causes the processor to perform the method according to any one of claims 1-8.
Citation Information
Patent Citations
Voice recognition method, wearable device, and system
CN112334977A
Data migration method, terminal equipment and readable storage medium
CN113918916A