Audio control method and device
The voice control method enhances voiceprint recognition accuracy by using multiple sensors and dynamic fusion coefficients to capture and process voice components, addressing high-frequency loss and noise interference, thereby improving user authentication and operation efficiency.
Patent Information
- Application Number
- JP2023558328
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-03-24
- Filing Date
- 2022-03-11
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2042-03-11
AI Technical Summary
Current bone vibration sensors in voiceprint recognition systems lose high-frequency components, leading to inaccurate voiceprint recognition, and existing systems struggle with noise interference, reducing accuracy and reliability.
A voice control method utilizing an in-ear voice sensor, an extra-ear voice sensor, and a bone vibration sensor to capture multiple voice components, enhancing high-frequency sound capture and employing dynamic fusion coefficients to improve recognition accuracy, especially in noisy environments.
The method significantly improves voiceprint recognition accuracy by compensating for lost high-frequency components and reducing noise interference, enabling efficient user authentication and operation execution with reduced power consumption.
Smart Images

Figure 0007794374000002 
Figure 0007794374000003 
Figure 0007794374000004
Abstract
Description
[Technical Field]
[0001] The present application relates to the field of audio processing technology, and in particular to a voice control method and device. [Background technology]
[0002] In the prior art, two voice sensors are usually used to capture two audio signals for voiceprint recognition to authenticate the speaking user. In other words, the speaking user is determined to be a preset user only when the voiceprint recognition results of both audio components match. A bone vibration sensor is a common audio sensor. When sound propagates through bones, the bones vibrate. The bone vibration sensor senses the bone vibrations and converts the vibration signals into electrical signals to capture the sound.
[0003] If one of the two audio sensors is a bone vibration sensor, current bone vibration sensors can usually only capture the low-frequency components (usually less than 1 kHz) of the speaker's audio signal, so the high-frequency components are lost, which do not contribute to voiceprint recognition, and therefore the voiceprint recognition is inaccurate. Summary of the Invention [Problem to be solved by the invention]
[0004] The present application provides a voice control method and device to solve the problem that high frequency components are lost and voiceprint recognition is inaccurate when using bone vibration sensors. [Means for solving the problem]
[0005] To achieve the above objectives, the following technical solutions are used in this application:
[0006] According to a first aspect, the present application provides a voice control method including: acquiring voice information of a user, the voice information including a first voice component, a second voice component, and a third voice component, the first voice component being captured by an in-ear voice sensor, the second voice component being captured by an extra-ear voice sensor, and the third voice component being captured by a bone vibration sensor; performing voiceprint recognition on each of the first voice component, the second voice component, and the third voice component; acquiring identification information of the user based on the first voiceprint recognition result of the first voice component, the second voiceprint recognition result of the second voice component, and the third voiceprint recognition result of the third voice component; and executing an operation instruction when the user's identification information matches preset information, the operation instruction being determined based on the voice information.
[0007] After a user wears a wearable device, the external auditory canal and the middle ear canal form a closed cavity, which has a specific amplification effect on the sound within the cavity, i.e., the cavity effect. Therefore, the sound captured by the in-ear sound sensor is clearer, and there is a significant enhancement effect, especially for high-frequency acoustic signals. Because the in-ear sound sensor is used when the wearable device captures sound, it can compensate for the distortion that occurs when some high-frequency signal components of the sound information are lost when the bone vibration sensor captures sound information. Therefore, the overall voiceprint capture effect and voiceprint recognition accuracy of the wearable device can be improved, resulting in an improved user experience.
[0008] Before performing voiceprint recognition, each voice component needs to be acquired. To improve the accuracy and anti-interference capability of voiceprint recognition, multiple voice components are acquired.
[0009] In a possible implementation, before performing voiceprint recognition on the first, second, and third voice components, the method further includes a step of performing keyword detection on the voice information or detecting a user input. Optionally, when the voice information includes a preset keyword, voiceprint recognition is performed on each of the first, second, and third voice components, or when a preset operation input by a user is received, voiceprint recognition is performed on each of the first, second, and third voice components. When the voice information does not include a preset keyword or has not received a preset operation input by a user, this indicates that the user does not currently need voiceprint recognition. In this case, the terminal or wearable device does not need to enable the voiceprint recognition function, and power consumption of the terminal or wearable device is reduced.
[0010] In a possible implementation, before performing keyword detection on the voice information or detecting user input, the method further includes a step of obtaining a wearing state detection result of the wearable device. Optionally, when the wearing state detection result passes, keyword detection is performed on the voice information or user input is detected. When the wearing state detection result does not pass, this means that the user is not currently wearing the wearable device, and of course, there is no need for voiceprint recognition. In this case, the terminal or wearable device does not need to enable the keyword detection function, and power consumption of the terminal or wearable device is reduced.
[0011] In a possible implementation, a specific process for performing voiceprint recognition on the first speech component includes: performing feature extraction on the first voice component to obtain a first voiceprint feature; and calculating a first similarity between the first voiceprint feature and a first enrollment voiceprint feature of the user, where the first enrollment voiceprint feature is obtained by performing feature extraction on the first enrollment voice using a first voiceprint model, and the first enrollment voiceprint feature represents a preset audio feature of the user, which is captured by an in-ear sound sensor. Voiceprint recognition is performed by calculating the similarity, thereby improving the accuracy of the voiceprint recognition.
[0012] In a possible implementation, the specific process of performing voiceprint recognition on the second speech component is: performing feature extraction on the second voice component to obtain a second voiceprint feature; and calculating a second similarity between the second voiceprint feature and a second enrollment voiceprint feature of the user, where the second enrollment voiceprint feature is obtained by performing feature extraction on the second enrollment voice by using a second voiceprint model, and the second enrollment voiceprint feature represents a preset audio feature of the user, which is captured by the extra-aural sound sensor. Voiceprint recognition is performed by calculating the similarity, thereby improving the accuracy of the voiceprint recognition.
[0013] In a possible implementation, the specific process of performing voiceprint recognition on the third speech component is: performing feature extraction on the third voice component to obtain a third voiceprint feature; and calculating a third similarity between the third voiceprint feature and a third enrollment voiceprint feature of the user, where the third enrollment voiceprint feature is obtained by performing feature extraction on the third enrollment voice using a third voiceprint model, and the third enrollment voiceprint feature represents a preset audio feature of the user, which is captured by the bone vibration sensor. Voiceprint recognition is performed by calculating the similarity, thereby improving the accuracy of the voiceprint recognition.
[0014] In one possible implementation, the step of acquiring user identification information based on the voiceprint recognition results of the first, second, and third voice components of the voice information may include: fusing all the voiceprint recognition results using a dynamic fusion coefficient to acquire the user identification information; determining a first fusion coefficient corresponding to the first similarity, a second fusion coefficient corresponding to the second similarity, and a third fusion coefficient corresponding to the third similarity; fusing the first, second, and third similarities based on the first, second, and third fusion coefficients to acquire a fusion similarity score; and determining that the user identification information matches the preset identification information if the fusion similarity score is greater than a first threshold. The method of acquiring a fusion similarity score by performing the steps of fusing and determining multiple similarities can effectively improve voiceprint recognition accuracy.
[0015] In a possible implementation, the steps of determining the first, second, and third fusion coefficients may specifically include the steps of obtaining decibels of the ambient sound based on a sound pressure sensor, determining a playback volume based on a playback signal of a speaker, and determining each of the first, second, and third fusion coefficients based on the decibels of the ambient sound and the playback volume, where the second fusion coefficient is negatively correlated with the decibels of the ambient sound, the first and third fusion coefficients are each negatively correlated with the decibels of the playback volume, and a sum of the first, second, and third fusion coefficients is a fixed value. Optionally, the sound pressure sensor and the speaker are sound pressure sensors and speakers of a wearable device.
[0016] In this embodiment of the present application, when similarities are fused, a dynamic fusion coefficient is used. In different application environments, voiceprint recognition results obtained for voice signals with different attributes are fused by using the dynamic fusion coefficient, so that the voice signals with different attributes compensate each other and improve the robustness and accuracy of voiceprint recognition. For example, in a noisy environment or when playing music using a headset, the recognition accuracy can be significantly improved. Voice signals with different attributes may be understood as voice signals obtained by using different sensors (in-ear voice sensor, extra-ear voice sensor, and bone vibration sensor).
[0017] In a possible implementation, the operation instruction includes an unlock instruction, a payment instruction, a power off instruction, an application starting instruction, or a call instruction. In this way, a user can complete a series of operations, such as user authentication or execution of a specific function, by inputting voice information only once, which can greatly improve user control efficiency and user experience.
[0018] According to a second aspect, the present application provides a voice control method applicable to a wearable device. In other words, the wearable device acquires voice information of a user, where the voice information includes a first voice component, a second voice component, and a third voice component, where the first voice component is captured by an in-ear voice sensor, the second voice component is captured by an extra-ear voice sensor, and the third voice component is captured by a bone vibration sensor. The wearable device performs voiceprint recognition on each of the first voice component, the second voice component, and the third voice component. The wearable device acquires user identification information based on the voiceprint recognition results of the first voice component of the voice information, the second voice component, and the third voice component of the voice information. When the user identification information matches preset information, the wearable device executes an operation instruction, where the operation instruction is determined based on the voice information.
[0019] After a user wears a wearable device, the external auditory canal and the middle ear canal form a closed cavity, which has a specific amplification effect on the sound within the cavity, i.e., a cavity effect. Therefore, the sound captured by the in-ear sound sensor is clearer, and there is a significant enhancement effect, especially for high-frequency acoustic signals. Because the in-ear sound sensor is used when the wearable device captures sound, it can compensate for the distortion that occurs when some high-frequency signal components of the sound information are lost when the bone vibration sensor captures sound information. Therefore, the overall voiceprint capture effect and voiceprint recognition accuracy of the wearable device can be improved, improving the user experience.
[0020] Before a wearable device can perform voiceprint recognition, it must first acquire each voice component. The wearable device acquires the three voice components by using different sensors, including an in-ear voice sensor, an extra-ear voice sensor, and a bone vibration sensor, to improve the accuracy and anti-interference capabilities of voiceprint recognition.
[0021] In a possible implementation, before the wearable device performs voiceprint recognition on the first voice component, the second voice component, and the third voice component, the method further includes: the wearable device performing keyword detection on the voice information or detecting a user input. Optionally, when the voice information includes a preset keyword, the wearable device performs voiceprint recognition on each of the first voice component, the second voice component, and the third voice component, or when the wearable device receives a preset operation input by the user, the wearable device performs voiceprint recognition on each of the first voice component, the second voice component, and the third voice component. When the voice information does not include a preset keyword and does not receive a preset operation input by the user, this indicates that the user does not currently need voiceprint recognition. In this case, the wearable device does not need to enable the voiceprint recognition function, and power consumption of the wearable device is reduced.
[0022] In a possible implementation, before the wearable device performs keyword detection on the voice information or detects user input, the method further includes: obtaining a wearing state detection result of the wearable device. Optionally, when the wearing state detection result is successful, performing keyword detection on the voice information or detecting user input. When the wearing state detection result is unsuccessful, this means that the user is not currently wearing the wearable device, and of course, there is no need for voiceprint recognition. In this case, the wearable device does not need to enable the keyword detection function, and the power consumption of the wearable device is reduced.
[0023] In a possible implementation, the specific process by which the wearable device performs voiceprint recognition on the first speech component is as follows:
[0024] The wearable device performs feature extraction on the first voice component to obtain a first voiceprint feature, and the wearable device calculates a first similarity between the first voiceprint feature and a first enrollment voiceprint feature of the user, where the first enrollment voiceprint feature is obtained by performing feature extraction on the first enrollment voice using a first voiceprint model, and the first enrollment voiceprint feature represents a preset audio feature of the user, which is captured by an in-ear sound sensor. Voiceprint recognition is performed by calculating the similarity, thereby improving the accuracy of the voiceprint recognition.
[0025] In a possible implementation, the specific process by which the wearable device performs voiceprint recognition on the second voice component is as follows:
[0026] The wearable device performs feature extraction on the second voice component to obtain a second voiceprint feature, and the wearable device calculates a second similarity between the second voiceprint feature and a second enrollment voiceprint feature of the user, where the second enrollment voiceprint feature is obtained by performing feature extraction on the second enrollment voice using a second voiceprint model, and the second enrollment voiceprint feature represents a preset audio feature of the user, which is captured by an extra-ear sound sensor. Voiceprint recognition is performed by calculating the similarity, thereby improving the accuracy of the voiceprint recognition.
[0027] In a possible implementation, the specific process by which the wearable device performs voiceprint recognition on the third voice component is as follows:
[0028] The wearable device performs feature extraction on the third voice component to obtain a third voiceprint feature, and the wearable device calculates a third similarity between the third voiceprint feature and a third enrollment voiceprint feature of the user, where the third enrollment voiceprint feature is obtained by performing feature extraction on the third enrollment voice using a third voiceprint model, and the third enrollment voiceprint feature represents a preset audio feature of the user, which is captured by the bone vibration sensor. Voiceprint recognition is performed by calculating the similarity, thereby improving the accuracy of the voiceprint recognition.
[0029] In a possible implementation, the wearable device obtaining the user's identification information based on the voiceprint recognition result of the first voice component of the voice information, the voiceprint recognition result of the second voice component of the voice information, and the voiceprint recognition result of the third voice component of the voice information may specifically be fusing all the voiceprint recognition results by using a dynamic fusion coefficient to obtain the user's identification information, specifically:
[0030] The wearable device may determine a first fusion coefficient corresponding to the first similarity, a second fusion coefficient corresponding to the second similarity, and a third fusion coefficient corresponding to the third similarity, and fuse the first similarity, the second similarity, and the third similarity to obtain a fusion similarity score based on the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient, and determine that the user's identification information matches the preset identification information if the fusion similarity score is greater than a first threshold. In this method of fusing multiple similarities and performing determination to obtain a fusion similarity score, the accuracy of voiceprint recognition can be effectively improved.
[0031] In a possible implementation, the determining of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient by the wearable device may specifically include obtaining decibels of the ambient sound based on a sound pressure sensor, determining a playback volume based on a playback signal from a speaker, and determining each of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient based on the decibels of the ambient sound and the playback volume, where the second fusion coefficient is negatively correlated with the decibels of the ambient sound, the first fusion coefficient and the third fusion coefficient are each negatively correlated with the decibels of the playback volume, and a sum of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient is a fixed value. Optionally, the sound pressure sensor and the speaker are sound pressure sensors and speakers of the wearable device.
[0032] In this embodiment of the present application, when similarities are fused, a dynamic fusion coefficient is used. In different application environments, voiceprint recognition results obtained for voice signals with different attributes are fused by using the dynamic fusion coefficient, so that the voice signals with different attributes compensate each other and improve the robustness and accuracy of voiceprint recognition. For example, in a noisy environment or when playing music using a headset, the recognition accuracy can be significantly improved. Voice signals with different attributes may be understood as voice signals obtained by using different sensors (in-ear voice sensor, extra-ear voice sensor, and bone vibration sensor).
[0033] In a possible implementation, the wearable device sends a command to the terminal, and the terminal executes an operation instruction corresponding to the voice information. The operation instruction may include an instruction to unlock, make a payment, power off, launch an application, or make a call. In this way, the user can complete a series of operations, such as authenticating the user or performing a specific function, by simply inputting voice information once, greatly improving the user's control efficiency and user experience over the wearable device.
[0034] According to a third aspect, the present application provides a voice control method. The voice control method is applied to a terminal. In other words, the voice control method is executed by the terminal. The method is particularly as follows, and includes: a terminal acquires voice information of a user, where the voice information includes a first voice component, a second voice component, and a third voice component, where the first voice component is captured by an in-ear voice sensor, the second voice component is captured by an extra-ear voice sensor, and the third voice component is captured by a bone vibration sensor; the terminal performs voiceprint recognition on each of the first voice component, the second voice component, and the third voice component; the terminal acquires user identification information based on a voiceprint recognition result of the first voice component of the voice information, the voiceprint recognition result of the second voice component of the voice information, and the voiceprint recognition result of the third voice component of the voice information; and when the user identification information matches preset information, the terminal executes an operation instruction, where the operation instruction is determined based on the voice information.
[0035] After a user wears a wearable device, the external auditory canal and the middle ear canal form a closed cavity, which has a specific amplification effect on the sound within the cavity, i.e., a cavity effect. Therefore, the sound captured by the in-ear sound sensor is clearer, and there is a significant enhancement effect, especially for high-frequency acoustic signals. Because the in-ear sound sensor is used when a wearable device captures sound, it can compensate for the distortion that occurs when a bone vibration sensor captures sound information and some high-frequency signal components of the sound information are lost. This can improve the overall voiceprint capture effect and voiceprint recognition accuracy of the device, thereby improving the user experience.
[0036] In a possible implementation, after obtaining the voice information input by the user, the wearable device transmits voice components corresponding to the voice information to the terminal, so that the terminal performs voiceprint recognition based on the voice components. When the voice control method is performed on the terminal side, the computing power of the terminal can be effectively utilized, so that the accuracy of identity authentication can be guaranteed even when the computing power of the wearable device is insufficient.
[0037] Before the wearable device can perform voiceprint recognition, the terminal must first acquire each voice component. The wearable device acquires three voice components by using different sensors, including an in-ear voice sensor, an extra-ear voice sensor, and a bone vibration sensor, and transmits the three voice components to the terminal, improving the accuracy and anti-interference capability of the terminal's voiceprint recognition.
[0038] In a possible implementation, before the terminal performs voiceprint recognition on the first voice component, the second voice component, and the third voice component, the method further includes: performing keyword detection on the voice information or detecting a user input. Optionally, when the voice information includes a preset keyword, the wearable device transmits a voice component corresponding to the voice information to the terminal, and the terminal performs voiceprint recognition on each of the first voice component, the second voice component, and the third voice component, or when the terminal receives a preset operation input by the user, the terminal performs voiceprint recognition on each of the first voice component, the second voice component, and the third voice component. When the voice information does not include a preset keyword and does not receive a preset operation input by the user, this indicates that the user does not currently need voiceprint recognition. In this case, the wearable device does not need to enable the voiceprint recognition function, and power consumption of the terminal is reduced.
[0039] In a possible implementation, before the wearable device performs keyword detection on the voice information or detects user input, the method further includes: obtaining a wearing state detection result of the wearable device. Optionally, when the wearing state detection result is successful, performing keyword detection on the voice information or detecting user input. When the wearing state detection result is unsuccessful, this means that the user is not currently wearing the wearable device, and of course, there is no need for voiceprint recognition. In this case, the wearable device does not need to enable the keyword detection function, and the power consumption of the wearable device is reduced.
[0040] In a possible implementation, the specific process by which the terminal performs voiceprint recognition on the first speech component is as follows:
[0041] The terminal performs feature extraction on the first voice component to obtain a first voiceprint feature, and the terminal calculates a first similarity between the first voiceprint feature and a first enrollment voiceprint feature of the user, where the first enrollment voiceprint feature is obtained by performing feature extraction on the first enrollment voice by using a first voiceprint model, and the first enrollment voiceprint feature represents a preset audio feature of the user, which is captured by an in-ear sound sensor. Voiceprint recognition is performed by calculating the similarity, thereby improving the accuracy of the voiceprint recognition.
[0042] In a possible implementation, the specific process by which the terminal performs voiceprint recognition on the second speech component is as follows:
[0043] The terminal performs feature extraction on the second voice component to obtain a second voiceprint feature, and the terminal calculates a second similarity between the second voiceprint feature and the user's second enrollment voiceprint feature, where the second enrollment voiceprint feature is obtained by performing feature extraction on the second enrollment voice using a second voiceprint model, and the second enrollment voiceprint feature represents the user's preset audio feature, which is captured by the extra-ear audio sensor. Voiceprint recognition is performed by calculating the similarity, thereby improving the accuracy of the voiceprint recognition.
[0044] In a possible implementation, the specific process by which the terminal performs voiceprint recognition on the third speech component is as follows:
[0045] The terminal performs feature extraction on the third voice component to obtain a third voiceprint feature, and the terminal calculates a third similarity between the third voiceprint feature and the user's third enrollment voiceprint feature, where the third enrollment voiceprint feature is obtained by performing feature extraction on the third enrollment voice using a third voiceprint model, and the third enrollment voiceprint feature represents the user's preset audio feature, which is captured by the bone vibration sensor. Voiceprint recognition is performed by calculating the similarity, thereby improving the accuracy of the voiceprint recognition.
[0046] In a possible implementation, the terminal obtaining the user's identification information based on the voiceprint recognition result of the first voice component of the voice information, the voiceprint recognition result of the second voice component of the voice information, and the voiceprint recognition result of the third voice component of the voice information may specifically be fusing all the voiceprint recognition results by using a dynamic fusion coefficient to obtain the user's identification information, specifically:
[0047] The terminal may determine a first fusion coefficient corresponding to the first similarity, a second fusion coefficient corresponding to the second similarity, and a third fusion coefficient corresponding to the third similarity, and fuse the first similarity, the second similarity, and the third similarity to obtain a fusion similarity score based on the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient, and determine that the user's identification information matches the preset identification information if the fusion similarity score is greater than a first threshold. In the method of fusing multiple similarities and performing determination to obtain a fusion similarity score, the accuracy of voiceprint recognition can be effectively improved.
[0048] In a possible implementation, the terminal's determination of the first, second, and third fusion coefficients may specifically involve the wearable device acquiring the decibels of the ambient sound based on a sound pressure sensor and determining a playback volume based on a playback signal from a speaker. After detecting the decibels of the ambient sound and the playback volume, the wearable device transmits data to the terminal. The terminal determines the first, second, and third fusion coefficients based on the decibels of the ambient sound and the playback volume. The second fusion coefficient is negatively correlated with the decibels of the ambient sound, and the first and third fusion coefficients are each negatively correlated with the decibels of the playback volume, and the sum of the first, second, and third fusion coefficients is a fixed value. Optionally, the sound pressure sensor and speaker are the sound pressure sensor and speaker of the wearable device.
[0049] In this embodiment of the present application, when similarities are fused, a dynamic fusion coefficient is used. In different application environments, voiceprint recognition results obtained for voice signals with different attributes are fused by using the dynamic fusion coefficient, so that the voice signals with different attributes compensate each other and improve the robustness and accuracy of voiceprint recognition. For example, in a large noise environment or when playing music using a headset, the recognition accuracy can be significantly improved. Voice signals with different attributes may be understood as voice signals obtained by using different sensors (in-ear voice sensor, extra-ear voice sensor, and bone vibration sensor).
[0050] In a possible implementation, the terminal executes an operation instruction corresponding to the voice information, such as an instruction to unlock, make a payment, power off, launch an application, or make a call. In this way, the user can complete a series of operations, such as authenticating the user or performing a specific function of the wearable device, by simply inputting the voice information once, greatly improving the user's control efficiency and user experience over the terminal.
[0051] According to a fourth aspect, the present application provides a voice control device including: a voice information acquisition unit configured to acquire voice information of a user, the voice information including a first voice component, a second voice component, and a third voice component, the first voice component being captured by an in-ear voice sensor, the second voice component being captured by an extra-ear voice sensor, and the third voice component being captured by a bone vibration sensor; a recognition unit configured to perform voiceprint recognition on each of the first voice component, the second voice component, and the third voice component; an identification information acquisition unit configured to acquire identification information of a user based on a voiceprint recognition result of the first voice component, the voiceprint recognition result of the second voice component, and the voiceprint recognition result of the third voice component; and an execution unit configured to execute an operation instruction when the identification information of the user matches preset information, the operation instruction being determined based on the voice information.
[0052] After a user wears a wearable device, the external auditory canal and the middle ear canal form a closed cavity, which has a specific amplification effect on the sound within the cavity, i.e., the cavity effect. Therefore, the sound captured by the in-ear sound sensor is clearer, with a significant enhancement effect, especially for high-frequency acoustic signals. The in-ear sound sensor is used when a wearable device captures sound, which can compensate for the distortion that occurs when some high-frequency signal components of the sound information are lost when the bone vibration sensor captures sound information. This can improve the overall voiceprint capture effect and voiceprint recognition accuracy of the wearable device, thereby improving the user experience. Before a voiceprint recognition result can be obtained, each voice component needs to be captured. Capturing multiple voice components improves the accuracy and anti-interference capabilities of voiceprint recognition.
[0053] In a possible implementation, the voice information acquiring unit is further configured to perform keyword detection on the voice information or detect user input. Optionally, when the voice information includes a preset keyword, voice recognition is performed on each of the first voice component, the second voice component, and the third voice component, or when a preset operation input by a user is received, voice recognition is performed on each of the first voice component, the second voice component, and the third voice component. When the voice information does not include the preset keyword and has not received a preset operation input by a user, this indicates that the user does not currently need voiceprint recognition. In this case, the terminal or wearable device does not need to enable the voiceprint recognition function, and power consumption of the terminal or wearable device is reduced.
[0054] In a possible implementation, the voice information acquisition unit is further configured to acquire a wearing state detection result of the wearable device. Optionally, when the wearing state detection result is successful, perform keyword detection on the voice information or detect user input. When the wearing state detection result is unsuccessful, this means that the user is not currently wearing the wearable device, and of course, there is no need for voiceprint recognition. In this case, the terminal or wearable device does not need to enable the keyword detection function, and the power consumption of the terminal or wearable device is reduced.
[0055] In a possible implementation, the recognition unit is specifically configured to: perform feature extraction on a first speech component to obtain a first voiceprint feature, and calculate a first similarity between the first voiceprint feature and a first enrollment voiceprint feature of the user, where the first enrollment voiceprint feature is obtained by performing feature extraction on the first enrollment speech using a first voiceprint model, where the first enrollment voiceprint feature represents a preset audio feature of the user, the preset audio feature being captured by an in-ear audio sensor; perform feature extraction on a second speech component to obtain a second voiceprint feature, and calculate a second similarity between the second voiceprint feature and a second enrollment voiceprint feature of the user, where the second enrollment voiceprint feature represents a preset audio feature of the user, the preset audio feature being captured by an in-ear audio sensor; the voiceprint recognition system is configured to: perform feature extraction on a second enrollment voice by using a third voice model, the second enrollment voiceprint feature representing a preset audio feature of the user, the preset audio feature being captured by the extra-aural sound sensor; perform feature extraction on a third voice component to obtain a third voiceprint feature, and calculate a third similarity between the third voiceprint feature and the third enrollment voiceprint feature of the user, the third enrollment voiceprint feature being obtained by performing feature extraction on the third enrollment voice by using a third voiceprint model, the third enrollment voiceprint feature representing a preset audio feature of the user, the preset audio feature being captured by the bone vibration sensor. Voiceprint recognition is performed by calculating the similarity, thereby improving the accuracy of the voiceprint recognition.
[0056] In a possible implementation, the identity acquisition unit may acquire the identity by using a dynamic fusion coefficient, specifically determining a first fusion coefficient corresponding to a first similarity, a second fusion coefficient corresponding to a second similarity, and a third fusion coefficient corresponding to a third similarity, fusing the first similarity, the second similarity, and the third similarity based on the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient to obtain a fusion similarity score, and determining that the user's identity matches the preset identity if the fusion similarity score is greater than a first threshold. The method of fusing multiple similarities and determining a fusion similarity score can effectively improve the accuracy of voiceprint recognition.
[0057] In a possible implementation, the identification information acquisition unit is specifically configured to acquire decibels of ambient sound based on a sound pressure sensor, determine a playback volume based on a playback signal from a speaker, and determine a first fusion coefficient, a second fusion coefficient, and a third fusion coefficient based on the decibels of ambient sound and the playback volume, wherein the second fusion coefficient is negatively correlated with the decibels of ambient sound, the first fusion coefficient and the third fusion coefficient are each negatively correlated with the decibels of the playback volume, and the sum of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient is a fixed value.
[0058] In this embodiment of the present application, when similarities are fused, a dynamic fusion coefficient is used. In different application environments, voiceprint recognition results obtained for voice signals with different attributes are fused by using the dynamic fusion coefficient, so that the voice signals with different attributes compensate each other and improve the robustness and accuracy of voiceprint recognition. For example, in a large noise environment or when playing music using a headset, the recognition accuracy can be significantly improved. Voice signals with different attributes may be understood as voice signals obtained by using different sensors (in-ear voice sensor, extra-ear voice sensor, and bone vibration sensor).
[0059] In a possible implementation, if the user is a preset user, the execution unit is specifically configured to execute an operation instruction corresponding to the voice information. The operation instruction includes an unlock instruction, a payment instruction, a power off instruction, an application launch instruction, or a call instruction. In this way, the user can complete a series of operations, such as user authentication or execution of a specific function, by only inputting the voice information once, thereby greatly improving the user's control efficiency and user experience.
[0060] The voice control device provided in the fourth aspect of the present application can be understood as a terminal or a wearable device, and the specific details can be understood depending on the entity that executes the voice control method, but this is not limited in the present application.
[0061] According to a fifth aspect, the present application provides a wearable device including an in-ear sound sensor, an out-of-ear sound sensor, a bone vibration sensor, a memory, and a processor. The in-ear sound sensor is configured to capture a first sound component of audio information, the out-of-ear sound sensor is configured to capture a second sound component of the audio information, and the bone vibration sensor is configured to capture a third sound component of the audio information. The memory is coupled to the processor. The memory is configured to store computer program code. The computer program code includes computer instructions. When the processor executes the computer instructions, the wearable device performs a voice control method according to any one of the first aspect or possible implementations of the first aspect or the third aspect or possible implementations of the third aspect.
[0062] According to a sixth aspect, the present application provides a terminal including a memory and a processor. The memory is coupled to the processor. The memory is configured to store computer program code. The computer program code includes computer instructions. When the processor executes the computer instructions, the terminal performs a voice control method according to any one of the first aspect or possible implementations of the first aspect or the third aspect or possible implementations of the third aspect.
[0063] According to a seventh aspect, the present application provides a chip system for use in an electronic device, the chip system including one or more interface circuits and one or more processors, the interface circuits and the processors being connected to each other via lines, the interface circuits being configured to receive signals from a memory of the electronic device and send the signals to the processor, the signals including computer instructions stored in the memory, and when the processor executes the computer instructions, the electronic device performs the voice control method according to the first aspect or any one of the possible implementations of the first aspect.
[0064] According to an eighth aspect, the present application provides a computer storage medium comprising computer instructions that, when executed on a voice control device, enable the voice control device to perform a voice control method according to the first aspect or any one of the possible implementations of the first aspect.
[0065] According to a ninth aspect, the present application provides a computer program product, the computer program product including computer instructions that, when executed on a voice control device, enable the voice control device to perform a voice control method according to the first aspect or any one of the possible implementations of the first aspect.
[0066] It can be understood that the wearable device according to the fifth aspect, the terminal according to the sixth aspect, the chip system according to the seventh aspect, the computer storage medium according to the eighth aspect, and the computer program product according to the ninth aspect are each configured to execute the corresponding methods provided above. Therefore, for advantageous effects achieved by the wearable device according to the fifth aspect, the terminal according to the sixth aspect, the chip system according to the seventh aspect, the computer storage medium according to the eighth aspect, and the computer program product according to the ninth aspect, please refer to the advantageous effects of the corresponding methods above. Details will not be described again here. [Brief explanation of the drawings]
[0067] [Figure 1] 1 is a schematic diagram of the hardware structure of a mobile phone according to an embodiment of the present invention;
[0068] [Figure 2] 2 is a schematic diagram of the software structure of a mobile phone according to one embodiment of the present invention;
[0069] [Figure 3] 1 is a schematic diagram of the structure of a wearable device according to an embodiment of the present application;
[0070] [Figure 4] 1 is a schematic diagram of a voice control system according to an embodiment of the present application;
[0071] [Figure 5] FIG. 2 is a schematic diagram of the structure of a server according to an embodiment of the present invention.
[0072] [Figure 6] 1 is a schematic flowchart of voiceprint recognition according to an embodiment of the present application;
[0073] [Figure 7] 1 is a schematic diagram of a voice control method according to an embodiment of the present application;
[0074] [Figure 8] FIG. 2 is a schematic diagram of a sensor placement area according to one embodiment of the present invention.
[0075] [Figure 9] FIG. 1 is a schematic diagram of a payment interface according to an embodiment of the present application.
[0076] [Figure 10] FIG. 1 is a schematic diagram of another voice control method according to an embodiment of the present application.
[0077] [Figure 11]FIG. 2 is a schematic diagram of a setting interface of a mobile phone according to one embodiment of the present invention.
[0078] [Figure 12] 1 is a schematic diagram of a voice control device according to one embodiment of the present invention;
[0079] [Figure 13] FIG. 1 is a schematic diagram of a wearable device according to an embodiment of the present application.
[0080] [Figure 14] 1 is a schematic diagram of a terminal according to an embodiment of the present application;
[0081] [Figure 15] FIG. 1 is a schematic diagram of a chip system according to an embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION
[0082] The following describes the technical solutions in the present application with reference to the accompanying drawings. It is clear that the described embodiments are only a part, but not all, of the embodiments of the present application. It can be understood by those skilled in the art that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application can also be applied to similar technical problems.
[0083] The terms "first" and "second" referred to below are intended for descriptive purposes only and should not be understood as an indication or implication of relative importance or the quantity of the indicated technical features. Accordingly, a feature qualified as "first" or "second" may explicitly or implicitly include one or more features. Data used in such methods are interchangeable where appropriate, and it is therefore understood that the embodiments described herein may be implemented in an order other than that shown or described herein. In describing embodiments, unless otherwise specified, "plurality" means two or more. Additionally, the terms "comprise" and "have" and other variations are intended to cover non-exclusive inclusions; for example, a process, method, system, product, or device that includes a list of steps or modules is not necessarily limited to those steps or modules and may include other steps or modules not expressly listed or inherent in such process, method, product, or device. The names or numbers of steps in this application do not imply that the steps in a method procedure must be performed in the temporal / logical sequence indicated by the names or numbers. The sequence of execution of steps in a named or numbered procedure may be varied based on the technical goal to be achieved, provided that the same or similar technical effect can be achieved.
[0084] With the development of audio processing technology, voiceprint recognition has become an important hot topic in the audio processing field. A voiceprint is a sound wave spectrum that carries audio information and is displayed by electroacoustic equipment. Voiceprints can be stable, measurable, unique, etc. In adults, a person's voice can remain stable for a long period of time. The size and shape of the vocal tract used when people speak vary greatly from one another. Therefore, the voiceprint graphs of any two people are different, and the distribution of resonance peaks in the spectrograms of different people's sounds is different. Voiceprint recognition involves comparing two speakers' utterances of the same phoneme to determine whether the speakers are the same person, thereby implementing the function of "recognizing people by listening to their voices."
[0085] Voiceprint recognition (VR), also known as speaker recognition, is a biometric recognition technology that extracts voiceprint information from a speech signal provided by a speaker. From an application perspective, voiceprint recognition may include: Speaker Identification (SI), where speaker identification is used to determine a specific person speaking a specific voice among multiple people, which is a "multisizemotomus" problem; and Speaker Verification (SV), where speaker verification is used to confirm whether a specific voice is spoken by a specific person, which is a "one-to-one decision" problem. This application mainly relates to speaker verification technology.
[0086] The voiceprint recognition technology may be applied to a terminal user identification scenario, and may also be applied to a householder identification scenario for home security, which is not limited in this application.
[0087] In typical voiceprint recognition technologies, voiceprint recognition is performed by capturing one or two voice signals. Specifically, a user is determined as a preset user only when the voiceprint recognition results of both voice components match. However, there are two problems. Problem 1 is that voice components captured in a multi-person speaking scenario or in a background with strong interfering environmental noise may interfere with the voiceprint recognition results, resulting in inaccurate or even incorrect identity authentication. When voice components are captured in an interfering environment, voiceprint recognition performance deteriorates and the identity authentication result is incorrectly determined. That is, existing voiceprint recognition technologies cannot effectively suppress noise from various directions, resulting in reduced voiceprint recognition accuracy.
[0088] Problem 2: If one of the two sound sensors is a bone vibration sensor, current bone vibration sensors can usually only capture the low-frequency components of the speaker's sound signal (usually below 1 kHz), so the high-frequency components are lost. This is not conducive to voiceprint recognition, and therefore the voiceprint recognition is inaccurate and even incorrect, because voiceprint recognition requires describing the speaker's voice-making characteristics in each frequency band.
[0089] In consideration of this, an embodiment of the present application provides a voice control method. It can be understood that the method in this embodiment may be performed by a terminal. The terminal can establish a connection to a wearable device, obtain voice information captured by the wearable device, and perform voiceprint recognition on the voice information. The method in this embodiment may instead be performed by the wearable device. The wearable device includes a processor with computing capabilities and can directly perform voiceprint recognition on the captured voice information. The method in this embodiment may also be performed by a server. The server can establish a connection to the wearable device, obtain voice information captured by the wearable device, and perform voiceprint recognition on the voice information. In actual application, the execution entity of the method in this embodiment may be determined based on the computing capabilities of the chip in the wearable device. For example, when the computing capabilities of the chip in the wearable device are high, the wearable device may execute the method in this embodiment. When the computing capabilities of the chip in the wearable device are low, a terminal device connected to the wearable device may execute the method in this embodiment, or a server connected to the wearable device may execute the method in this embodiment. For ease of explanation, the following describes this embodiment of the present application in detail using an example in which a terminal connected to a wearable device performs the method in this embodiment, an example in which a wearable device performs the method in this embodiment, and an example in which a server connected to a wearable device performs the method in this embodiment.
[0090] A terminal device, also referred to as a user equipment (UE), a mobile station (MS), a mobile terminal (MT), etc., is a device that can connect wired or wirelessly to a wearable device to provide a user with voice and / or data connectivity, such as a handheld device or an in-vehicle device with a wireless connection function. Currently, some examples of terminal devices include a mobile phone, a tablet computer, a notebook computer, a palmtop computer, a mobile internet device (MID), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc. This is not limited to this embodiment of the present application.
[0091] When the voice control method is performed by a terminal, the voice control method may be implemented by using an application installed on the terminal and for recognizing voiceprints.
[0092] The application used for voiceprint recognition may be an embedded application installed on the terminal (i.e., a system application of the terminal) or a downloadable application. An embedded application is an application provided as part of the terminal (e.g., a mobile phone). A downloadable application is an application that can provide Internet Protocol Multimedia Subsystem (IMS) connectivity for downloadable applications. A downloadable application is an application that may be pre-installed on the terminal or a third-party application that may be downloaded by the user and installed on the terminal.
[0093] For ease of understanding, the following will first describe a terminal, a wearable device, and a server to which the method in the embodiment of the present application is applied. Please refer to FIG. 1. An example in which the terminal is a mobile phone is used. FIG. 1 shows the hardware structure of the mobile phone. As shown in FIG. 1, the mobile phone 10 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) port 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, a subscriber identification module (SIM) card interface 195, etc.
[0094] The sensor module 180 may include a pressure sensor 180A, a gyro sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, an optical proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, and the like.
[0095] It can be understood that the structure shown in this embodiment of the present application does not constitute a specific limitation on the mobile phone. In some other embodiments of the present application, the mobile phone may include more or fewer components than those shown in the drawings, or may combine some components, or may separate some components, or may have a different component arrangement. The components shown in the drawings may be implemented by hardware, software, or a combination of software and hardware.
[0096] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent components or may be integrated into one or more processors. The processor 110 may execute the voiceprint recognition algorithm provided in the embodiments of the present application.
[0097] The controller can be the central and command center of the mobile phone, and can generate operation control signals according to the instruction operation code and the time series signal, and complete the control of the instruction reading and execution.
[0098] Memory may also be located within processor 110 and configured to store instructions and data. In some embodiments, the memory within processor 110 is cache memory. The memory may store instructions or data that have been used or periodically used by processor 110. When processor 110 needs to use the instructions or data again, the processor may retrieve the instructions or data directly from memory. This avoids repeated accesses, reduces latency for processor 110, and improves system efficiency.
[0099] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identification module (SIM) interface, a universal serial bus (USB) port, and / or the like. The terminal may establish a wired communication connection to the wearable device via the interface. The terminal may obtain, via the interface, a first sound component captured by the wearable device using an in-ear sound sensor, a second sound component captured using an extra-ear sound sensor, and a third sound component captured using a bone vibration sensor.
[0100] The I2C interface is a bidirectional synchronous serial bus and includes a serial data line (SDA) and a serial clock line (SCL). The I2S interface can be configured for audio communication. The PCM interface can also be used to perform audio communication and sample, quantize, and encode analog signals. The UART interface is a universal serial data bus and is configured for asynchronous communication. The bus may be a bidirectional communication bus. The UART interface converts transmitted data between serial and parallel communication. The MIPI interface can be configured to connect the processor 110 to peripheral components such as the display 194 or the camera 193. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal. The USB port 130 is an interface compliant with the USB standard specification and can be, for example, a mini USB port, a micro USB port, a USB Type-C port, etc. The USB port 130 may be configured to connect to a charger to charge a mobile phone, or to transmit data between a mobile phone and a peripheral device, or to connect to a headset to play audio by using the headset. The interface may further be configured to connect to another electronic device, for example an AR device.
[0101] It can be understood that the interface connection relationships between modules shown in this embodiment of the present application are merely examples for explanation and do not constitute limitations on the structure of the mobile phone. In some other embodiments of the present application, different interface connection methods or combinations of multiple interface connection methods in the above embodiments may alternatively be used for the mobile phone.
[0102] The charging management module 140 is configured to receive charging input from a charger. The charger may be a wireless charger or a wired charger. The power management module 141 is configured to be connected to the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to provide power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, the wireless communication module 160, etc. The power management module 141 may be further configured to monitor parameters such as battery capacity, battery cycle count, or battery health status (electrical leakage or impedance).
[0103] The wireless communication function of the mobile phone may be implemented by using antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, baseband processor, and the like.
[0104] Antenna 1 and Antenna 2 are configured to transmit and receive electromagnetic signals. Each antenna of the mobile phone may be configured to cover one or more communication frequency bands. Different antennas may be further multiplexed to improve antenna utilization. For example, Antenna 1 may be multiplexed as a diversity antenna in a wireless local area network. In some other embodiments, the antennas may be used in combination with tuning switches.
[0105] Mobile communication module 150 may provide wireless communication solutions including 2G / 3G / 4G / 5G, etc., for mobile phones. Mobile communication module 150 may include at least one filter, switch, power amplifier, low-noise amplifier (LNA), etc. Mobile communication module 150 may receive electromagnetic waves via antenna 1, process the received electromagnetic waves, such as filtering and amplifying them, and send the electromagnetic waves to a modem processor for demodulation. Mobile communication module 150 further amplifies signals modulated by the modem processor and converts the signals into electromagnetic waves for emission via antenna 1. In some embodiments, at least some functional modules in mobile communication module 150 may be located within processor 110. In some embodiments, at least some functional modules of mobile communication module 150 may be located in the same device as at least some modules of processor 110. The modem processor may include a modulator and a demodulator.
[0106] The wireless communication module 160 is applied to a mobile phone and may provide wireless communication solutions, including wireless local area networks (WLANs) (e.g., wireless fidelity (WI-FI) networks), Bluetooth (BT), GNSS, frequency modulation (FM), near field communication (NFC) technology, and infrared (IR) technology. The wireless communication module 160 may be one or more components integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs modulation and filtering on the electromagnetic wave signals, and transmits the processed signals to the processor 110. The wireless communication module 160 may further receive signals to be transmitted from the processor 110, perform frequency modulation and amplification on the signals, and convert the signals into electromagnetic waves for emission via the antenna 2. The terminal may establish a communication connection to the wearable device by using the wireless communication module 160. The terminal may acquire, via the wireless communication module 160, a first sound component captured by the wearable device by using an in-ear sound sensor, a second sound component captured by using an extra-ear sound sensor, and a third sound component captured by using a bone vibration sensor.
[0107] For example, in this embodiment of the present application, GNSS may include GPS, GLONASS, BDS, QZSS, SBAS, and / or GALILEO.
[0108] The mobile phone implements display functionality by using a GPU, a display 194, an application processor, etc. The GPU is a microprocessor for image processing and is connected to the display 194 and the application processor. The GPU is configured to perform mathematical and geometric calculations and render images. The processor 110 may include one or more GPUs and may execute program instructions to generate or modify display information. The display 194 is configured to display images, videos, etc. The display 194 includes a display panel.
[0109] The mobile phone may implement a photography function by using an ISP, a camera 193, a video codec, a GPU, a display 194, an application processor, etc. The ISP may be configured to process data fed back by the camera 193. The camera 193 is configured to capture still images or video. An optical image of an object is generated through a lens and projected onto a photosensitive element. The digital signal processor 7 is configured to process digital signals and may process other digital signals in addition to the digital image signal. The video codec is configured to compress or decompress digital video.
[0110] An NPU is a neural network (NN) computing processor. Based on the structure of biological neural networks, such as the communication mode between human brain neurons, the NPU can rapidly process input information and can also continuously self-learn. The NPU can be used to implement applications such as intelligent recognition in mobile phones, such as image recognition, face recognition, speech recognition, and text understanding.
[0111] The external memory interface 120 can be configured to connect to an external memory card, such as a microSD card, to expand the storage capabilities of the mobile phone. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. Files such as music and videos are stored on the external memory card.
[0112] The internal memory 121 may be configured to store computer-executable program code. The executable program code includes instructions. The processor 110 executes the instructions stored in the internal memory 121 to execute various functional applications and data processing of the mobile phone. The code stored in the internal memory 121 may be used to implement a voice control method provided in an embodiment of the present application. For example, when a user inputs voice information into the wearable device, the wearable device captures a first voice component using an in-ear voice sensor, a second voice component using an out-of-ear voice sensor, and a third voice component using a bone vibration sensor. The mobile phone obtains the first, second, and third voice components from the wearable device via a communication connection, performs voiceprint recognition on each of the first, second, and third voice components, and authenticates the user based on the first voiceprint recognition result of the first voice component, the second voiceprint recognition result of the second voice component, and the third voiceprint recognition result of the third voice component. If the user is identified as a preset user as a result of the user authentication, the mobile phone executes an operation instruction corresponding to the voice information.
[0113] The mobile phone may implement audio functionality by using an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset jack 170D, an application processor, etc. The terminal may establish a communication connection to the wearable device by using a wireless communication module 160. The terminal may obtain, via the wireless communication module 160, a first sound component captured by the wearable device by using an in-ear sound sensor, a second sound component captured by using an out-of-ear sound sensor, and a third sound component captured by using a bone vibration sensor.
[0114] The audio module 170 is configured to convert digital audio information into analog audio signals for output and also to convert analog audio input into digital audio signals. The speaker 170A, also called a "horn," is configured to convert audio electrical signals into sound signals. The receiver 170B, also called an "earpiece," is configured to convert audio electrical signals into sound signals. The microphone 170C, also called a "microphone" or "mic," is configured to convert sound signals into electrical signals. The headset jack 170D is configured to connect to a wired headset. The headset jack 170D may be the USB port 130 or a 3.2 mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0115] The buttons 190 include a power button, a volume button, etc. The buttons 190 may be mechanical buttons or touch buttons. The mobile phone may receive key inputs and generate key signal inputs related to user settings and function inputs of the mobile phone. The motor 191 may generate vibration prompts. The motor 191 may be configured to provide incoming call vibration prompts and touch vibration feedback. The indicator 192 may be an indicator and may be configured to indicate a charging status and power change, or may be configured to indicate a message, a missed call, a notification, etc. The SIM card interface 195 is configured to be connected to a SIM card. The SIM card may be inserted into or removed from the SIM card interface 195 to implement contact with or separation from the mobile phone. The mobile phone may support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 may support a nano-SIM card, a micro-SIM card, a SIM card, etc.
[0116] 1, the mobile phone 100 may further include a camera, a flash, a microprojection device, a near field communication (NFC) device, etc., which will not be described in detail here.
[0117] A layered architecture, an event-driven architecture, a microscopic architecture, a microservice architecture, or a cloud architecture can be used for the software system of the mobile phone. In this embodiment of the present application, the layered architecture of the Android system is used as an example to describe the software structure of the mobile phone.
[0118] FIG. 2 is a block diagram of the software structure of a mobile phone according to one embodiment of the present application.
[0119] In a layered architecture, software is divided into layers, each with a distinct role and task. These layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0120] The application layer may include a series of application packages.
[0121] As shown in Fig. 2, the application package may include applications such as Camera, Gallery, Calendar, Phone, Map, Navigation, WLAN, Bluetooth, Music, Videos, and Messages. An application used for voiceprint recognition may also be included. The application used for voiceprint recognition may be built into the terminal or downloaded from an external website.
[0122] The application framework layer provides an application programming interface (API) and a programming framework for applications in the application layer.
[0123] The application framework layer includes some predetermined functions.
[0124] As shown in FIG. 2, the application framework layer may include a window manager, a content provider, a view system, a telephony manager, a resource manager, a notification manager, and the like.
[0125] A window manager is configured to manage window programs. The window manager may obtain the size of the display, determine whether a status bar is present, perform screen locking, take screenshots, etc.
[0126] Content providers are configured to store and retrieve data and make it accessible by applications. Data may include video, images, audio, calls made and answered, browsing history and bookmarks, address books, etc.
[0127] A view system includes visual controls, such as controls for displaying text and controls for displaying images. The view system can be configured to build applications. A display interface can include one or more views. For example, a display interface that includes an SMS message notification icon can include a view displaying text and a view displaying images.
[0128] The telephone manager is configured to provide communication functions of the mobile phone, such as managing the call status (answer, reject, etc.).
[0129] The resource manager provides various resources to applications, such as localized strings, icons, pictures, layout files, and video files.
[0130] A notification manager may be configured to allow applications to display notification information in the status bar and to communicate notification messages. The notification manager may automatically disappear after a short pause without requiring user interaction. For example, the notification manager may be configured to notify of download completion, provide message notifications, etc. The notification manager may alternatively be a notification that appears in the system's top status bar in the form of a graph or scrollbar text, such as a notification for an application running in the background, or a notification that appears on the screen in the form of a dialog window. For example, text information may be prompted in the status bar, a prompt tone may be generated, the electronic device may vibrate, or an indicator may flash.
[0131] The Android runtime includes the kernel library and virtual machine, and is responsible for scheduling and managing the Android system.
[0132] The kernel library includes two parts: the functions that need to be called in the Java language and the Android kernel library.
[0133] The application layer and the application framework layer run on a virtual machine. The virtual machine executes the java files of the application layer and the application framework layer as binary files. The virtual machine is configured to implement functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage capture.
[0134] The system library may include multiple functional modules, such as a surface manager, media libraries, a 3D graphics processing library (eg, OpenGL ES), and a 2D graphics engine (eg, SGL).
[0135] The surface manager is configured to manage the display subsystem and provide a blend of 2D and 3D layers for multiple applications.
[0136] The media library supports playback and recording of multiple commonly used audio and video formats, static image files, etc. The media library may support multiple audio and video encoding formats, such as MPEG-4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0137] The 3D graphics processing library is configured to implement 3D graphics drawing, image rendering, compositing, layer processing, and the like.
[0138] The 2D graphics engine is a drawing engine for 2D drawing.
[0139] The kernel layer is a layer between the hardware and the software, and includes at least a display driver, a camera driver, an audio driver, and a sensor driver.
[0140] In the following, an example of the operating process of the software and hardware of a mobile phone will be described in relation to a capture and photography scenario.
[0141] When the touch sensor 180K receives a touch operation, a corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into an original input event (including information such as touch coordinates and a timestamp of the touch operation). The original input event is stored in the kernel layer. The application framework layer obtains the original input event from the kernel layer and identifies the control corresponding to the input event. For example, the touch operation is a single-tap touch operation, and the control corresponding to the single-tap operation is the control of a camera application icon. The camera application invokes an interface in the application framework layer, which results in the camera application being opened. Next, the camera driver is launched by invoking the kernel layer, and a still image or video is captured via the camera 193.
[0142] The voice control method in the embodiment of the present application may be applied to a wearable device. In other words, the wearable device may execute the voice control method in the embodiment of the present application. The wearable device may be a device with a voice capture function, such as a wireless headset, a wired headset, smart glasses, a smart helmet, or a smart watch. This is not limited in this embodiment of the present application.
[0143] For example, the wearable device provided in this embodiment of the present application may be a TWS (true wireless stereo) headset, where TWS technology is based on the development of Bluetooth chip technology. Based on the working principle of the wearable device, a mobile phone is connected to a primary headset, and the primary headset is quickly connected to a secondary headset wirelessly. In this way, the left audio channel and the right audio channel are used separately.
[0144] With the development of TWS technology and artificial intelligence technology, TWS smart headsets have begun to play a role in the fields of wireless connection, voice interaction, intelligent noise reduction, health monitoring, hearing enhancement / protection, etc. Noise reduction, hearing protection, intelligent translation, health monitoring, bone vibration ID, and anti-loss are the main technological trends of TWS headsets.
[0145] FIG. 3 is a diagram of the structure of a wearable device. The wearable device 30 may specifically include an in-ear sound sensor 301, an extra-ear sound sensor 302, and a bone vibration sensor 303. The in-ear sound sensor 301 and the extra-ear sound sensor may each be an air conduction microphone, and the bone vibration sensor may be a sensor capable of capturing vibration signals generated when a user speaks, such as a bone conduction microphone, an optical vibration sensor, an acceleration sensor, or an air conduction microphone. An air conduction microphone captures sound information by transmitting vibration signals generated when a user speaks through the air to the microphone, which captures and converts the sound signals into electrical signals. A bone conduction microphone captures sound information by transmitting vibration signals generated when a user speaks through bones and slight vibrations of the bones in the head and neck that occur when a person speaks to the microphone, which captures and converts the sound signals into electrical signals.
[0146] It can be understood that the voice control method provided in the embodiment of the present application needs to be applied to a wearable device with a voiceprint recognition function, in other words, the wearable device 30 needs to have a voiceprint recognition function.
[0147] The in-ear sound sensor 301 of the wearable device 30 provided in this embodiment of the present application means that when the wearable device is in use by a user, the in-ear sound sensor is located inside the user's ear canal, or the sound detection direction of the in-ear sound sensor is inside the ear canal. The in-ear sound sensor is configured to capture sound transmitted through vibrations of the outside air and the air in the ear canal when the user speaks, and this sound is the in-ear sound signal component. The out-of-ear sound sensor 302 means that when the wearable device is in use by a user, the out-of-ear sound sensor is located outside the user's ear canal, or the sound detection direction of the out-of-ear sound sensor is in a direction other than the inside of the ear canal, i.e., the all-outside air direction. The out-of-ear sound sensor is exposed to the environment and configured to capture sound emitted by the user and transmitted through vibrations of the outside air. This sound is the out-of-ear sound signal component or the ambient sound component. The bone vibration sensor 303 means that when the wearable device is in use by a user, the bone vibration sensor 303 is configured to contact the user's skin and capture vibration signals transmitted through the user's bones, or configured to capture audio information components transmitted through bone vibrations when the user speaks at a specific time. Optionally, for both the in-ear microphone and the extra-ear microphone, microphones with different directionality, such as a heart-shaped microphone, an omnidirectional microphone, or an 8-type microphone, may be selected based on the microphone position to acquire audio signals from different directions.
[0148] After a user puts on a headset, the external auditory canal and the middle ear canal form a closed cavity, which has a special amplification effect on the sound inside the cavity, i.e., a cavity effect, so that the sound captured by the in-ear sound sensor is clearer, and there is a significant enhancement effect, especially for high-frequency acoustic signals, which can compensate for the distortion caused when some high-frequency signal components of the sound information are lost when the bone vibration sensor captures the sound information, thereby improving the overall voiceprint capture effect and voiceprint recognition accuracy of the headset and improving the user experience.
[0149] It can be seen that when the in-ear sound sensor 301 picks up an in-ear sound signal, there is usually in-ear residual noise, and when the extra-ear sound sensor 302 picks up an extra-ear sound signal, there is usually extra-ear noise.
[0150] In this embodiment of the present application, when a user wears the wearable device 30 and speaks, the wearable device 30 can not only capture audio information transmitted from the user and transmitted through the air by using the in-ear audio sensor 301 and the out-of-ear audio sensor 302, but also capture audio information transmitted by the user and transmitted through the bones by using the bone vibration sensor 303.
[0151] It can be understood that multiple in-ear sound sensors 301, extra-ear sound sensors 302 and bone vibration sensors 303 may be present in the wearable device 30. This is not limited in the present application. The in-ear sound sensors 301, extra-ear sound sensors 302 and bone vibration sensors 303 may be built into the wearable device 30.
[0152] As further shown in FIG. 3, the wearable device 30 may further include components such as a communication module 304, a speaker 305, a calculation module 306, a storage module 307, and a power source 309.
[0153] When the terminal or server executes the voice control method in the embodiment of the present application, the communication module 304 can establish a communication connection to the terminal or server. The communication module 304 may include a communication interface. The communication interface may be wired or wireless, and the wireless interface may be Bluetooth or Wi-Fi. The communication module 304 may be configured to transfer the first sound component captured by the wearable device 30 using the in-ear sound sensor 301, the second sound component captured using the out-of-ear sound sensor 302, and the third sound component captured using the bone vibration sensor 303 to the terminal or server.
[0154] When the wearable device 30 executes the voice control method in the embodiment of the present application, the calculation module 306 can execute the voice control method provided in the embodiment of the present application. When a user inputs voice information into the wearable device, the wearable device 30 captures a first voice component by using the in-ear voice sensor 301, captures a second voice component by using the out-of-ear voice sensor 302, and captures a third voice component by using the bone vibration sensor 303, performs voiceprint recognition on each of the first, second, and third voice components, and authenticates the user based on the first voiceprint recognition result of the first voice component, the second voiceprint recognition result of the second voice component, and the third voiceprint recognition result of the third voice component. If the user is a preset user as a result of the user authentication, the wearable device executes an operation instruction corresponding to the voice information.
[0155] The storage module 307 is configured to store application code for executing the method in this embodiment of the present application, and the calculation module 306 controls the execution.
[0156] The code stored in the storage module 307 may be used to execute a voice control method provided in an embodiment of the present application. For example, when a user inputs voice information into the wearable device, the wearable device 30 captures a first voice component using the in-ear voice sensor 301, captures a second voice component using the out-of-ear voice sensor 302, captures a third voice component using the bone vibration sensor 303, performs voiceprint recognition on each of the first, second, and third voice components, and authenticates the user based on the first voiceprint recognition result of the first voice component, the second voiceprint recognition result of the second voice component, and the third voiceprint recognition result of the third voice component. If the user is a preset user as a result of the user authentication, the wearable device executes an operation instruction corresponding to the voice information.
[0157] It can be understood that microphones and bone vibration sensors may be randomly combined. Wearable device 30 may further include pressure sensors, acceleration sensors, optical sensors, etc. Wearable device 30 may have more or fewer components than those shown in FIG. 3, may combine two or more components, or may have a different configuration of components. The various components shown in FIG. 3 may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing or application specific integrated circuits.
[0158] The voice control method provided in the embodiment of the present application may be applied to a voice control system including a wearable device 30 and a terminal 10. The voice control system is shown in FIG. 4. In the voice control system, when a user inputs voice information into the wearable device, the wearable device 30 may individually capture a first voice component by using an in-ear voice sensor 301, a second voice component by using an extra-ear voice sensor 302, and a third voice component by using a bone vibration sensor 303. The terminal 10 acquires the first, second, and third voice components from the wearable device, performs voiceprint recognition on the first, second, and third voice components, and authenticates the user based on the first voiceprint recognition result of the first voice component, the second voiceprint recognition result of the second voice component, and the third voiceprint recognition result of the third voice component. If the user authentication result indicates that the user is a preset user, the terminal 10 executes an operation instruction corresponding to the voice information.
[0159] The voice control method in the embodiment of the present application may further be applied to a server, in other words, the server may execute the voice control method in the embodiment of the present application.
[0160] The server may be a desktop server, a rack server, a cabinet server, a blade server, or another type of server, or the server may be a cloud server, such as a public cloud or a private cloud, which is not limited in this embodiment of the present application.
[0161] 5 is a diagram showing the configuration of a server 50. The server 50 includes at least one processor 501, at least one memory 502, and at least one communication interface 503. The processor 501, the memory 502, and the communication interface 503 are connected via a communication bus 504 and communicate with each other.
[0162] The processor 501 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the aforementioned solution programs.
[0163] The memory 502 may be a read-only memory (ROM) or another type of static storage device capable of storing static information and instructions, a random access memory (RAM) or another type of dynamic storage device capable of storing information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other compact disc storage, an optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, and Blu-ray discs, etc.), a disc storage medium or other disc storage device, or any other medium that can be used to carry or store program code expected in the form of instructions or data structures and that can be accessed by a computer. However, the memory 502 is not limited thereto. The memory may exist independently and be connected to the processor via a bus. Alternatively, the memory may be integral to the processor.
[0164] The memory 502 is configured to store application code for executing the methods in the embodiments of the present application, and the processor 501 controls the execution.
[0165] The code stored in the memory 502 may be used to execute a voice control method provided in an embodiment of the present application. For example, when a user inputs voice information into the wearable device, the wearable device captures a first voice component by using an in-ear voice sensor, a second voice component by using an extra-ear voice sensor, and a third voice component by using a bone vibration sensor. The server acquires the first, second, and third voice components from the wearable device via a communication connection, performs voiceprint recognition on each of the first, second, and third voice components, and authenticates the user based on the first voiceprint recognition result of the first voice component, the second voiceprint recognition result of the second voice component, and the third voiceprint recognition result of the third voice component. If the user is identified as a preset user as a result of the user authentication, the server executes an operation instruction corresponding to the voice information.
[0166] The communication interface 503 is configured to communicate with another device or communication network, such as an Ethernet, a Radio Access Network (RAN), or a Wireless Local Area Network (WLAN).
[0167] 1 to 5, a specific implementation of applying the voice control method in the present application to a terminal will be described using an example in which the wearable device is a Bluetooth headset and the terminal is a mobile phone. In this method, user voice information is first acquired. The voice information includes a first voice component, a second voice component, and a third voice component. In this embodiment of the present application, the user may input voice information into the Bluetooth headset while wearing the Bluetooth headset. In this case, the Bluetooth headset may capture the first voice component by using an in-ear voice sensor, the second voice component by using an out-of-ear voice sensor, and the third voice component by using a bone vibration sensor based on the voice information input by the user.
[0168] The Bluetooth headset acquires a first audio component, a second audio component, and a third audio component from the audio information. The mobile phone acquires the first audio component, the second audio component, and the third audio component from the Bluetooth headset via a Bluetooth connection to the Bluetooth headset. In a possible implementation, the mobile phone may perform keyword detection on audio information input by a user to the Bluetooth headset, or the mobile phone may detect user input. Optionally, when the audio information includes a preset keyword, voiceprint recognition is performed on each of the first audio component, the second audio component, and the third audio component. Upon receiving a preset operation input by the user, voiceprint recognition is performed on each of the first audio component, the second audio component, and the third audio component. The user input may be an input performed by the user on the mobile phone by using a touchscreen or a button. For example, the user taps the unlock button on the mobile phone. Optionally, before performing keyword detection on the audio information or detecting a user input, the mobile phone may further acquire a wearing state detection result from the Bluetooth headset. Optionally, if the wearing state detection result is successful, the mobile phone performs keyword detection on the audio information or detects a user input.
[0169] After performing voiceprint recognition on each of the first voice component, the second voice component, and the third voice component, the mobile phone obtains a first voiceprint recognition result corresponding to the first voice component, a second voiceprint recognition result corresponding to the second voice component, and a third voiceprint recognition result corresponding to the third voice component.
[0170] When the first voiceprint feature matches the first registered voiceprint feature, the second voiceprint feature matches the second registered voiceprint feature, and the third voiceprint feature matches the third registered voiceprint feature, this indicates that the voice information captured by the Bluetooth headset in this case was input by a preset user. For example, the mobile phone may calculate a first matching degree between the first voiceprint feature and the first registered voiceprint feature, a second matching degree between the second voiceprint feature and the second registered voiceprint feature, and a third matching degree between the third voiceprint feature and the third registered voiceprint feature based on a specific algorithm. A higher matching degree indicates that the voiceprint feature matches the corresponding registered voiceprint feature better, indicating a higher probability that the speaking user is a preset user. For example, when the average value of the first matching degree, the second matching degree, and the third matching degree is greater than 80 points, the mobile phone may determine that the first voiceprint feature matches the first registered voiceprint feature, the second voiceprint feature matches the second registered voiceprint feature, and the third voiceprint feature matches the third registered voiceprint feature. Alternatively, when the first matching degree, the second matching degree, and the third matching degree are each greater than 85 points, the mobile phone may determine that the first voiceprint feature matches the first enrolled voiceprint feature, the second voiceprint feature matches the second enrolled voiceprint feature, and the third voiceprint feature matches the third enrolled voiceprint feature. The first enrolled voiceprint feature is obtained by performing feature extraction using a first voiceprint model, and the first enrolled voiceprint feature represents a preset user's voiceprint feature captured by an in-ear sound sensor. The second enrolled voiceprint feature is obtained by performing feature extraction using a second voiceprint model, and the second enrolled voiceprint feature represents a preset user's voiceprint feature captured by an out-of-ear sound sensor. The third enrolled voiceprint feature is obtained by performing feature extraction using a third voiceprint model, and the third enrolled voiceprint feature represents a preset user's voiceprint feature captured by a bone vibration sensor.
[0171] It may be understood that the algorithm type and judgment conditions are not limited herein as long as the technical effect of this embodiment of the present application can be achieved. Furthermore, the mobile phone may execute operation instructions corresponding to the voice information, such as an unlock instruction, a payment instruction, a power-off instruction, an application launch instruction, and a call instruction. In this way, the mobile phone can execute corresponding operations based on the operation instructions, so that the user can control the mobile phone by using voice. It may also be understood that the conditions for identity authentication are not limited. For example, if the first matching degree, the second matching degree, and the third matching degree are each greater than a predetermined threshold, identity authentication may be successful and the sound-making user may be considered to be a preset user. If the first matching degree, the second matching degree, and the third matching degree are each greater than a specific threshold, identity authentication may be successful and the sound-making user may be considered to be a preset user. Alternatively, if the fusion matching degree obtained by performing a matching degree fusion on the first matching degree, the second matching degree, and the third matching degree in a specific manner is greater than a predetermined threshold, identity authentication may be successful and the sound-making user may be considered to be a preset user. In this embodiment of the present application, identity authentication refers to obtaining a user's identification information and determining whether the user's identification information matches the preset identification information. If the identification information matches the preset identification information, the authentication is deemed successful; otherwise, if the identification information does not match the preset identification information, the authentication is deemed unsuccessful.
[0172] A preset user is a user who can pass the identity authentication means preset by a mobile phone. For example, when the identity authentication means preset by a terminal are password entry, fingerprint recognition, and voiceprint recognition, a user who successfully enters a password, or a user whose fingerprint information and registered voiceprint characteristics have been pre-stored in the terminal and who has successfully authenticated the user, may be considered a preset user of the terminal. Of course, a terminal may have one or more preset users, and any user other than the preset user may be considered an authorized user of the terminal. After passing a specific identity authentication means, an unauthorized user may be changed to a preset user. This is not limited to this embodiment of the present application.
[0173] In possible embodiments, the first enrollment voiceprint features are obtained by performing feature extraction using a first voiceprint model, the first enrollment voiceprint features representing voiceprint features of a preset user captured by an in-ear sound sensor; the second enrollment voiceprint features are obtained by performing feature extraction using a second voiceprint model, the second enrollment voiceprint features representing voiceprint features of a preset user captured by an out-of-ear sound sensor; and the third enrollment voiceprint features are obtained by performing feature extraction using a third voiceprint model, the third enrollment voiceprint features representing voiceprint features of a preset user captured by a bone vibration sensor.
[0174] In a possible implementation, the algorithm for calculating the degree of matching may be to calculate similarities, wherein the mobile phone performs feature extraction on the first voice component to obtain a first voiceprint feature, separately calculates a first similarity between the first voiceprint feature and a first registered voiceprint feature of a preset user that has been pre-stored, a second similarity between the second voiceprint feature and a second registered voiceprint feature of a preset user that has been pre-stored, and a third similarity between the third voiceprint feature and a third registered voiceprint feature of a preset user that has been pre-stored, and performs user authentication based on the first similarity, the second similarity, and the third similarity.
[0175] In a possible implementation, a method for performing identity authentication on a user may be as follows: the mobile phone separately determines a first fusion coefficient corresponding to a first similarity, a second fusion coefficient corresponding to a second similarity, and a third fusion coefficient corresponding to a third similarity based on the decibels of the ambient sound and the playback volume of the wearable device, and fuses the first similarity, the second similarity, and the third similarity based on the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient to obtain a fused similarity score. If the fused similarity score is greater than a first threshold, the mobile phone determines that the user who inputs voice information into the Bluetooth headset is a preset user.
[0176] In a possible implementation, the decibels of the ambient sound may be detected by a sound pressure sensor in the Bluetooth headset and sent to the mobile phone, and the playback volume may be obtained by detecting the playback signal by the speaker of the Bluetooth headset and sent to the mobile phone, or may be obtained by the mobile phone by calling data on the mobile phone, i.e., by using the volume interface program interface of the underlying system.
[0177] In a possible implementation, the second fusion coefficient is negatively correlated with the decibels of the ambient sound, the first fusion coefficient and the third fusion coefficient are each negatively correlated with the decibels of the playback volume, and the sum of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient is a fixed value. Specifically, when the sum of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient is a preset fixed value, a larger decibel of the ambient sound indicates a smaller second fusion coefficient. In this case, the first fusion coefficient and the third fusion coefficient are adaptively increased, while the sum of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient remains unchanged. A larger playback volume indicates a smaller first fusion coefficient and a smaller third fusion coefficient. In this case, the second fusion coefficient is adaptively increased, while the sum of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient remains unchanged. It can be seen that based on the variable fusion coefficient, recognition accuracy can be considered in different application scenarios (such as when a noisy environment is high or when music is played using a headset).
[0178] After the mobile phone determines that the user who inputs the voice information into the Bluetooth headset is a preset user, the mobile phone may automatically execute an operation instruction corresponding to the voice information, such as a mobile phone unlocking operation or a payment confirmation operation.
[0179] In this embodiment of the present application, when a user inputs voice information into a wearable device to control a terminal, the wearable device may acquire voice information generated within the ear canal when the user speaks, and voice information and bone vibration information generated outside the ear canal. In this case, three-channel voice information (i.e., a first voice component, a second voice component, and a third voice component) is generated in the wearable device. In this manner, the terminal (or the wearable device or the server) may perform voiceprint recognition on each of the three channels of voice information. When the voiceprint recognition results of the three channels of voice information all match the registered voiceprint features of a preset user, it may be determined that the user who input the voice information is the current preset user. Alternatively, when the fusion result obtained after performing weighted fusion on the voiceprint recognition results of the three channels of voice information is greater than a predetermined threshold, it may be determined that the user who input the voice information is the current preset user. It is clear that, compared with the voiceprint recognition process of one channel of voice information or the voiceprint recognition process of two channels of voice information, the triple voiceprint recognition process of three channels of voice information can significantly improve the accuracy and security of user identity authentication. In particular, by adding one microphone to each ear, the voiceprint recognition process for two-channel audio information from the extra-aural audio sensor and the bone vibration sensor can solve the problem of high-frequency signals being lost in the audio signal captured by the bone vibration sensor.
[0180] In addition, the wearable device can capture voice information input by the user through bone conduction only after the user wears the wearable device. Therefore, when voiceprint recognition performed on the voice information captured by the wearable device through bone conduction is successful, this indicates that the voice information was generated when the preset user wearing the wearable device speaks, thereby preventing a case where an unauthorized user maliciously controls the preset user's terminal based on the preset user's recording.
[0181] For ease of understanding, the following will specifically describe the voice control method provided in the embodiments of the present application with reference to the accompanying drawings. In the following embodiments, the description will be provided by using an example in which a mobile phone functions as a terminal and a Bluetooth headset functions as a wearable device.
[0182] First, a brief description of voiceprint recognition technology will be given.
[0183] In practical applications, voiceprint recognition technology typically includes two steps: an enrollment step and a verification step. A typical voiceprint recognition application procedure is shown in FIG. 6. In the enrollment step, an enrollment voice 601 is captured, preprocessed by a preprocessing module 602, and input to a pre-trained voiceprint model 603 for feature extraction to obtain enrollment voiceprint features 604. The enrollment voiceprint features may also be understood as preset user enrollment voiceprint features. It may be understood that the enrollment voice may be extracted by different types of sensors, such as an extra-aural voice sensor, an in-aural voice sensor, or a bone vibration sensor. The voiceprint model 603 is acquired in advance through training performed based on training data. The voiceprint model 603 may be built in the terminal before the terminal is shipped from the factory, or may be trained by an application that instructs the user. The training method may be a method in the prior art, but this is not a limitation in this application. In the verification procedure, a test voice 605 of the speaking user is captured in the voiceprint recognition process, preprocessed by a preprocessing module 606, and input into a pre-trained voiceprint model 607 for feature extraction to obtain test voice voiceprint features 608. The test voiceprint features may also be understood as enrollment voiceprint features of a preset user. After performing voiceprint recognition 609 by performing voiceprint recognition based on the enrollment voiceprint features 604 and the test voice voiceprint features 608, voiceprint recognition results are obtained, namely, successful authentication 6010 and unsuccessful authentication 6011. The successful authentication 6010 means that the speaking user of the test voice 605 and the speaking user of the enrollment voice 601 are the same person. In other words, the speaking user of the test voice 605 is the preset user. The unsuccessful authentication 6011 means that the speaking user of the test voice 605 and the speaking user of the enrollment voice 601 are not the same person. In other words, the user who speaks the test voice 605 is an unauthorized user. It can be understood that in different application scenarios, the processes of sound pre-processing, feature extraction, and voiceprint model training vary to different degrees. In addition, the pre-processing module is an optional module, and the pre-processing includes filtering, noise reduction, or enhancement of the voice signal. This is not limited in the present application.
[0184] 7 is a schematic flowchart of a voice control method according to an embodiment of the present application, by using an example in which the terminal is a mobile phone and the wearable device is a Bluetooth headset. The Bluetooth headset includes an in-ear sound sensor, an out-of-ear sound sensor and a bone vibration sensor. As shown in FIG. 7, the voice control method may include the following steps:
[0185] S701: A mobile phone establishes a connection with a Bluetooth headset.
[0186] The connection method may be a Bluetooth connection, a Wi-Fi connection, or a wired connection. When a mobile phone establishes a Bluetooth connection to a Bluetooth headset, the user may enable the Bluetooth function of the Bluetooth headset when the user expects to use the Bluetooth headset. In this case, the Bluetooth headset may transmit a pairing broadcast to the outside. If the Bluetooth function of the mobile phone is not enabled, the user must enable the Bluetooth function of the mobile phone. If the Bluetooth function of the mobile phone is enabled, the mobile phone may receive the pairing broadcast and prompt the user when a related Bluetooth device is discovered through scanning. After the user selects a Bluetooth headset on the mobile phone, the mobile phone may pair with the Bluetooth headset and establish a Bluetooth connection. Thereafter, the mobile phone and the Bluetooth headset may communicate with each other through the Bluetooth connection. Of course, if the mobile phone was successfully paired with the Bluetooth headset before the current Bluetooth connection was established, the mobile phone may automatically establish a Bluetooth connection to the Bluetooth headset discovered through scanning.
[0187] In addition, if the user expects the headset being used to have a Wi-Fi function, the user may operate the mobile phone to establish a Wi-Fi connection to the headset. Alternatively, if the user expects the headset being used to be a wired headset, the user may insert the headset cable plug into the corresponding headset jack of the mobile phone to establish a wired connection. This is not limited to this embodiment of the present application.
[0188] S702 (optional): The Bluetooth headset detects whether the Bluetooth headset is in a worn state.
[0189] In the wearing detection method, the wearing state of the user may be detected by a photoelectric detection method based on the principle of light detection: when the user wears the headset, the light detected by the photoelectric sensor inside the headset is blocked, and a switch control signal is output, determining that the user is wearing the headset.
[0190] Specifically, an optical proximity sensor and an acceleration sensor may be disposed in the Bluetooth headset, the optical proximity sensor being disposed on the side that contacts the user when the user wears the Bluetooth headset, and the optical proximity sensor and the acceleration sensor may be periodically enabled to obtain currently detected measurements.
[0191] After a user puts on the Bluetooth headset, the light emitted to the optical proximity sensor is blocked. Therefore, when the light intensity detected by the optical proximity sensor is less than a preset light intensity threshold, the Bluetooth headset may determine that the Bluetooth headset is currently in a worn state. In addition, after a user puts on the Bluetooth headset, the Bluetooth headset may move with the user. Therefore, when the acceleration value detected by the acceleration sensor is greater than the preset acceleration threshold, the Bluetooth headset may determine that the Bluetooth headset is currently in a worn state. Alternatively, when the light intensity detected by the optical proximity sensor is less than the preset light intensity threshold, if it is detected that the acceleration value currently detected by the acceleration sensor is greater than the preset acceleration threshold, the Bluetooth headset may determine that the Bluetooth headset is currently in a worn state.
[0192] Furthermore, since a sensor that captures audio information through bone conduction, such as a bone vibration sensor or an optical vibration sensor, is further disposed within the Bluetooth headset, in a possible implementation, the Bluetooth headset may further capture vibration signals generated in the current environment by using the bone vibration sensor. When the Bluetooth headset is in a worn state, the Bluetooth headset is in direct contact with the user. Therefore, the vibration signal captured by the bone vibration sensor is stronger than the vibration signal captured in an unaware state. In this case, if the energy of the vibration signal captured by the bone vibration sensor is greater than an energy threshold, the Bluetooth headset may determine that the Bluetooth headset is in a worn state. Alternatively, since the spectral features, such as harmonics and resonance, of the vibration signal captured when the user is wearing the Bluetooth headset are significantly different from the vibration signal captured when the user is not wearing the Bluetooth headset, the Bluetooth headset may determine that the Bluetooth headset is in a worn state if the vibration signal captured by the bone vibration sensor meets preset spectral features. In both of these cases, the detection result of the user's wearing state can be understood as being successful. This allows the use of an optical proximity sensor or an acceleration sensor to reduce the probability that the Bluetooth headset will not be able to accurately detect the wearing state in a scenario where the user puts the Bluetooth headset in a pocket, for example.
[0193] The energy threshold or preset spectral characteristics are obtained through statistical capture after various vibration signals generated through speaking, movement, etc. are captured after a large number of users wear Bluetooth headsets, and are clearly different from the energy or spectral characteristics of the audio signal detected by the bone vibration sensor when the user is not wearing the Bluetooth headset. In addition, because the power consumption of an audio sensor (e.g., an air conduction microphone) external to the Bluetooth headset is usually high, it is not necessary to enable the in-ear audio sensor, the extra-ear audio sensor, and / or the bone vibration sensor before the Bluetooth headset detects that the Bluetooth headset is currently in a worn state. In order to reduce the power consumption of the Bluetooth headset, after detecting that the Bluetooth headset is currently in a worn state, the Bluetooth headset may enable the in-ear audio sensor, the extra-ear audio sensor, and / or the bone vibration sensor to capture audio information generated when the user speaks.
[0194] After the Bluetooth headset detects that the Bluetooth headset is currently in a worn state or after the wearing state detection result is successful, steps S703 to S707 may be continuously executed. Before the Bluetooth headset detects that the Bluetooth headset is currently in a worn state or before the wearing state detection result is successful, the Bluetooth headset may enter a sleep state, and steps S703 to S707 may be continuously executed until the Bluetooth headset detects that the Bluetooth headset is currently in a worn state. In other words, only when it is detected that the user is wearing the Bluetooth headset, that is, only when it is detected that the user intends to use the Bluetooth headset, the Bluetooth headset can trigger a process of performing a capture to obtain voice information input by the user, a voiceprint recognition process, etc., to reduce the power consumption of the Bluetooth headset. Of course, step S702 is optional. Specifically, the Bluetooth headset may continue to execute steps S703 to S707 regardless of whether the user is wearing the Bluetooth headset. This is not limited to this embodiment of the present application.
[0195] In a possible implementation, if the Bluetooth headset captures an audio signal before detecting whether the Bluetooth headset is in a wearing state, after detecting that the Bluetooth headset is currently in a wearing state or after the wearing state detection result is passed, the audio signal captured by the Bluetooth headset is stored, and steps S703 to S707 are continued to be executed, or when the Bluetooth headset does not detect that the Bluetooth headset is currently in a wearing state or after the wearing state detection result is passed, the Bluetooth headset deletes the captured audio signal.
[0196] S703: When the Bluetooth headset is in a worn state, the Bluetooth headset performs capture by using an in-ear sound sensor to obtain a first sound component of the sound information input by the user, captures a second sound component of the sound information by using an out-of-ear sound sensor, and captures a third sound component of the sound information by using a bone vibration sensor.
[0197] When the Bluetooth headset is determined to be in a worn state, the Bluetooth headset activates the voice detection module and performs capture by using the in-ear voice sensor, the extra-ear voice sensor, and the bone vibration sensor to acquire voice information input by the user and obtain the first, second, and third voice components of the voice information. For example, the in-ear voice sensor and the extra-ear voice sensor are air conduction microphones, and the bone vibration sensor is a bone conduction microphone. In the process of using the Bluetooth headset, the user may input voice information such as "Hey Celia, use WeChat Pay." In this case, because the air conduction microphone is exposed to air, the Bluetooth headset can use the air conduction microphone to receive vibration signals generated through air vibrations after the user speaks (i.e., the first, second, and third voice components of the voice information). In addition, because the bone conduction microphone can contact the user's ear bones through the skin, the Bluetooth headset can use the bone conduction microphone to receive vibration signals generated through vibrations of the ear bones and skin after the user speaks (i.e., the third voice component of the voice information).
[0198] 8 is a schematic diagram of a sensor placement area. The Bluetooth headset provided in this embodiment of the present application includes an in-ear sound sensor, an extra-ear sound sensor, and a bone vibration sensor. The in-ear sound sensor means that when the headset is in use by a user, the in-ear sound sensor is located inside the user's ear canal, or the sound detection direction of the in-ear sound sensor is inside the ear canal, and the in-ear sound sensor is placed in the in-ear sound sensor placement area 801. The in-ear sound sensor is configured to capture sound transmitted by vibrations of the outside air and the air in the ear canal when the user speaks, and the sound is an in-ear sound signal component. The extra-ear sound sensor means that when the headset is in use by a user, the extra-ear sound sensor is located outside the user's ear canal, or the sound detection direction of the extra-ear sound sensor is in a direction other than the inside of the ear canal, i.e., the all-outside air direction, and the extra-ear sound sensor is placed in the extra-ear sound sensor placement area 802. The extra-ear sound sensor is exposed to the environment and configured to capture sounds emitted by the user and transmitted by vibrations in the external air. The sounds are extra-ear sound signal components or ambient sound components. The bone vibration sensor means that when the headset is in use by the user, the bone vibration sensor is in contact with the user's skin and configured to capture vibration signals transmitted through the user's bones, or configured to capture sound information components transmitted through bone vibrations when the user speaks at a specific time. The placement area of the bone vibration sensor is not limited as long as it can detect the user's bone vibrations when the user is wearing the headset. It can be understood that the in-ear sound sensor may be placed at any position within area 801, and the extra-ear sound sensor may be placed at any position within area 802. This is not limited in the present application. Note that the area division method in FIG. 8 is merely an example, and the area division method may be any method as long as it can detect sounds inside the ear canal at the position where the in-ear sound sensor is placed and can detect sounds in the external air direction at the position where the extra-ear sound sensor is placed.
[0199] In some embodiments of the present application, after detecting the voice information input by the user, the Bluetooth headset may further distinguish between a voice signal and background noise in the voice information based on a voice activity detection (VAD) algorithm. In particular, the Bluetooth headset may input each of a first voice component, a second voice component, and a third voice component of the voice information into a corresponding VAD algorithm to obtain a first VAD value corresponding to the first voice component, a second VAD value corresponding to the second voice component, and a third VAD value corresponding to the third voice component. The VAD values may indicate whether the voice information is a normal voice signal of a speaker or a noise signal. For example, the VAD value may be set to a range of 0 to 100. When the VAD value is greater than a certain VAD threshold, this may indicate that the voice information is a normal voice signal of a speaker, or when the VAD value is less than the certain VAD threshold, this may indicate that the voice information is a noise signal. In another example, the VAD value may be set to 0 or 1. A VAD value of 1 indicates that the voice information is a normal voice signal of a speaker, and a VAD value of 0 indicates that the voice information is a noise signal.
[0200] In this case, the Bluetooth headset may determine whether the audio information is a noise signal based on three VAD values, i.e., the first VAD value, the second VAD value, and the third VAD value. For example, when the first VAD value, the second VAD value, and the third VAD value are each 1, the Bluetooth headset may determine that the audio information is not a noise signal but is a normal voice signal of the speaker. In another example, when the first VAD value, the second VAD value, and the third VAD value are each greater than a preset value, the Bluetooth headset may determine that the audio information is not a noise signal but is a normal voice signal of the speaker.
[0201] In addition, when the third VAD value is 1 or the third VAD value is greater than the preset value, this may indicate to some extent that the currently captured voice information is transmitted by a live user. Therefore, the Bluetooth headset may alternatively determine whether the voice information is a noise signal based only on the third VAD value. It can be understood that in some cases, the Bluetooth headset may alternatively determine whether the voice information is a noise signal based only on the first VAD value or the second VAD value, or the Bluetooth headset may alternatively determine whether the voice information is a noise signal based on any two of the first VAD value, the second VAD value, and the third VAD value.
[0202] Voice activity detection is performed for each of the first voice component, the second voice component, and the third voice component. If the Bluetooth headset determines that the voice information is a noise signal, the Bluetooth headset may discard the voice information. If the Bluetooth headset determines that the voice information is not a noise signal, the Bluetooth headset may continue to perform steps S704 to S707. In other words, only when the user inputs valid voice information into the Bluetooth headset is the Bluetooth headset triggered to perform subsequent processing, such as voiceprint recognition, to reduce the power consumption of the Bluetooth headset.
[0203] In addition, after obtaining the first VAD value, the second VAD value, and the third VAD value corresponding to the first audio component, the second audio component, and the third audio component, respectively, the Bluetooth headset may further calculate each noise value of the audio information based on a noise estimation algorithm (e.g., a minimum statistics algorithm or a minimum-controlled recursive average algorithm). For example, the Bluetooth headset may be provided with a storage space specially used for storing the noise values, and the Bluetooth headset may update the new noise value in the storage space every time after calculating a new noise value. In other words, the latest calculated noise value is always stored in the storage space.
[0204] In this way, after determining that the voice information is valid voice information based on the VAD algorithm, the Bluetooth headset may perform noise reduction processing on each of the first voice component, the second voice component, and the third voice component based on the noise value in the storage space, so that the recognition results obtained when the Bluetooth headset subsequently performs voiceprint recognition on each of the first voice component, the second voice component, and the third voice component are more accurate.
[0205] S704: The Bluetooth headset transmits the first audio component, the second audio component, and the third audio component to the mobile phone via the Bluetooth connection.
[0206] After obtaining the first voice component, the second voice component, and the third voice component, the Bluetooth headset may send the first voice component, the second voice component, and the third voice component to the mobile phone, so that the mobile phone executes steps S705 to S707 to perform voiceprint recognition, user identity authentication, etc. on the voice information input by the user.
[0207] S705: The mobile phone performs voiceprint recognition on each of the first voice component, the second voice component, and the third voice component, and obtains a first voiceprint recognition result corresponding to the first voice component, a second voiceprint recognition result corresponding to the second voice component, and a third voiceprint recognition result corresponding to the third voice component.
[0208] The principle of voiceprint recognition is to compare the registered voiceprint features of the preset user with the voiceprint features extracted from the voice information input by the user, and make a judgment based on a specific algorithm. The judgment result is the voiceprint recognition result.
[0209] Specifically, the mobile phone may pre-store enrollment voiceprint features of one or more preset users, each of which has three enrollment voiceprint features: a first enrollment voiceprint feature of the user that is obtained by performing feature extraction on the first enrollment voice captured when the in-ear sound sensor is operating, a second enrollment voiceprint feature of the user that is obtained by performing feature extraction on the second enrollment voice captured when the out-of-ear sound sensor is operating, and a third enrollment voiceprint feature of the user that is obtained by performing feature extraction on the third enrollment voice captured when the bone conduction microphone is operating.
[0210] The first, second, and third enrollment voiceprint features need to be obtained in two stages. The first stage is a background model training stage. In the first stage, a developer may capture the speech of related text (e.g., "Hey Celia") produced when multiple speakers wearing Bluetooth headsets utter. The mobile phone may then perform preprocessing (e.g., filtering and noise reduction) on the speech of the related text to extract voiceprint features from the speech. The voiceprint features may be spectrogram features, filter bank-based features, mel-frequency cepstral coefficient (MFCC) features, perceptual linear prediction (PLP) features, constant Q cepstral coefficient (CQCC) features, etc. Instead of directly extracting the above voiceprint features, the mobile phone may extract two or more of the above voiceprint features and obtain a fused voiceprint feature through splicing. After the mobile phone extracts the voiceprint features, a background model for voiceprint recognition is established based on machine learning algorithms such as GMM (Gaussian mixed model), SVM (support vector machine) or deep neural network frameworks, including but not limited to DNN (deep neural network) algorithms, RNN (recurrent neural network) algorithms, LSTM (long short-term memory) algorithms, TDNN (time delay neural network) and ResNet (deep residual network).It can be seen that in the above steps, a UBM (universal background model) is constructed by training a large amount of speech. The UBM can be adaptively trained, and the parameters of the UBM can be adjusted based on the requirements of different manufacturers or users.
[0211] After acquiring the background model, the mobile phone stores the acquired background model. It can be understood that the storage location can be the mobile phone, the wearable device, or a server depending on the execution entity of the method. It should be noted that one or more background models can be stored, and the stored background models can be acquired based on the same or different algorithms. The stored background models can implement voiceprint model fusion. For example, a Resnet (i.e., deep residual network) can be used for training to acquire a voiceprint model of a first background speaker, a TDNN (time delay neural network) can be used for training to acquire a voiceprint model of a second background speaker, and an RNN (i.e., recurrent neural network) can be used for training to acquire a voiceprint model of a third background speaker. In this embodiment of the present application, it can be understood that a model can be established for each of the air conduction microphone and the bone vibration microphone, and multiple models can be fused. The mobile phone or Bluetooth headset can separately establish multiple voiceprint models based on the background model and by referring to the characteristics of different voice sensors in the wearable device connected to the mobile phone. For example, a first voiceprint model corresponding to an in-ear sound sensor of the Bluetooth headset, a second voiceprint model corresponding to an out-of-ear sound sensor of the Bluetooth headset, and a third voiceprint model corresponding to a bone vibration sensor of the Bluetooth headset are established. The mobile phone may store the first, second, and third voiceprint models locally within the mobile phone or may transmit the first, second, and third voiceprint models to the Bluetooth headset for storage.
[0212] The second stage is a process in which, when a user uses the voiceprint recognition function of a mobile phone for the first time, the user inputs an enrollment voice, and the mobile phone extracts the user's first, second, and third enrollment voiceprint features by using the in-ear voice sensor, the out-of-ear voice sensor, and the bone vibration sensor of a Bluetooth headset connected to the mobile phone. In this stage, the enrollment process may be performed by using a voiceprint recognition option in the device biometric recognition function built into the mobile phone's system, or by calling a system program using a downloaded app. For example, when preset user 1 uses a voice assistant app installed on the mobile phone for the first time, the voice assistant app may prompt the user to put on the Bluetooth headset and say the enrollment voice, "Hey Celia." Similarly, because the Bluetooth headset includes an in-ear voice sensor, an out-of-ear voice sensor, and a bone vibration sensor, the Bluetooth headset may acquire a first enrollment voice component of the enrollment voice captured using the in-ear voice sensor, a second enrollment voice component captured using the out-of-ear voice sensor, and a third enrollment voice component captured using the bone vibration sensor. Furthermore, after the Bluetooth headset transmits the first, second, and third enrollment voice components to the mobile phone, the mobile phone individually performs feature extraction on the first enrollment voice component by using the first voiceprint model to obtain a first enrollment voiceprint feature, performs feature extraction on the second enrollment voice component by using the second voiceprint model to obtain a second enrollment voiceprint feature, and performs feature extraction on the third enrollment voice component by using the third voiceprint model to obtain a third enrollment voiceprint feature. The mobile phone may locally store the first, second, and third enrollment voiceprint feature of preset user 1, or may transmit the first, second, and third enrollment voiceprint feature of preset user 1 to the Bluetooth headset for storage.
[0213] Optionally, when extracting the first, second, and third registered voiceprint features of preset user 1, the mobile phone may further use the currently connected Bluetooth headset as a preset Bluetooth device. For example, the mobile phone may locally store the identifier of the preset Bluetooth device (e.g., the MAC address of the Bluetooth headset). In this case, the mobile phone may receive and execute associated operation instructions sent by the preset Bluetooth device. If an unauthorized Bluetooth device sends an operation instruction to the mobile phone, the mobile phone may discard the operation instruction to improve security. One mobile phone may manage one or more preset Bluetooth devices. As shown in FIG. 11(a), a user may access a setting interface 1101 of the voiceprint recognition function from the setting function and tap a setting button 1105. After that, the user may access a preset device management interface 1106 shown in FIG. 11(b). The user may add or delete preset Bluetooth devices to or from the preset device management interface 1106.
[0214] In step S705, after obtaining the first, second, and third voice components of the voice information, the mobile phone individually extracts voiceprint features of the first voice component to obtain a first voiceprint feature, extracts voiceprint features of the second voice component to obtain a second voiceprint feature, extracts voiceprint features of the third voice component to obtain a third voiceprint feature, matches the first registered voiceprint feature of preset user 1 with the first voiceprint feature, matches the second registered voiceprint feature of preset user 1 with the second voiceprint feature, and matches the third registered voiceprint feature of preset user 1 with the third voiceprint feature. For example, the mobile phone may calculate, based on a specific algorithm, a first matching degree between the first registered voiceprint feature and the first voice component (i.e., the first voiceprint recognition result), a second matching degree between the second registered voiceprint feature and the second voice component (i.e., the second voiceprint recognition result), and a third matching degree between the third registered voiceprint feature and the third voice component (i.e., the third voiceprint recognition result). Generally, the higher the matching degree, the higher the similarity between the voiceprint features of the voice information and the voiceprint features of preset user 1, indicating a higher probability that the user who input the voice information is preset user 1.
[0215] For example, when the average value of the first matching degree, the second matching degree, and the third matching degree is greater than 80 points, the mobile phone may determine that the first voiceprint feature matches the first registered voiceprint feature, the second voiceprint feature matches the second registered voiceprint feature, and the third voiceprint feature matches the third registered voiceprint feature. Alternatively, when the first matching degree, the second matching degree, and the third matching degree are each greater than 85 points, the mobile phone may determine that the first voiceprint feature matches the first registered voiceprint feature, the second voiceprint feature matches the second registered voiceprint feature, and the third voiceprint feature matches the third registered voiceprint feature.
[0216] The first enrollment voiceprint feature is obtained by performing feature extraction using a first voiceprint model, and the first enrollment voiceprint feature represents a voiceprint feature of a preset user captured by an in-ear sound sensor. The second enrollment voiceprint feature is obtained by performing feature extraction using a second voiceprint model, and the second enrollment voiceprint feature represents a voiceprint feature of a preset user captured by an extra-ear sound sensor. The third enrollment voiceprint feature is obtained by performing feature extraction using a third voiceprint model, and the third enrollment voiceprint feature represents a voiceprint feature of a preset user captured by a bone vibration sensor. It can be understood that the function of the voiceprint model is to extract voiceprint features of an input voice. When the input voice is an enrollment voice, the voiceprint model can extract the enrollment voiceprint feature of the enrollment voice. When the input voice is a voice uttered by a user at a specific time, the voiceprint model can extract the voiceprint feature of the voice. Optionally, the voiceprint feature acquisition method may alternatively be a fusion method, including a voiceprint model fusion method and a voiceprint feature fusion method.
[0217] In a possible implementation, the algorithm for calculating the degree of matching may be to calculate similarities, wherein the mobile phone performs feature extraction on the first voice component to obtain a first voiceprint feature, and separately calculates a first similarity between the first voiceprint feature and a first registered voiceprint feature of a preset user that has been pre-stored, a second similarity between the second voiceprint feature and a second registered voiceprint feature of a preset user that has been pre-stored, and a third similarity between the third voiceprint feature and a third registered voiceprint feature of a preset user that has been pre-stored.
[0218] If the mobile phone stores registered voiceprint features of multiple preset users, the mobile phone may further sequentially calculate a first matching degree between the first voice component and another preset user (e.g., preset user 2 or preset user 3) and a second matching degree between the second voice component and another preset user by the above-mentioned method. Furthermore, the Bluetooth headset may determine the preset user with the highest matching degree (e.g., preset user A) as the currently speaking user.
[0219] In addition, before performing voiceprint recognition on the first, second, and third voice components, the mobile phone may further predetermine whether voiceprint recognition needs to be performed on the first, second, and third voice components. The determination method may be to perform keyword detection on the voice information. When the voice information includes a preset keyword, the mobile phone performs voiceprint recognition on each of the first, second, and third voice components. Alternatively, the determination method may be to detect a user input. Upon receiving a preset operation input by the user, the mobile phone performs voiceprint recognition on each of the first, second, and third voice components. A specific method of keyword detection is that after performing voice recognition on the keyword, if the similarity is greater than a preset threshold, keyword detection is deemed successful.
[0220] In a possible implementation, if the Bluetooth headset or the mobile phone can identify preset keywords from the voice information input by the user, such as keywords related to user privacy or fund behavior, such as "transfer," "payment," "**bank," or "chat record," this indicates high security requirements when the user controls the mobile phone using voice. Therefore, the mobile phone may perform step S705 to perform voiceprint recognition. In another example, if the Bluetooth headset detects a preset operation, such as tapping the Bluetooth headset or simultaneously pressing the volume up and volume down buttons, input by the user and used to enable the voiceprint recognition function, this indicates that the user needs to verify their identity through voiceprint recognition. Therefore, the Bluetooth headset may notify the mobile phone to perform step S705, i.e., to perform voiceprint recognition.
[0221] Alternatively, keywords corresponding to different security levels may be preset in the mobile phone. For example, the highest security level keywords may include "pay" and "payment," while the higher security level keywords may include "photographing" and "calling," and the lowest security level keywords may include "listening to a song" and "navigation." In this way, when it is detected that the captured audio information contains a keyword with the highest security level, the mobile phone may be triggered to perform voiceprint recognition on each of the first audio component, the second audio component, and the third audio component, i.e., on all three captured audio sources, thereby improving the security of controlling the mobile phone by using voice. When it is detected that the captured audio information contains a keyword with a high security level, the mobile phone may be triggered to perform voiceprint recognition only on the first audio component, the second audio component, or the third audio component, because the security requirements imposed when a user controls the mobile phone by using voice are medium. When the captured audio information is detected to contain a keyword with the lowest security level, the mobile phone does not need to perform voiceprint recognition on the first audio component, the second audio component, or the third audio component.
[0222] Of course, if the voice information captured by the Bluetooth headset does not contain the keyword, this indicates that the currently captured voice information is only the voice information transmitted by the user in normal conversation, so the mobile phone does not need to perform voiceprint recognition on the first voice component, the second voice component or the third voice component, thereby reducing the power consumption of the mobile phone.
[0223] Alternatively, the mobile phone may further preset one or more wake-up words to wake up the mobile phone and enable the voiceprint recognition function. For example, the wake-up word may be "Hey Celia." After the user inputs voice information into the Bluetooth headset, the Bluetooth headset or the mobile phone may identify whether the voice information is a wake-up voice including the wake-up word. For example, the Bluetooth headset may transmit the first voice component, the second voice component, and the third voice component of the captured voice information to the mobile phone. If the mobile phone further identifies that the voice information includes the wake-up word, the mobile phone may enable the voiceprint recognition function (e.g., the mobile phone may power on the voiceprint recognition chip). Subsequently, if the voice information captured by the Bluetooth headset includes a keyword, the mobile phone may perform the voiceprint recognition of the method in step S705 by using the enabled voiceprint recognition function.
[0224] In another example, after capturing the audio information, the Bluetooth headset may further identify whether the audio information includes a wake-up word. If the audio information includes a wake-up word, this indicates that the user may need to use the voiceprint recognition function later. In this case, the Bluetooth headset sends an activation instruction to the mobile phone, and the mobile phone then activates the voiceprint recognition function in response to the activation instruction.
[0225] S706: The mobile phone performs user authentication based on the first voiceprint recognition result, the second voiceprint recognition result, and the third voiceprint recognition result.
[0226] In step S706, after obtaining a first voiceprint recognition result corresponding to the first voice component, a second voiceprint recognition result corresponding to the second voice component, and a third voiceprint recognition result corresponding to the third voice component through voiceprint recognition, the mobile phone may combine the three voiceprint recognition results to authenticate the user who input the voice information, thereby improving the accuracy and security of user authentication.
[0227] For example, the first matching degree between the first registered voiceprint feature of the preset user and the first voiceprint feature is the first voiceprint recognition result, the second matching degree between the second registered voiceprint feature of the preset user and the second voiceprint feature is the second voiceprint recognition result, and the third matching degree between the third registered voiceprint feature of the preset user and the third voiceprint feature is the third voiceprint recognition result. During user identity authentication, if the first matching degree, the second matching degree, and the third matching degree satisfy a preset authentication policy, for example, when the first matching degree is greater than a first threshold, the second matching degree is greater than a second threshold, and the third matching degree is greater than a third threshold (the third threshold, the second threshold, and the first threshold may be the same or different), the mobile phone may determine that the user who sent the first voice component, the second voice component, and the third voice component is a preset user; alternatively, if the first matching degree, the second matching degree, or the third matching degree does not satisfy the preset authentication policy, the mobile phone may determine that the user who sent the first voice component, the second voice component, and the third voice component is an unauthorized user.
[0228] In another example, the mobile phone may calculate a weighted average value of the first matching degree and the second matching degree, and when the weighted average value is greater than a preset threshold, the mobile phone may determine that the user who sent the first voice component, the second voice component, and the third voice component is a preset user, or when the weighted average value is not greater than the preset threshold, the mobile phone may determine that the user who sent the first voice component, the second voice component, and the third voice component is an unauthorized user.
[0229] Alternatively, the mobile phone may use different authentication policies in different voiceprint recognition scenarios. For example, when the captured voice information includes a keyword with the highest security level, the mobile phone may set each of the first, second, and third thresholds to 99 points. In this case, the mobile phone determines that the currently speaking user is a preset user only when the first matching degree, the second matching degree, and the third matching degree are all greater than 99 points. When the captured voice information includes a keyword with a low security level, the mobile phone may set each of the first, second, and third thresholds to 85 points. In this case, the mobile phone may determine that the currently speaking user is a preset user only when the first matching degree, the second matching degree, and the third matching degree are all greater than 85 points. In other words, for voiceprint recognition scenarios with different security levels, the mobile phone may perform user identity authentication based on authentication policies for the different security levels.
[0230] In addition, if the mobile phone stores voiceprint models of one or more preset users, for example, if the mobile phone stores enrolled voiceprint features of preset user A, preset user B, and preset user C, the enrolled voiceprint features of each preset user include a first enrolled voiceprint feature, a second enrolled voiceprint feature, and a third enrolled voiceprint feature. In this case, the mobile phone may match the captured first voice component, the second voice component, and the third voice component with the enrolled voiceprint features of each preset user in the above-mentioned method. Furthermore, the mobile phone may determine that the preset user (e.g., preset user A) that satisfies the authentication policy and has the highest degree of matching is the user currently speaking.
[0231] In this way, after receiving the first, second, and third voice components of the voice information transmitted by the Bluetooth headset, the mobile phone can fuse the first, second, and third voice components and then perform voiceprint recognition. For example, the mobile phone calculates the degree of matching between the fused first, second, and third voice components and a preset user voiceprint model. Furthermore, the mobile phone can also perform user authentication based on the degree of matching. In this authentication method, the preset user voiceprint models are fused into one, thereby reducing the complexity and required storage space of the voiceprint model. In addition, since the voiceprint feature information of the second voice component is used, dual voiceprint assurance and liveness detection functions are achieved.
[0232] As another example, the algorithm for calculating the degree of matching may be a similarity calculation. The mobile phone performs feature extraction on the first voice component to obtain a first voiceprint feature, calculates a first similarity between the first voiceprint feature and a first registered voiceprint feature of a preset user that has been pre-stored, calculates a second similarity between the second voiceprint feature and a second registered voiceprint feature of a preset user that has been pre-stored, and calculates a third similarity between the third voiceprint feature and a third registered voiceprint feature of a preset user that has been pre-stored, and performs user authentication based on the first similarity, the second similarity, and the third similarity. Methods for calculating the similarity include Euclidean distance, cosine similarity, Pearson correlation coefficient, adjusted cosine similarity, Hamming distance, Manhattan distance, etc. This is not limited to this application.
[0233] The method for performing identity authentication on a user may be as follows: the mobile phone separately determines a first fusion coefficient corresponding to the first similarity, a second fusion coefficient corresponding to the second similarity, and a third fusion coefficient corresponding to the third similarity based on the decibels of the ambient sound and the playback volume of the Bluetooth headset, and fuses the first similarity, the second similarity, and the third similarity based on the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient to obtain a fusion similarity score. If the fusion similarity score is greater than a first threshold, the mobile phone determines that the user who inputs voice information into the Bluetooth headset is a preset user.
[0234] In a possible implementation, the decibels of the ambient sound are detected by a sound pressure sensor in the Bluetooth headset and sent to the mobile phone, and the playback volume may be obtained by detecting the playback signal by the speaker in the Bluetooth headset and sending it to the mobile phone, or may be obtained by the mobile phone by calling data on the mobile phone.
[0235] In a possible implementation, the second fusion coefficient is negatively correlated with the decibels of the ambient sound, the first fusion coefficient and the third fusion coefficient are each negatively correlated with the decibels of the playback volume, and the sum of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient is a fixed value. Specifically, when the sum of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient is a preset fixed value, a larger decibel of the ambient sound indicates a smaller second fusion coefficient. In this case, the first fusion coefficient and the third fusion coefficient are adaptively increased to keep the sum of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient unchanged. A higher playback volume indicates a smaller first fusion coefficient and a smaller third fusion coefficient. In this case, the second fusion coefficient is adaptively increased to keep the sum of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient unchanged. In this implementation, it can be understood that the fusion coefficients are dynamic. In other words, the fusion coefficient dynamically changes based on the ambient sound and the playback volume, and the fusion coefficient is dynamically determined based on the decibels of the ambient sound detected by the microphone and the playback volume detected by the in-ear sensor. If the decibels of the ambient sound are high, this indicates a high ambient noise level, and the Bluetooth headset may be considered to be more affected by the ambient noise. Therefore, in the audio control method provided in the present application, the fusion coefficients corresponding to the extra-ear sensors and bone vibration sensors of the Bluetooth headset need to be small, and the fusion similarity score result depends more on the in-ear sensors, which are less affected by the ambient noise. Conversely, if the playback volume is high, this indicates a high noise level of the playback sound in the ear canal, and the in-ear sensors of the Bluetooth headset may be considered to be more affected by the playback sound. Therefore, in the audio control method provided in the present application, the fusion coefficients corresponding to the in-ear sensors need to be small, and the fusion similarity score result depends more on the extra-ear sensors and bone vibration sensors, which are less affected by the playback sound.
[0236] Specifically, when designing a system, a lookup table may be set based on the above principle. In specific use, the fusion coefficient may be determined based on the monitored volume of the Bluetooth headset and the decibels of the ambient sound by searching the table. For example, Table 1-1 shows an example. The fusion coefficients of the similarity scores of the audio signals captured by the in-ear audio sensor and the bone vibration sensor are represented by a1 and a2, respectively, and the fusion coefficient of the similarity score of the audio signal captured by the out-of-ear audio sensor is represented by b1. When the ambient sound exceeds 60 dB, it can be considered that the external environment is noisy and the audio signal captured by the out-of-ear audio sensor contains a lot of ambient noise, and the fusion coefficient corresponding to the audio signal captured by the out-of-ear audio sensor can have a low value or can be directly set to 0. When the playback volume of the speaker inside the headset exceeds 80% of the full volume, it can be considered that the volume inside the headset is too high and the fusion coefficient corresponding to the audio signal captured by the in-ear audio sensor can have a low value or can be directly set to 0. When external ambient noise is too high (e.g., ambient noise exceeds 60 dB) and the speaker volume is too high (e.g., headset speaker volume exceeds 60% of full volume), the interference with the captured audio signal is too high and voiceprint recognition fails. It can be understood that, in certain applications, "20% volume," "40% volume," "20 dB ambient sound," and "40 dB ambient sound" can represent ranges. For example, "20% volume" indicates "10% to 30% volume," "40% volume" indicates "30% to 50% volume," "20 dB ambient sound" indicates "10 dB to 30 dB ambient sound," and "40 dB ambient sound" indicates "30 dB to 50 dB ambient sound." [Table 1]
[0237] It can be understood that the above specific designs are merely examples. Specific parameter settings, specific threshold settings, and coefficients corresponding to different decibels of ambient sound and speaker volume can be designed and modified based on actual situations, which is not limited in the present application. It should be noted that the fusion coefficient provided in this embodiment of the present application can be understood as a "dynamic fusion coefficient." In other words, the fusion coefficient can be dynamically adjusted based on different decibel values of ambient sound and speaker volume.
[0238] For example, in another possible implementation, in S706, the policy of authenticating the user based on the fusion of the first, second, and third voiceprint recognition results may be changed to a method of directly fusing voice features, extracting voiceprint features based on the fused voice features and a voiceprint model, calculating the similarity between the voiceprint features and the user's pre-stored enrolled voiceprint features, and authenticating the user. Specifically, voice features feaE1 and feaE2 for each frame are extracted from the current user's voice signal captured by the in-ear voice sensor and the out-of-ear voice sensor. An audio feature feaB1 for each frame is extracted from the current user's voice signal, which is the voice signal captured by the bone voiceprint sensor. The fusion of the acoustic features feaE1, feaE2, and feaB1 includes, but is not limited to, the following methods: performing normalization on feaE1, feaE2, and feaB1 to obtain feaE1', feaE2', and feaB1', and then splicing feaE1', feaE2', and feaB1' into a feature vector fea = [feaE1', feaE2', feaB1']. Using a voiceprint model, voiceprint feature extraction is performed on the feature vector fea to obtain the voiceprint features of the current user. Similarly, the voiceprint features of the enrolled user can be obtained from the enrollment voice of the enrolled user by referring to the above-mentioned method. A similarity comparison is performed between the voiceprint features of the current user and the voiceprint features of the enrolled user to obtain a similarity score, and the relationship between the similarity score and a preset threshold is determined to obtain an authentication result.
[0239] For example, in another possible implementation, the policy of performing user authentication based on the fusion of the first similarity, the second similarity, and the third similarity in S706 may be changed to a method of fusing the first voiceprint feature, the second voiceprint feature, and the third voiceprint feature to obtain a fused voiceprint feature, calculating the similarity between the fused voiceprint feature and a pre-stored enrolled fused voiceprint feature of a preset user, and performing user authentication. Specifically, by using a voiceprint model, feature extraction is performed on the current user's voice signal captured by the in-ear sound sensor and the extra-ear sound sensor to obtain voiceprint features e1 and e2. By using the voiceprint model, feature extraction is performed on the current user's voice signal, which is the voice signal captured by the bone voiceprint sensor, to obtain voiceprint feature b1. The voiceprint features e1, e2, and b1 are spliced and fused to obtain a spliced voiceprint feature m1 = [e1, e2, b1] of the current user. Similarly, the spliced voiceprint feature of an enrolled user can be obtained from the enrolled voice of the enrolled user by referring to the above-mentioned method. A similarity comparison is performed between the spliced voiceprint feature of the current user and the spliced voiceprint feature of the enrolled user to obtain a similarity score, and a relationship between the similarity score and a preset threshold is determined to obtain an authentication result.
[0240] S707: If the user is a preset user, the mobile phone executes the operation instruction corresponding to the voice information.
[0241] In the authentication process of step S706, if the authentication is successful, the mobile phone determines that the speaking user who input the voice information in step S702 is a preset user, and the mobile phone may execute an operation instruction corresponding to the voice information; otherwise, if the authentication is unsuccessful, the mobile phone will not execute any subsequent operation instruction. It is understood that the operation instruction includes, but is not limited to, an operation to unlock the mobile phone or a payment confirmation operation. For example, when the voice information is "Hey Celia, pay by using WeChat," the operation instruction corresponding to the voice information is to display the payment interface of the WeChat app. In this way, after generating the operation instruction to display the payment interface of the WeChat app, the mobile phone may automatically open the WeChat app and display the payment interface of the WeChat app.
[0242] In addition, since the mobile phone determines that the user is a preset user, as shown in FIG. 9, if the mobile phone is currently in a locked state, the mobile phone may unlock the screen, and then execute the operation instruction to display the payment interface of the WeChat app, to display the payment interface of the WeChat app 901.
[0243] For example, the voice control method provided in steps S701 to S707 may be a function provided by a voice assistant app. When the Bluetooth headset interacts with the mobile phone, if it determines through voiceprint recognition that the currently speaking user is a preset user, the mobile phone may send data such as generated operation instructions or voice information to the voice assistant app running on the application layer. Furthermore, the voice assistant app invokes related interfaces or services on the application framework layer to execute operation instructions corresponding to the voice information.
[0244] It can be seen that the voice control method provided in this embodiment of the present application can identify the user's identity based on the voiceprint, while unlocking the mobile phone and executing operation instructions corresponding to the voice information. In other words, the user only needs to input voice information once to complete a series of operations such as user authentication, unlocking the mobile phone, and enabling the functions of the mobile phone, which can greatly improve the user's control efficiency and user experience over the mobile phone.
[0245] In steps S701-S707, a mobile phone is used as an execution entity to perform operations such as voiceprint recognition and user identity authentication. It can be understood that steps S701-S707 can alternatively be completed in whole or in part by a Bluetooth headset to reduce the implementation complexity of the mobile phone and the power consumption of the mobile phone. As shown in Figure 10, the voice control method can include the following steps:
[0246] S1001: The mobile phone establishes a Bluetooth connection to the Bluetooth headset.
[0247] S1002 (optional): The Bluetooth headset detects whether the Bluetooth headset is in a worn state.
[0248] S1003: When the Bluetooth headset is in a worn state, the Bluetooth headset performs capture by using a first audio sensor to obtain a first audio component of audio information input by a user, captures a second audio component of the audio information by using a second audio sensor, and captures a third audio component of the audio information by using a bone vibration sensor first audio sensor.
[0249] In steps S1001 to S1003, for specific methods of establishing a Bluetooth connection between the Bluetooth headset and the mobile phone, detecting whether the Bluetooth headset is being worn, and detecting the first, second, and third audio components of the audio information, please refer to the relevant descriptions of steps S701 to S703, and the details will not be described again here.
[0250] It should be noted that after obtaining the first sound component, the second sound component, and the third sound component, the Bluetooth headset may further perform enhancement, noise reduction, filtering, etc. on the detected first sound component and the detected second sound component, which is not limited in this embodiment of the present application.
[0251] In some embodiments of the present application, the Bluetooth headset has an audio playback function, so that when the speaker of the Bluetooth headset is operating, the air conduction microphone and the bone conduction microphone on the Bluetooth headset can receive echo signals of the audio source played by the speaker. Therefore, after obtaining the first audio component and the second audio component, the Bluetooth headset can further cancel the echo signals in each of the first audio component and the second audio component based on an adaptive echo cancellation algorithm (AEC) to improve the accuracy of subsequent voiceprint recognition.
[0252] S1004: The Bluetooth headset performs voiceprint recognition on each of the first voice component, the second voice component, and the third voice component, and obtains a first voiceprint recognition result corresponding to the first voice component, a second voiceprint recognition result corresponding to the second voice component, and a third voiceprint recognition result corresponding to the third voice component.
[0253] Unlike steps S701 to S707, in step S1004, the Bluetooth headset may pre-store one or more voiceprint models and registered voiceprint features of preset users. After acquiring the first, second, and third voice components in this manner, the Bluetooth headset may perform voiceprint recognition on the first, second, and third voice components by using the voiceprint models locally stored in the Bluetooth headset, individually acquire voiceprint features corresponding to these voice components, and compare the voiceprint features corresponding to the acquired voice components with the corresponding registered voiceprint features. Voiceprint recognition is thus performed. For a specific method by which the Bluetooth headset performs voiceprint recognition on each of the first, second, and third voice components, please refer to the specific method by which the mobile phone performs voiceprint recognition on each of the first, second, and third voice components in step S705. Therefore, a detailed description thereof will be omitted here.
[0254] S1005: The Bluetooth headset performs user authentication based on the first voiceprint recognition result, the second voiceprint recognition result, and the third voiceprint recognition result.
[0255] For the process in which the Bluetooth headset authenticates the user based on the first, second, and third voiceprint recognition results, please refer to the related description of the process in which the mobile phone authenticates the user based on the first, second, and third voiceprint recognition results in step S706, and the details will not be described again here.
[0256] S1006: If the user is a preset user, the Bluetooth headset sends an operation instruction corresponding to the voice information to the mobile phone via the Bluetooth connection.
[0257] S1007: The mobile phone executes the operation instruction.
[0258] If the Bluetooth headset determines that the speaking user who inputs the voice information is a preset user, the Bluetooth headset may generate an operation instruction corresponding to the voice information. For the operation instruction, please refer to the example of the operation instruction of the mobile phone in step S707. The details will not be described again here.
[0259] In addition, since the Bluetooth headset determines that the user is a preset user, when the mobile phone is in a locked state, the Bluetooth headset may further send a message or an unlock instruction to the mobile phone indicating that the user identity authentication has been successful, so that the mobile phone can unlock the screen and then execute an operation instruction corresponding to the voice information. Of course, the Bluetooth headset may alternatively send the captured voice information to the mobile phone, and the mobile phone may generate a corresponding operation instruction based on the voice information and execute the operation instruction.
[0260] In some embodiments of the present application, when transmitting voice information or corresponding operation instructions to the mobile phone, the Bluetooth headset may further transmit a device identifier (e.g., a MAC address) of the Bluetooth headset to the mobile phone. Since the mobile phone stores identifiers of preset Bluetooth devices that have been successfully authenticated, the mobile phone may determine whether the currently connected Bluetooth headset is a preset Bluetooth device based on the received device identifier. If the Bluetooth headset is a preset Bluetooth device, the mobile phone may further execute the operation instructions transmitted by the Bluetooth headset or perform voice recognition or the like on the voice information transmitted by the Bluetooth headset; otherwise, if the Bluetooth headset is not a preset Bluetooth device, the mobile phone may discard the operation instructions transmitted by the Bluetooth headset to avoid security issues arising when an unauthorized Bluetooth device maliciously controls the mobile phone.
[0261] Alternatively, the mobile phone and the preset Bluetooth device may pre-agreed on a passcode or password for transmitting operation instructions. In this way, when transmitting voice information and corresponding operation instructions to the mobile phone, the Bluetooth headset may also transmit the pre-agreed on passcode or password to the mobile phone, so that the mobile phone determines whether the currently connected Bluetooth headset is a preset Bluetooth device.
[0262] Alternatively, the mobile phone and the preset Bluetooth device may pre-agree on an encryption algorithm and a decryption algorithm for transmitting the operation instruction. In this way, before transmitting the voice information or the corresponding operation instruction to the mobile phone, the Bluetooth headset may encrypt the operation instruction based on the agreed-upon encryption algorithm. After receiving the encrypted operation instruction, if the operation instruction can be obtained through decryption based on the agreed-upon decryption algorithm, this indicates that the currently connected Bluetooth headset is a preset Bluetooth device, and the mobile phone may further execute the operation instruction sent by the Bluetooth headset. Alternatively, if the mobile phone cannot obtain the operation instruction through decryption based on the agreed-upon decryption algorithm, this indicates that the currently connected Bluetooth headset is an unauthorized Bluetooth device, and the mobile phone may discard the operation instruction sent by the Bluetooth headset.
[0263] It should be noted that steps S701 to S707 and steps S1001 to S1007 are merely two implementations of the voice control method provided in the present application. It can be understood that those skilled in the art may set the specific steps performed by the Bluetooth headset and the specific steps performed by the mobile phone in the above-described embodiment based on actual application scenarios or practical experience. This is not limited to the embodiments of the present application. In addition, the voice control method provided in the present application may alternatively be performed by a server, i.e., the Bluetooth headset establishes a connection to the server, and the server implements the functions of the mobile phone in the above-described embodiment. The specific process will not be described again.
[0264] For example, after performing voiceprint recognition on the first voice component, the second voice component, and the third voice component, the Bluetooth headset may alternatively transmit the obtained first voiceprint recognition result, the second voiceprint recognition result, and the third voiceprint recognition result to the mobile terminal, and then the mobile terminal performs user identity authentication, etc. based on the voiceprint recognition results.
[0265] In another example, after obtaining the first, second, and third audio components, the Bluetooth headset may alternatively first determine whether voiceprint recognition needs to be performed on the first, second, and third audio components. If voiceprint recognition needs to be performed on the first, second, and third audio components, the Bluetooth headset may send the first, second, and third audio components to the mobile phone, so that the mobile phone can complete subsequent voiceprint recognition, user authentication, etc. Alternatively, if voiceprint recognition does not need to be performed on the first, second, and third audio components, the Bluetooth headset does not need to send the first, second, and third audio components to the mobile phone, thereby avoiding increased power consumption that would occur when the mobile phone processes the first, second, and third audio components.
[0266] 11(a), the user may further access the setting interface 1101 of the mobile phone to enable or disable the voice control function. When the user enables the voice control function, the user may use the setting button 1102 to set a keyword for triggering the voice control function, such as "Hey Celia" or "Pay," or the user may use the setting button 1103 to manage preset user voiceprint models, such as adding or deleting preset user voiceprint models, or the user may use the setting button 1104 to set operation instructions that can be supported by the voice assistant, such as paying, making a call, ordering a meal, etc. In this way, the user can have a customized voice control experience.
[0267] In some embodiments of the present application, the embodiments of the present application disclose a voice control device. As shown in FIG. 12, the voice control device includes a voice information acquiring unit 1201, a recognition unit 1202, an identification information acquiring unit 1203, and an execution unit 1204. It can be understood that the voice control device may be a terminal or a wearable device. The voice control device may be fully integrated into the wearable device, and the wearable device and the terminal may form a voice control system. In other words, some units are arranged in the wearable device, and some units are arranged in the terminal.
[0268] In a possible implementation, for example, the audio control device may be fully integrated into a Bluetooth headset. The audio information acquiring unit 1201 is configured to acquire audio information of a user. In this embodiment of the present application, a user may input audio information into the Bluetooth headset when wearing the Bluetooth headset. In this case, the Bluetooth headset may capture a first audio component by using an in-ear audio sensor, a second audio component by using an out-of-ear audio sensor, and a third audio component by using a bone vibration sensor based on the audio information input by the user.
[0269] The recognition unit 1202 is configured to perform voiceprint recognition on each of the first speech component, the second speech component, and the third speech component, and obtain a first voiceprint recognition result corresponding to the first speech component, a second voiceprint recognition result corresponding to the second speech component, and a third voiceprint recognition result corresponding to the third speech component.
[0270] In a possible implementation, the recognition unit 1202 may be further configured to perform keyword detection on voice information input by a user into the Bluetooth headset, and to perform voiceprint recognition on each of the first voice component, the second voice component, and the third voice component when the voice information includes a preset keyword, or the recognition unit 1202 may be configured to detect a user input and perform voiceprint recognition on each of the first voice component, the second voice component, and the third voice component upon receiving a preset operation input by the user. The user input may be a user input into the Bluetooth headset by using a touch screen or a button. For example, the user taps the unlock button on the Bluetooth headset. Optionally, before the recognition unit 1202 performs keyword detection on the voice information or detects a user input, Audio information The acquisition unit 1201 may further acquire the wearing state detection result, and if the wearing state detection result is successful, the recognition unit 1202 performs keyword detection on the voice information or detects user input.
[0271] In a possible implementation, the recognition unit 1202 is specifically configured to: perform feature extraction on a first voice component to obtain a first voiceprint feature, and calculate a first similarity between the first voiceprint feature and a first enrollment voiceprint feature of a preset user, where the first enrollment voiceprint feature is obtained by performing feature extraction on the first enrollment voice using a first voiceprint model, and the first enrollment voiceprint feature represents a voice feature of the preset user, which is a voice feature captured by an in-ear voice sensor; perform feature extraction on a second voice component to obtain a second voiceprint feature, and calculate a second similarity between the second voiceprint feature and a second enrollment voiceprint feature of the preset user, where , a second enrollment voiceprint feature is obtained by performing feature extraction on the second enrollment voice by using a second voiceprint model, the second enrollment voiceprint feature being a voice feature of a preset user and representing a voice feature captured by an extra-aural voice sensor; performing feature extraction on a third voice component to obtain a third voiceprint feature, and calculating a third similarity between the third enrollment voiceprint feature and the third enrollment voiceprint feature of the preset user, wherein the third enrollment voiceprint feature is obtained by performing feature extraction on the third enrollment voice by using a third voiceprint model, the third enrollment voiceprint feature being a voice feature of a preset user and representing a voice feature captured by a bone vibration sensor.
[0272] In a possible implementation, the first enrollment voiceprint feature is obtained by performing feature extraction using a first voiceprint model, the first enrollment voiceprint feature being a voiceprint feature of a preset user and representing a voiceprint feature captured by an in-ear sound sensor, the second enrollment voiceprint feature is obtained by performing feature extraction using a second voiceprint model, the second enrollment voiceprint feature being a voiceprint feature of a preset user and representing a voiceprint feature captured by an out-of-ear sound sensor, and the third enrollment voiceprint feature is obtained by performing feature extraction using a third voiceprint model, the third enrollment voiceprint feature being a voiceprint feature of a preset user and representing a voiceprint feature captured by a bone vibration sensor.
[0273] The identification information acquisition unit 1203 is configured to acquire user identification information and perform user authentication. Specifically, the identification information acquisition unit 1203 is configured to separately determine a first fusion coefficient corresponding to a first similarity, a second fusion coefficient corresponding to a second similarity, and a third fusion coefficient corresponding to a third similarity based on the decibels of the ambient sound and the playback volume, and then fuse the first similarity, the second similarity, and the third similarity based on the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient to obtain a fusion similarity score. If the fusion similarity score is greater than a first threshold, the mobile phone determines that the user who inputs the audio information into the Bluetooth headset is a preset user. The decibels of the ambient sound can be detected by a sound pressure sensor of the Bluetooth headset, and the playback volume can be obtained by detecting a playback signal from a speaker of the Bluetooth headset.
[0274] In a possible implementation, the second fusion coefficient is negatively correlated with the decibels of the ambient sound, the first fusion coefficient and the third fusion coefficient are each negatively correlated with the decibels of the playback volume, and the sum of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient is a fixed value. Specifically, when the sum of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient is a preset fixed value, a larger decibel of the ambient sound indicates a smaller second fusion coefficient. In this case, the first fusion coefficient and the third fusion coefficient are adaptively increased, while the sum of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient is kept unchanged. A higher playback volume indicates a smaller first fusion coefficient and a smaller third fusion coefficient. In this case, the second fusion coefficient is adaptively increased, while the sum of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient is kept unchanged. It can be understood that based on the variable fusion coefficient, recognition accuracy can be considered in different application scenarios (such as a noisy environment or playing music using a headset).
[0275] After the mobile phone determines that the user who inputs the voice information into the Bluetooth headset is a preset user, or after the authentication is successful, the execution unit 1204 is configured to execute an operation instruction corresponding to the voice information, such as an unlock instruction, a payment instruction, a power off instruction, an application launch instruction, or a call instruction, etc.
[0276] Compared with the prior art, the voice control method provided in this embodiment of the present application adds a method for capturing voiceprint features by using an in-ear voice sensor. After a user wears a headset including an in-ear voice sensor, a closed cavity is formed by the external ear canal and the middle ear canal, which has a specific amplification effect on the sound within the cavity, i.e., a cavity effect. Therefore, the sound captured by the in-ear voice sensor is clearer, with a significant enhancement effect, especially for high-frequency voice signals. This can compensate for the distortion caused when high-frequency signal components of some voice information are lost when the bone vibration sensor captures voice information, improving the overall voiceprint capture effect and voiceprint recognition accuracy of the headset and improving the user experience. In addition, this embodiment of the present application uses a dynamic fusion coefficient when fusing similarities. By using the dynamic fusion coefficient, voiceprint recognition results obtained for voice signals with different attributes can be fused for different application environments and scenarios, allowing the voice signals with different attributes to compensate for each other and improving the robustness and accuracy of voiceprint recognition. For example, recognition accuracy can be significantly improved in noisy environments or when playing music using a headset. Audio signals with different attributes may also be understood as audio signals obtained by using different sensors (in-ear audio sensor, extra-ear audio sensor and bone vibration sensor).
[0277] Another embodiment of the present application further provides a wearable device. Figure 13 is a schematic diagram of a wearable device 130 according to an embodiment of the present application. The wearable device shown in Figure 13 includes a memory 1301, a processor 1302, a communication interface 1303, a bus 1304, an in-ear sound sensor 1305, an out-of-ear sound sensor 1306, and a bone vibration sensor 1307. The memory 1301, the processor 1302, and the communication interface 1303 are communicatively connected to each other via the bus 1304. The memory 1301 is coupled to the processor 1302. Memory 1301 The processor is configured to store computer program code, which includes computer instructions. 1302 When the executes the computer instructions, the wearable device can perform the voice control methods described in the above embodiments.
[0278] The in-ear sound sensor 1305 is configured to capture a first sound component of the sound information, the out-of-ear sound sensor 1306 is configured to capture a second sound component of the sound information, and the bone vibration sensor 1307 is configured to capture a third sound component of the sound information.
[0279] The memory 1301 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1301 may store a program. When the program stored in the memory 1301 is executed by the processor 1302, the processor 1302 and the communication interface 1303 are configured to execute steps of the voice control method in the embodiment of the present application.
[0280] The processor 1302 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), a graphics processing unit (GPU) or one or more integrated circuits, and is configured to execute associated programs to implement functions that need to be performed by units in a voice control device in an embodiment of the present application or to perform a voice control method in a method embodiment of the present application.
[0281] The processor 1302 may alternatively be an integrated circuit chip and have signal processing capabilities. In the implementation process, the steps of the voice control method in the present application may be completed by using a hardware integrated logic circuit in the processor 1302 or by using instructions in the form of software. The processor 1302 may alternatively be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or another programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The processor 1302 may implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc. The steps of the methods disclosed with reference to the embodiments of the present application may be directly executed and completed by a hardware decoding processor, or may be executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be arranged in a storage medium that is mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, or a register. The storage medium is arranged in the memory 1301. The processor 1302 reads information in the memory 1301 and, in combination with the hardware of the processor 1302, completes the functions that need to be performed by the units included in the voice control device in the embodiments of the present application, or executes the voice control method in the method embodiments of the present application.
[0282] The communication interface 1303 may perform wired or wireless communication by using a transceiver device, including but not limited to a transceiver, so that the wearable device 1300 can communicate with another device or a communication network. For example, the wearable device may establish a communication connection to a terminal device via the communication interface 1303.
[0283] The bus 1304 may include a path for transmitting information between various components of the device 1300 (eg, the memory 1301, the processor 1302, and the communication interface 1303).
[0284] Another embodiment of the present application further provides a terminal. Figure 14 is a schematic diagram of a terminal according to one embodiment of the present application. The terminal shown in Figure 14 includes a touchscreen 1401, a processor 1402, a memory 1403, one or more computer programs 1404, a bus 1405, and a communication interface 1408. The touchscreen 1401 includes a touch-sensitive surface 1406 and a display 1407, and the terminal may further include one or more applications (not shown). The components may be connected via one or more communication buses 1405.
[0285] The memory 1403 is coupled to the processor 1402. The memory 1403 is configured to store computer program code. The computer program code includes computer instructions. When the processor 1402 executes the computer instructions, the terminal can perform the voice control method described in the previous embodiments.
[0286] Touchscreen 1401 is configured to interact with a user and can receive user input. A user provides input to the mobile phone on touch-sensitive surface 1406. For example, the user taps an unlock button displayed on touch-sensitive surface 1406 of the mobile phone.
[0287] The memory 1403 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1403 may store a program. When the program stored in the memory 1403 is executed by the processor 1402, the processor 1402 and the communication interface 1408 is configured to perform the steps of the voice control method in the embodiment of the present application.
[0288] The processor 1402 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), a graphics processing unit (GPU) or one or more integrated circuits and is configured to execute associated programs to implement functions that need to be performed by units in a voice control device in an embodiment of the present application or to perform a voice control method in a method embodiment of the present application.
[0289] The processor 1402 may alternatively be an integrated circuit chip and have signal processing capabilities. In the implementation process, the steps of the voice control method in the present application may be completed by using a hardware integrated logic circuit in the processor 1402 or by using instructions in the form of software. The processor 1402 may alternatively be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or another programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The processor 1402 may implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc. The steps of the methods disclosed with reference to the embodiments of the present application may be directly executed and completed by a hardware decoding processor, or may be executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be arranged in a storage medium that is mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, or a register. The storage medium is arranged in the memory 1403. The processor 1402 reads information in the memory 1403 and, in combination with the hardware of the processor 1402, completes the functions that need to be performed by the units included in the voice control device in the embodiments of the present application, or executes the voice control method in the method embodiments of the present application.
[0290] The communication interface 1408 may perform wired or wireless communication by using a transceiver device, including but not limited to a transceiver, so that the terminal 1400 can communicate with another device or a communication network. For example, the terminal may establish a communication connection to a wearable device via the communication interface 1408.
[0291] The bus 1304 may include a path for transmitting information between various components of the device 1400 (eg, the touchscreen 1401, the memory 1403, the processor 1402, and the communication interface 1408).
[0292] 13 and 14 only show the memory, processor, communication interface, etc. of wearable device 1300 and terminal 1400, but in a specific implementation process, those skilled in the art should understand that wearable device 1300 and terminal 1400 may each further include other components necessary for normal operation. In addition, based on specific requirements, those skilled in the art should understand that wearable device 1300 and terminal 1400 may each further include hardware components for implementing other additional functions. In addition, those skilled in the art should understand that wearable device 1300 and terminal 1400 may each include only the components necessary to implement an embodiment of the present application, and do not necessarily need to include all the components shown in FIG. 13 or 14.
[0293] Another embodiment of the present application further provides a chip system. FIG. 15 is a schematic diagram of the chip system. The chip system includes at least one processor 1501, at least one interface circuit 1502, and a bus 1503. The processor 1501 and the interface circuit 1502 may be interconnected via a line. For example, the interface circuit 1502 may be configured to receive a signal from another device (e.g., a memory of a voice control device). In another example, the interface circuit 1502 may be configured to send a signal to another device (e.g., the processor 1501). For example, the interface circuit 1502 may read an instruction stored in a memory and send the instruction to the processor 1501. When the instruction is executed by the processor 1501, the voice control device can be enabled to perform the steps in the above-mentioned embodiment. Of course, the chip system may further include another individual device. This is not particularly limited in this embodiment of the present application.
[0294] Another embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions that, when executed on a voice control device, cause the voice control device to perform the steps performed by the recognition device in the method procedures shown in the above-described method embodiments.
[0295] Another embodiment of the present application further provides a computer program product, which stores computer instructions that, when executed on a recognition device of a voice control device, cause the recognition device to perform the steps performed by the recognition device in the method procedures set forth in the method embodiments described above.
[0296] In some embodiments, the disclosed methods may be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or encoded on another non-transitory medium or article of manufacture.
[0297] In one embodiment, a computer program product is provided using a signal-bearing medium. The signal-bearing medium may include one or more program instructions. When the program instructions are executed by one or more processors, the functions of the voice control method in the embodiment of the present application may be implemented. Thus, for example, one or more features in S701 to S707 of FIG. 7 may be carried by one or more instructions associated with the signal-bearing medium.
[0298] In some examples, the signal-bearing medium may include a computer-readable medium, including, but not limited to, a hard disk drive, a compact disc (CD), a digital video disc (DVD), digital tape, memory, read-only memory (ROM), random access memory (RAM), etc.
[0299] In some implementations, the signal-bearing medium may include computer-recordable media, including but not limited to memory, read / write (R / W) CDs, R / W DVDs, and the like.
[0300] In some implementations, signal-bearing media may include communications media, including, but not limited to, digital and / or analog communications media (e.g., fiber optic cables, wave guides, wired communications links, wireless communications links), and the like.
[0301] The signal-bearing medium may be conveyed by a wireless form of communication medium (e.g., a wireless communication medium conforming to the IEEE 802.16 standard or another transmission protocol). The one or more program instructions may be, for example, one or more computer-executable instructions or one or more logic-implemented instructions.
[0302] Based on the description of the implementation, those skilled in the art can clearly understand that the division into functional modules is used merely as an example for convenience and concise description. In actual application, functions can be allocated to different functional modules for implementation based on requirements. In other words, the internal structure of the device is divided into different functional modules to implement all or part of the above-mentioned functions. For specific operation processes of the system, device, and unit, please refer to the corresponding processes in the method embodiments. Details will not be described again here.
[0303] Those skilled in the art may recognize that, in combination with the examples described in the embodiments disclosed herein, the units and algorithm steps may be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether a function is performed by hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but the implementation should not be considered to go beyond the scope of this application.
[0304] The functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0305] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, the integrated unit may be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application may be essentially implemented in the form of a software product, or a part that contributes to the prior art, or all or part of the technical solutions. The computer software product is stored in a storage medium and includes some instructions for instructing a computer device (which may be a personal computer, a server, or a network device) or a processor to perform all or part of the steps of the methods described in the embodiments of the present application. The above-mentioned storage medium includes any medium that can store program code, such as a flash memory, a removable hard disk, a read-only memory, a random access memory, a magnetic disk, or an optical disk.
[0306] In some embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods may be implemented in other manners. For example, the described device embodiments are merely examples. For example, the division into units is merely a logical function division, and other divisions may be used in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some functions may be ignored or not performed. In addition, the shown or discussed mutual couplings or direct couplings or communication connections may be implemented by using some interfaces. Indirect couplings or communication connections between devices or units may be implemented in electronic, mechanical, or other forms.
[0307] Units described as separate parts may or may not be physically separate, and parts shown as units may or may not be physical units, and may be located in one location or distributed across multiple network units. Some or all of the units may be selected based on actual requirements to achieve the objectives of the solutions of the embodiments.
[0308] The above description is merely a specific implementation of the embodiments of the present application, but is not intended to limit the protection scope of the embodiments of the present application. Any modifications or replacements within the technical scope disclosed in the embodiments of the present application shall be included in the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application shall be subject to the protection scope of the claims.
Claims
1. 1. A method of audio control performed by a wearable device, the wearable device including a headset including an in-ear audio sensor, an out-of-ear audio sensor, and a bone vibration sensor, the method comprising: acquiring user audio information, the audio information including a first audio component, a second audio component, and a third audio component, the first audio component being captured by the in-ear audio sensor, the second audio component being captured by the extra-ear audio sensor, and the third audio component being captured by the bone vibration sensor; performing voiceprint recognition on each of the first speech component, the second speech component, and the third speech component, the performing voiceprint recognition comprising: performing feature extraction on the first speech component to obtain first voiceprint features, and calculating a first similarity between the first voiceprint features and first enrollment voiceprint features of the user, wherein the first enrollment voiceprint features are obtained by performing feature extraction on a first enrollment speech by using a first voiceprint model, and the first enrollment voiceprint features represent preset audio features of the user captured by the in-ear sound sensor; performing feature extraction on the second speech component to obtain second voiceprint features, and calculating a second similarity between the second voiceprint features and second enrollment voiceprint features of the user, wherein the second enrollment voiceprint features are obtained by performing feature extraction on a second enrollment speech by using a second voiceprint model, and the second enrollment voiceprint features represent preset audio features of the user captured by the extra-ear sound sensor; performing feature extraction on the third speech component to obtain a third voiceprint feature, and calculating a third similarity between the third voiceprint feature and a third enrollment voiceprint feature of the user, wherein the third enrollment voiceprint feature is obtained by performing feature extraction on a third enrollment speech by using a third voiceprint model, and the third enrollment voiceprint feature represents a preset audio feature of the user captured by the bone vibration sensor; performing voiceprint recognition, including: a step of acquiring identification information of the user based on a voiceprint recognition result of the first voice component, a voiceprint recognition result of the second voice component, and a voiceprint recognition result of the third voice component, wherein the step of acquiring the identification information of the user includes: determining a first fusion coefficient corresponding to the first similarity, a second fusion coefficient corresponding to the second similarity, and a third fusion coefficient corresponding to the third similarity, Obtaining decibels of ambient sound based on a sound pressure sensor; determining a playback volume based on a playback signal of the speaker of the headset; determining each of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient based on the decibels of the ambient sound and the playback volume, wherein the second fusion coefficient is negatively correlated with the decibels of the ambient sound, the first fusion coefficient and the third fusion coefficient are each negatively correlated with the decibels of the playback volume, and a sum of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient is a fixed value; fusing the first similarity, the second similarity, and the third similarity based on the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient to obtain a fusion similarity score; and determining that the identification information of the user matches preset identification information when the fusion similarity score is greater than a first threshold value; obtaining a user's identity, a step of executing an operation instruction when the identification information of the user matches preset information, the operation instruction being determined based on the voice information; A voice control method comprising:
2. Before performing voiceprint recognition on the first speech component, the second speech component, and the third speech component, the method comprises: performing keyword detection or detecting user input on the audio information; The voice control method of claim 1 further comprising:
3. Before performing keyword detection on the audio information or detecting user input, the method comprises: acquiring a detection result of the wearing state of the wearable device; The voice control method of claim 2 further comprising:
4. The operation instruction includes an unlock instruction, a payment instruction, a power-off instruction, an application launch instruction, or a call instruction. The voice control method according to any one of claims 1 to 3.
5. 1. A sound control device integrated into a wearable device, the wearable device including a headset including an in-ear sound sensor, an out-of-ear sound sensor, and a bone vibration sensor, the sound control device comprising: an audio information acquisition unit configured to acquire audio information of a user, the audio information including a first audio component, a second audio component, and a third audio component, the first audio component being captured by the in-ear audio sensor, the second audio component being captured by the extra-ear audio sensor, and the third audio component being captured by the bone vibration sensor; a recognition unit configured to perform voiceprint recognition on each of the first speech component, the second speech component and the third speech component, the recognition unit comprising: configured to perform feature extraction on the first speech component to obtain a first voiceprint feature and calculate a first similarity between the first voiceprint feature and a first enrollment voiceprint feature of the user, wherein the first enrollment voiceprint feature is obtained by performing feature extraction on a first enrollment speech by using a first voiceprint model, and the first enrollment voiceprint feature represents a preset audio feature of the user captured by the in-ear sound sensor; configured to perform feature extraction on the second speech component to obtain second voiceprint features, and calculate a second similarity between the second voiceprint features and second enrollment voiceprint features of the user, wherein the second enrollment voiceprint features are obtained by performing feature extraction on a second enrollment speech by using a second voiceprint model, and the second enrollment voiceprint features represent preset audio features of the user captured by the extra-ear sound sensor; a recognition unit configured to perform feature extraction on the third speech component to obtain a third voiceprint feature, and calculate a third similarity between the third voiceprint feature and a third enrollment voiceprint feature of the user, where the third enrollment voiceprint feature is obtained by performing feature extraction on a third enrollment voice by using a third voiceprint model, and the third enrollment voiceprint feature indicates a preset audio feature of the user captured by the bone vibration sensor; an identification information obtaining unit configured to obtain identification information of the user based on a voiceprint recognition result of the first voice component, a voiceprint recognition result of the second voice component, and a voiceprint recognition result of the third voice component, wherein the identification information obtaining unit: determining a first fusion coefficient corresponding to the first similarity, a second fusion coefficient corresponding to the second similarity, and a third fusion coefficient corresponding to the third similarity; Obtaining decibels of ambient sound based on a sound pressure sensor; determining a playback volume based on a playback signal of the speaker of the headset; determining each of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient based on the decibels of the ambient sound and the playback volume, wherein the second fusion coefficient is negatively correlated with the decibels of the ambient sound, the first fusion coefficient and the third fusion coefficient are each negatively correlated with the decibels of the playback volume, and a sum of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient is a fixed value; an identification information obtaining unit configured to obtain a fusion similarity score by fusing the first similarity, the second similarity, and the third similarity according to the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient, and determine that the identification information of the user matches preset identification information when the fusion similarity score is greater than a first threshold; an execution unit configured to execute an operation instruction when the identification information of the user matches preset information, and the operation instruction is determined based on the voice information; 11. A voice control device comprising:
6. The voice information acquisition unit: further configured to perform keyword detection on the audio information or detect user input. The voice control device according to claim 5 .
7. The voice information acquisition unit: Further configured to acquire a wearing state detection result of the wearable device. The voice control device according to claim 6.
8. The operation instruction includes an unlock instruction, a payment instruction, a power-off instruction, an application launch instruction, or a call instruction. The voice control device according to any one of claims 5 to 7.
9. A wearable device, the wearable device comprising an in-ear sound sensor, an out-of-ear sound sensor, a bone vibration sensor, a memory and a processor; the in-ear sound sensor is configured to capture a first sound component of audio information, the out-of-ear sound sensor is configured to capture a second sound component of the audio information, and the bone vibration sensor is configured to capture a third sound component of the audio information; The memory is coupled to the processor, the memory being configured to store computer program code, the computer program code including computer instructions, and execution of the computer instructions by the processor causes the wearable device to perform the voice control method of any one of claims 1 to 4. Wearable device.
10. A terminal comprising a memory and a processor, the memory coupled to the processor, the memory configured to store computer program code, the computer program code comprising computer instructions, and wherein execution of the computer instructions by the processor causes the terminal to perform the voice control method of any one of claims 1 to 4. Terminal.
11. A chip system applied to an electronic device, the chip system comprising one or more processors and a memory, the memory storing computer instructions which, when executed by the processor, cause the electronic device to perform the voice control method according to any one of claims 1 to 4. Chip system.
12. A computer readable storage medium containing computer instructions, which when executed on a voice control device, enable the voice control device to perform the voice control method of any one of claims 1 to 4. A computer-readable storage medium.
13. A computer program comprising computer instructions, which when executed on a voice control device enable the voice control device to perform the voice control method of any one of claims 1 to 4. Computer program.
Citation Information
Patent Citations
Voice control method, wearable apparatus, and terminal
EP3790006A1
Speech authentication device
JP2007017840A
Portable type personal identification method and electronic commerce method
JP2008033144A
Speaker recognition system
JP2008224911A
Signal processing method and system
JP2011525724A