Voice Control Method, System, Electronic Device and Storage Medium Based on Voiceprint

By determining the voiceprint and wake-up words of the voiceprint to be recognized in offline speech recognition technology, and dynamically adjusting the threshold of the confidence value according to the existence or absence of the voiceprint library, the problem of high misrecognition rate in the prior art is solved, and the accuracy and user experience of voice control are improved.

CN115910049BActive Publication Date: 2025-06-13PATEO CONNECT (NANJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111165986.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-30
Publication Date
2025-06-13
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

Existing offline voice recognition technology is prone to misrecognition in noisy environments, multimedia playback and multi-person conversations, resulting in unstable control functions.

Method used

By determining the voiceprint, wake-up word and confidence value corresponding to the voice to be identified, and dynamically adjusting the threshold of the confidence value according to the existence or not of the voiceprint library, the accuracy of the wake-up operation is improved.

Benefits of technology

It effectively reduces the misrecognition rate caused by environmental interference sound, improves the accuracy and security of voice control, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115910049B_ABST
    Figure CN115910049B_ABST
Patent Text Reader

Abstract

The present invention discloses a voice control method, system, electronic device and storage medium based on voiceprint. The voice control method includes the following steps: obtaining the voice to be recognized; determining the voiceprint to be recognized, wake-up word and confidence value corresponding to the voice to be recognized; when the voiceprint to be recognized is included in the preset voiceprint library and the confidence value is greater than the first trigger threshold, triggering a wake-up operation according to the wake-up word; when the voiceprint to be recognized is not included in the preset voiceprint library and the confidence value is greater than the second trigger threshold, triggering a wake-up operation according to the wake-up word; the first trigger threshold is less than the second trigger threshold. The voice control method of the present invention effectively reduces the probability of misrecognition caused by ambient interference sounds and improves the accuracy and security of voice control and the user experience by requiring the confidence value to be greater than a higher trigger threshold to trigger a wake-up operation when the voiceprint to be recognized is not included in the preset voiceprint library.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This field relates to the field of voice control technology, and particularly to a voice control method, system, electronic device, and storage medium based on voiceprint. Background Art

[0002] Currently, voice control has been widely applied to various electronic devices commonly used in daily life, such as mobile phones, car infotainment systems, smart speakers, etc.; voice recognition is the core technology of voice control technology. Only when the performance and accuracy of voice recognition are high can voice control correctly and quickly complete the actions desired by users. Offline recognition is a voice recognition technology used without the need to connect to the Internet. Compared with online recognition and control, it has a faster response and has advantages over online recognition in certain scenarios.

[0003] For existing offline voice recognition technologies, the basic implementation method is to pre-register some common terms in the offline voice recognition engine, such as "open music", "open phone", etc., and set a threshold and corresponding actions for each term. Every time the user says a sentence, the offline voice engine will calculate a confidence value according to the recognition algorithm. If this value is greater than the preset threshold, it is considered that what the user said is this term, thus hitting the term and executing the corresponding action.

[0004] However, there are still some technical problems in the existing offline voice recognition control function, resulting in a relatively high misrecognition rate. For example, in a noisy environment, such as when it is relatively noisy around, some registered terms are easily misawakened; in the case of multimedia playback, if the played content contains a registered term, it will be misawakened; in the case of multi-person conversations, if the conversation content contains a registered term, it is easily misawakened. Summary of the Invention

[0005] An object of the present invention is to provide a voice control method based on voiceprint. The advantage is that by determining the voiceprint to be recognized, wake-up word, and confidence value corresponding to the voice to be recognized, when the preset voiceprint library includes the voiceprint to be recognized, the confidence value only needs to be greater than a lower trigger threshold to trigger the wake-up operation, and when the preset voiceprint library does not include the voiceprint to be recognized, the confidence value needs to be greater than a higher trigger threshold to trigger the wake-up operation, effectively reducing the probability of misrecognition caused by ambient interference sounds, improving the accuracy and security of voice control, and enhancing the user experience.

[0006] An object of the present invention is to provide a voice control method based on voiceprint. The advantage is that by setting a higher threshold to filter out interfering sounds with low confidence values, and after the voice to be recognized passes the threshold test, while saving the voiceprint to be recognized, maintaining a lower trigger threshold associated with the voiceprint to facilitate the user's next voice wake-up operation, enhancing the intelligence of the voice control system and improving the user experience.

[0007] An object of the present invention is to provide a voice control method based on voiceprint, which has the advantage that when the preset voiceprint library includes the voiceprint to be recognized but the wake-up operation cannot be triggered because the confidence value is less than or equal to the first trigger threshold, the first trigger threshold of the voiceprint is dynamically reduced, so that the user can trigger the wake-up operation in a normal speaking manner, enhancing the intelligence of the voice control system and improving the user experience.

[0008] An object of the present invention is to provide a voice control method based on voiceprint, which has the advantage that by associating the first trigger threshold and the second trigger threshold with the voiceprint to be recognized respectively, each voiceprint to be recognized has a corresponding first trigger threshold and second trigger threshold, and different first trigger thresholds and second trigger thresholds can be set according to the characteristics of the voiceprint to be recognized, enhancing the accuracy of the voice control system and improving the user experience.

[0009] An object of the present invention is to provide a voice control method based on voiceprint, which has the advantage that by associating the first trigger threshold and the second trigger threshold with the wake-up word respectively, each wake-up word has a corresponding first trigger threshold and second trigger threshold, and different first trigger thresholds and second trigger thresholds can be set according to the characteristics of the wake-up word, enhancing the accuracy of the voice control system and improving the user experience.

[0010] In the first aspect of the present invention, a voice control method based on voiceprint is provided. The voice control method includes the following steps:

[0011] Obtain the voice to be recognized;

[0012] Determine the voiceprint to be recognized, wake-up word and confidence value corresponding to the voice to be recognized;

[0013] When the preset voiceprint library includes the voiceprint to be recognized and the confidence value is greater than the first trigger threshold, trigger a wake-up operation according to the wake-up word;

[0014] When the preset voiceprint library does not include the voiceprint to be recognized and the confidence value is greater than the second trigger threshold, trigger a wake-up operation according to the wake-up word;

[0015] The first trigger threshold is less than the second trigger threshold.

[0016] In the second aspect of the present invention, a voice control system based on voiceprint is provided. The voice control system includes:

[0017] An acquisition module for acquiring the voice to be recognized;

[0018] A determination module for determining the voiceprint to be recognized, wake-up word and confidence value corresponding to the voice to be recognized;

[0019] A triggering module, configured to trigger a wake-up operation according to the wake-up word when the preset voiceprint library includes the voiceprint to be recognized and the confidence value is greater than a first triggering threshold;

[0020] The triggering module is further configured to trigger a wake-up operation according to the wake-up word when the preset voiceprint library does not include the voiceprint to be recognized and the confidence value is greater than a second triggering threshold;

[0021] The first triggering threshold is less than the second triggering threshold.

[0022] A third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned voice control method is implemented.

[0023] A fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned voice control method is implemented. Description of the Drawings

[0024] Figure 1 It is a flowchart of the voice control method according to Embodiment 1 of the present invention.

[0025] Figure 2 It is a flowchart of the voice control method according to Embodiment 2 of the present invention.

[0026] Figure 3 It is a schematic diagram of the modules of the voice control system according to Embodiment 3 of the present invention.

[0027] Figure 4 It is a schematic diagram of the modules of the voice control system according to Embodiment 4 of the present invention.

[0028] Figure 5 It is a schematic diagram of the hardware structure of the electronic device according to Embodiment 5 of the present invention. Detailed Embodiments

[0029] The present invention will be further described below by way of embodiments, but the present invention is not limited to the scope of the described embodiments.

[0030] Embodiment 1

[0031] Please refer to Figure 1 which is a flowchart of the voice control method based on voiceprint in this embodiment. Specifically, as Figure 1 shown, the voice control method includes the following steps:

[0032] S101. Obtain the voice to be recognized.

[0033] S102. Determine the voiceprint to be recognized, wake-up word, and confidence value corresponding to the voice to be recognized.

[0034] The wake-up word is a pre-registered piece of voice and corresponds to a wake-up operation. For example, "Turn on music" corresponds to the wake-up operation of a playback device, and "Turn on the phone" corresponds to the wake-up operation of a communication device. The confidence value is used to represent the confidence level of the voice to be recognized hitting the wake-up word. It should be noted that in the application scenario of offline voice control, the confidence value is determined by comparing the physical attributes of the voice to be recognized with the physical attributes of the wake-up word, rather than determining the wake-up word and confidence value by semantic recognition of the voice to be recognized.

[0035] S103. Determine whether the voiceprint to be recognized is included in the preset voiceprint library.

[0036] By determining whether the voiceprint to be recognized is included in the preset voiceprint library, it is determined whether the voiceprint to be recognized has the permission to trigger the wake-up operation. The voiceprint to be recognized can be pre-registered and stored in the voiceprint library in advance; the voiceprint to be recognized can also obtain permission during use and be stored in the voiceprint library, and there is no need to obtain permission again during the next use.

[0037] S104. When the voiceprint to be recognized is included in the preset voiceprint library and the confidence value is greater than the first trigger threshold, trigger the wake-up operation according to the wake-up word.

[0038] If the voiceprint to be recognized is included in the preset voiceprint library, it means that the voiceprint to be recognized has obtained the permission to trigger the wake-up operation through pre-registration or other means. When the confidence value corresponding to the voice to be recognized is greater than the first trigger threshold, trigger the wake-up operation according to the wake-up word corresponding to the voice to be recognized. The first trigger threshold can be the confidence value corresponding to normal speech.

[0039] S105. When the voiceprint to be recognized is not included in the preset voiceprint library and the confidence value is greater than the second trigger threshold, trigger the wake-up operation according to the wake-up word. The first trigger threshold is less than the second trigger threshold.

[0040] If the voiceprint to be recognized is not included in the preset voiceprint library, it means that the voiceprint to be recognized has not obtained the permission to trigger the wake-up operation. At this time, the voice to be recognized needs to obtain permission first. The second trigger threshold can be set. The second trigger threshold is higher than the first trigger threshold for normal wake-up. Only when the confidence value of the voice to be recognized is very high can permission be obtained and the wake-up operation be triggered. The purpose of this is that when the user uses it for the first time, since the corresponding voiceprint and trigger threshold are not saved in the voiceprint library, a relatively high trigger threshold is required to ensure that the voices and voiceprints passing the threshold test are correct wake-up words, while ignoring the voices and voiceprints with relatively low confidence values, thereby reducing the false wake-up rate. The user can be prompted by voice to register and obtain permission by speaking clearly.

[0041] The voice control method based on voiceprint in this embodiment determines the voiceprint to be recognized, wake-up word, and confidence value corresponding to the voice to be recognized. When the voiceprint to be recognized is included in the preset voiceprint library, the confidence value only needs to be greater than a relatively low trigger threshold to trigger the wake-up operation. When the voiceprint to be recognized is not included in the preset voiceprint library, the confidence value needs to be greater than a relatively high trigger threshold to trigger the wake-up operation, effectively reducing the probability of misrecognition caused by ambient interference sounds, improving the accuracy and security of voice control, and enhancing the user experience.

[0042] Embodiment 2

[0043] As Figure 2 shown, the voice control method based on voiceprint in this embodiment is a further improvement of Embodiment 1. Specifically:

[0044] In an alternative embodiment, the voice control method further includes:

[0045] S106. When the voiceprint to be recognized is not included in the preset voiceprint library and the confidence value is greater than the second trigger threshold, after triggering the wake-up operation according to the wake-up word, the voiceprint to be recognized and the corresponding trigger threshold are also saved to the preset voiceprint library; the trigger threshold is less than the second trigger threshold.

[0046] When the confidence value of the voice to be recognized is greater than the second trigger threshold, the voice to be recognized has passed the threshold test for registration and obtained permissions. Therefore, after triggering the wake-up operation according to the wake-up word, the voiceprint to be recognized is saved. Since the second trigger threshold is a relatively high threshold set to filter out interfering sounds with low confidence values, when saving the voiceprint to be recognized, the threshold can be dynamically adjusted and a threshold lower than the second trigger threshold can be saved at the same time.

[0047] Specifically, the step of saving the corresponding trigger threshold includes:

[0048] The second trigger threshold is lowered by a preset first step value to obtain the corresponding trigger threshold. The first step value can be determined according to actual needs or usage conditions. Preferably, the first step value is 0.05.

[0049] In an alternative embodiment, after step S103, it further includes:

[0050] S201. Determine whether the confidence value is greater than the first trigger threshold.

[0051] S202. When the preset voiceprint library includes the voiceprint to be recognized and the confidence value is less than or equal to the first trigger threshold, lower the first trigger threshold stored in the preset voiceprint library. Since the first trigger threshold may be higher than the confidence value corresponding to normal speech of a person, resulting in the user being unable to trigger the wake-up operation in the normal speech manner and having to speak more clearly or louder to trigger the wake-up operation. Therefore, when the confidence value is less than or equal to the first trigger threshold, the threshold can be dynamically adjusted and the first trigger threshold stored in the preset voiceprint library can be lowered so that the user can trigger the wake-up operation in the normal speech manner.

[0052] Specifically, step S202 includes:

[0053] S2021. Lower the first trigger threshold by a preset second step value. The second step value can be determined according to actual needs or usage conditions. Preferably, the second step value is 0.05. The first step value and the second step value can be the same or different, and the present invention does not limit this.

[0054] S2022. Determine whether the lowered first trigger threshold is greater than the minimum threshold. The minimum threshold is the lowest threshold preset to trigger the wake-up operation. If the set first trigger threshold is lower than the minimum threshold, the reliability of recognition will be reduced. Therefore, the first trigger threshold cannot be lower than the minimum threshold.

[0055] S2023. If the lowered first trigger threshold is greater than the minimum threshold, save the lowered first trigger threshold.

[0056] S2024. If the lowered first trigger threshold is less than or equal to the minimum threshold, save the minimum threshold as the first trigger threshold.

[0057] In an alternative embodiment, the first trigger threshold and the second trigger threshold respectively have a corresponding relationship with the voiceprint to be recognized. That is, each voiceprint to be recognized has a corresponding first trigger threshold and second trigger threshold. Different first trigger thresholds and second trigger thresholds can be set according to the characteristics of the voiceprint to be recognized, or the first trigger threshold and the second trigger threshold can be adjusted according to actual usage conditions. However, the corresponding first trigger thresholds of different voiceprints to be recognized may be the same or different, and the present invention does not limit this; the corresponding second trigger thresholds of different voiceprints to be recognized may be the same or different, and the present invention does not limit this.

[0058] In addition, the first trigger threshold and the second trigger threshold are respectively in a corresponding relationship with the wake-up word. That is, each wake-up word has a corresponding first trigger threshold and a second trigger threshold. Different first trigger thresholds and second trigger thresholds can be set according to the characteristics of the wake-up word, or adjusted according to the actual usage situation. However, the first trigger thresholds corresponding to different wake-up words may be the same or different, and the present invention does not limit this. The second trigger thresholds corresponding to different wake-up words may be the same or different, and the present invention does not limit this either.

[0059] In an alternative embodiment, the first trigger threshold and the second trigger threshold are respectively negatively correlated with the length of the wake-up word. The longer the length of the wake-up word, the lower the probability of false wake-up, and accordingly, lower first trigger threshold and second trigger threshold can be set. The shorter the length of the wake-up word, the higher the probability of false wake-up, and accordingly, higher first trigger threshold and second trigger threshold can be set.

[0060] The voice control method based on voiceprint in this embodiment determines the voiceprint to be recognized, the wake-up word, and the confidence value corresponding to the voice to be recognized. When the voiceprint to be recognized is included in the preset voiceprint library, the confidence value only needs to be greater than the lower trigger threshold to trigger the wake-up operation. If the confidence value is still less than or equal to the lower trigger threshold, the trigger threshold is lowered to dynamically reduce the trigger threshold so that the registered voiceprint can easily trigger the wake-up operation, which is convenient for the user's regular use. When the voiceprint to be recognized is not included in the preset voiceprint library, the confidence value needs to be greater than the higher trigger threshold to trigger the wake-up operation, effectively reducing the probability of misrecognition caused by ambient interference sounds, improving the accuracy and security of voice control, and enhancing the user experience.

[0061] Embodiment 3

[0062] Please refer to Figure 3 , which is a schematic diagram of the modules of the voice control system based on voiceprint in this embodiment. Specifically, as Figure 3 shown, the voice control system includes: an acquisition module 1 for acquiring the voice to be recognized; a determination module 2 for determining the voiceprint to be recognized, the wake-up word, and the confidence value corresponding to the voice to be recognized; a trigger module 3 for triggering the wake-up operation according to the wake-up word when the voiceprint to be recognized is included in the preset voiceprint library and the confidence value is greater than the first trigger threshold. The trigger module 3 is also used to trigger the wake-up operation according to the wake-up word when the voiceprint to be recognized is not included in the preset voiceprint library and the confidence value is greater than the second trigger threshold. The first trigger threshold is less than the second trigger threshold.

[0063] Specifically, the wake-up word is a pre-registered piece of voice and corresponds to a wake-up operation. For example, "Turn on music" corresponds to the wake-up operation of a playback device, and "Turn on phone" corresponds to the wake-up operation of a communication device. The confidence value is used to represent the confidence level of the voice to be recognized hitting the wake-up word. It should be noted that in the application scenario of offline voice control, the confidence value is determined by comparing the physical attributes of the voice to be recognized with the physical attributes of the wake-up word, rather than by semantic recognition of the voice to be recognized and then determining the wake-up word and the confidence value.

[0064] It is determined whether the voiceprint to be recognized has the permission to trigger the wake-up operation by judging whether the preset voiceprint library includes the voiceprint to be recognized. The voiceprint to be recognized can be pre-registered and stored in the voiceprint library in advance; the voiceprint to be recognized can also obtain permission during use and be stored in the voiceprint library, and there is no need to obtain permission again during the next use.

[0065] If the preset voiceprint library includes the voiceprint to be recognized, it means that the voiceprint to be recognized has obtained the permission to trigger the wake-up operation through pre-registration or other means. When the confidence value corresponding to the voice to be recognized is greater than the first trigger threshold, the wake-up operation is triggered according to the wake-up word corresponding to the voice to be recognized. The first trigger threshold can be the confidence value corresponding to normal speech.

[0066] If the preset voiceprint library does not include the voiceprint to be recognized, it means that the voiceprint to be recognized has not obtained the permission to trigger the wake-up operation. At this time, the voice to be recognized needs to obtain permission first. A second trigger threshold can be set, and the second trigger threshold is higher than the first trigger threshold for normal wake-up. Only when the confidence value of the voice to be recognized is very high can permission be obtained and the wake-up operation be triggered. The purpose of this is that when the user uses it for the first time, since there is no corresponding voiceprint and corresponding trigger threshold saved in the voiceprint library, the trigger threshold needs to be relatively high to ensure that the voices and voiceprints passing the threshold test are correct wake-up words, while ignoring the voices and voiceprints with relatively low confidence values, thereby reducing the false wake-up rate. The user can be prompted by voice to register and obtain permission by speaking clearly.

[0067] The voice control system based on voiceprint in this embodiment determines the voiceprint to be recognized, the wake-up word, and the confidence value corresponding to the voice to be recognized. When the preset voiceprint library includes the voiceprint to be recognized, the wake-up operation can be triggered as long as the confidence value is greater than the lower trigger threshold. When the preset voiceprint library does not include the voiceprint to be recognized, the confidence value needs to be greater than the higher trigger threshold to trigger the wake-up operation, effectively reducing the probability of misrecognition caused by ambient interference sounds, improving the accuracy and security of voice control, and enhancing the user experience.

[0068] Embodiment 4

[0069] As Figure 4As shown, the voice control system based on voiceprint in this embodiment is a further improvement on Embodiment 3. Specifically:

[0070] In an alternative embodiment, the voice control system further includes: an adjustment module 4, which is configured to, when the voiceprint to be recognized is not included in the preset voiceprint library and the confidence value is greater than the second trigger threshold, after triggering the wake-up operation according to the wake-up word, save the voiceprint to be recognized and the corresponding trigger threshold to the preset voiceprint library; the trigger threshold is less than the second trigger threshold.

[0071] When the confidence value of the voice to be recognized is greater than the second trigger threshold, the voice to be recognized has passed the threshold test for registration and obtained the permission. Therefore, after triggering the wake-up operation according to the wake-up word, the voiceprint to be recognized is saved. Since the second trigger threshold is a relatively high threshold set to filter out interfering sounds with low confidence values, when saving the voiceprint to be recognized, the threshold can be dynamically adjusted and a threshold lower than the second trigger threshold can be saved at the same time.

[0072] Specifically, the adjustment module 4 is further configured to lower the second trigger threshold by a preset first step value to obtain the corresponding trigger threshold. The first step value can be determined according to actual needs or usage conditions. Preferably, the first step value is 0.05.

[0073] In an alternative embodiment, the adjustment module 4 is further configured to determine whether the confidence value is greater than the first trigger threshold.

[0074] The adjustment module 4 is further configured to, when the voiceprint to be recognized is included in the preset voiceprint library and the confidence value is less than or equal to the first trigger threshold, lower the first trigger threshold stored in the preset voiceprint library. Since the first trigger threshold may be higher than the confidence value corresponding to normal speech of a person, resulting in the user being unable to trigger the wake-up operation in the normal speech manner and having to speak more clearly or louder to trigger the wake-up operation, when the confidence value is less than or equal to the first trigger threshold, the threshold can be dynamically adjusted and the first trigger threshold stored in the preset voiceprint library can be lowered so that the user can trigger the wake-up operation in the normal speech manner.

[0075] Specifically, the adjustment module 4 is further configured to lower the first trigger threshold by a preset second step value. The second step value can be determined according to actual needs or usage conditions. Preferably, the second step value is 0.05. The first step value and the second step value can be the same or different, and the present invention does not limit this.

[0076] The adjustment module 4 is further configured to determine whether the adjusted first trigger threshold is greater than the minimum threshold. The minimum threshold is the lowest threshold preset to trigger the wake-up operation. If the set first trigger threshold is lower than the minimum threshold, the recognition reliability will be reduced. Therefore, the first trigger threshold cannot be lower than the minimum threshold.

[0077] If the first trigger threshold after being down-regulated is greater than the minimum threshold, the adjustment module 4 is further configured to save the down-regulated first trigger threshold.

[0078] If the first trigger threshold after being down-regulated is less than or equal to the minimum threshold, the adjustment module 4 is further configured to save the minimum threshold as the first trigger threshold.

[0079] In an alternative embodiment, the first trigger threshold and the second trigger threshold respectively have a corresponding relationship with the voiceprint to be recognized. That is, each voiceprint to be recognized has a corresponding first trigger threshold and a second trigger threshold. Different first trigger thresholds and second trigger thresholds can be set according to the characteristics of the voiceprint to be recognized, or the first trigger threshold and the second trigger threshold can be adjusted according to the actual usage situation. However, the corresponding first trigger thresholds of different voiceprints to be recognized may be the same or different, and the present invention does not limit this. The corresponding second trigger thresholds of different voiceprints to be recognized may be the same or different, and the present invention does not limit this.

[0080] In addition, the first trigger threshold and the second trigger threshold respectively have a corresponding relationship with the wake-up word. That is, each wake-up word has a corresponding first trigger threshold and a second trigger threshold. Different first trigger thresholds and second trigger thresholds can be set according to the characteristics of the wake-up word, or the first trigger threshold and the second trigger threshold can be adjusted according to the actual usage situation. However, the corresponding first trigger thresholds of different wake-up words may be the same or different, and the present invention does not limit this. The corresponding second trigger thresholds of different wake-up words may be the same or different, and the present invention does not limit this.

[0081] In an alternative embodiment, the first trigger threshold and the second trigger threshold are respectively negatively correlated with the length of the wake-up word. The longer the length of the wake-up word, the lower the probability of being misawakened, and accordingly, lower first trigger threshold and second trigger threshold can be set. The shorter the length of the wake-up word, the higher the probability of being misawakened, and accordingly, higher first trigger threshold and second trigger threshold can be set.

[0082] The voice control system based on voiceprint in this embodiment determines the voiceprint to be recognized, the wake-up word and the confidence value corresponding to the voice to be recognized. When the preset voiceprint library includes the voiceprint to be recognized, the confidence value only needs to be greater than the lower trigger threshold to trigger the wake-up operation. If the confidence value is still less than or equal to the lower trigger threshold, the trigger threshold is down-regulated to dynamically reduce the trigger threshold so that the registered voiceprint can easily trigger the wake-up operation, facilitating the regular use of the user. When the preset voiceprint library does not include the voiceprint to be recognized, the confidence value needs to be greater than the higher trigger threshold to trigger the wake-up operation, effectively reducing the probability of misrecognition caused by ambient interference sounds, improving the accuracy and security of voice control, and enhancing the user experience.

[0083] Example 5

[0084] Figure 5 FIG. is a schematic structural diagram of an electronic device provided in Example 5 of the present invention. The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the voice control method of Example 1 or Example 2 is implemented. Figure 5 The electronic device 30 shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0085] As Figure 5 shown, the electronic device 30 may be presented in the form of a general-purpose computing device, for example, it may be a server device. The components of the electronic device 30 may include, but are not limited to: at least one of the above-mentioned processors 31, at least one of the above-mentioned memories 32, and a bus 33 connecting different system components (including the memory 32 and the processor 31).

[0086] The bus 33 includes a data bus, an address bus, and a control bus.

[0087] The memory 32 may include volatile memory, such as a random access memory (RAM) 321 and / or a cache memory 322, and may further include a read-only memory (ROM) 323.

[0088] The memory 32 may further include a program / utilities 325 having a set (at least one) of program modules 324. Such program modules 324 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0089] The processor 31 executes various functional applications and data processing by running the computer program stored in the memory 32, such as the voice control method of Example 1 or Example 2 of the present invention.

[0090] The electronic device 30 can also communicate with one or more external devices 34 (such as a keyboard, a pointing device, etc.). Such communication can be carried out through the input / output (I / O) interface 35. Moreover, the model generating device 30 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN) and / or a public network, such as the Internet) through the network adapter 36. As shown in the figure, the network adapter 36 communicates with other modules of the model generating device 30 through the bus 33. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the model generating device 30, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (redundant array of independent disks) systems, tape drives, and data backup storage systems, etc.

[0091] It should be noted that, although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present invention, the features and functions of two or more of the above-described units / modules can be embodied in one unit / modules. Conversely, the features and functions of one unit / modules described above can be further divided and embodied by multiple units / modules.

[0092] Embodiment 6

[0093] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the voice control method of Embodiment 1 or Embodiment 2.

[0094] Among them, the more specific forms that the readable storage medium can adopt can include but not limited to: portable disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories, optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0095] In a possible implementation manner, the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to make the terminal device execute the voice control method of Embodiment 1 or Embodiment 2.

[0096] Among them, the program code for executing the present invention can be written in any combination of one or more programming languages. The program code can be completely executed on the user device, partially executed on the user device, executed as an independent software package, partially executed on the user device and partially executed on a remote device, or completely executed on a remote device.

[0097] Although the specific embodiments of the present invention have been described above, those skilled in the art should understand that this is only an example, and the protection scope of the present invention is defined by the appended claims. Without departing from the principles and essence of the present invention, those skilled in the art can make various changes or modifications to these embodiments, but these changes and modifications all fall within the protection scope of the present invention.

Claims

1. A voice control method based on voiceprint, characterized in that, the voice control method includes the following steps: Obtain the voice to be recognized; Determine the voiceprint to be recognized, wake-up word and confidence value corresponding to the voice to be recognized; When the preset voiceprint library includes the voiceprint to be recognized and the confidence value is greater than the first trigger threshold, trigger a wake-up operation according to the wake-up word; When the preset voiceprint library does not include the voiceprint to be recognized and the confidence value is greater than the second trigger threshold, trigger a wake-up operation according to the wake-up word; The first trigger threshold is less than the second trigger threshold.

2. The voice control method according to claim 1, the voice control method also includes: When the preset voiceprint library does not include the voiceprint to be recognized and the confidence value is greater than the second trigger threshold, after triggering a wake-up operation according to the wake-up word, also save the voiceprint to be recognized and the corresponding trigger threshold to the preset voiceprint library; The trigger threshold is less than the second trigger threshold.

3. The voice control method according to claim 2, the step of saving the corresponding trigger threshold includes: Lower the second trigger threshold by a preset first step value to obtain the corresponding trigger threshold.

4. The voice control method according to claim 1, the voice control method also includes: When the preset voiceprint library includes the voiceprint to be recognized and the confidence value is less than or equal to the first trigger threshold, lower the first trigger threshold stored in the preset voiceprint library.

5. The voice control method according to claim 4, the step of lowering the first trigger threshold stored in the preset voiceprint library includes: Lower the first trigger threshold by a preset second step value; If the lowered first trigger threshold is greater than the minimum threshold, save the lowered first trigger threshold; If the lowered first trigger threshold is less than or equal to the minimum threshold, save the minimum threshold as the first trigger threshold.

6. The voice control method according to claim 1, the first trigger threshold and the second trigger threshold respectively have a corresponding relationship with the voiceprint to be recognized.

7. The voice control method according to claim 1, the first trigger threshold and the second trigger threshold respectively have a corresponding relationship with the wake-up word.

8. The voice control method according to claim 7, the first trigger threshold and the second trigger threshold are respectively negatively correlated with the length of the wake-up word.

9. A voice control system based on voiceprint, characterized in that, the voice control system includes: An acquisition module for acquiring the voice to be recognized; A determination module for determining the voiceprint to be recognized, wake-up word and confidence value corresponding to the voice to be recognized; A trigger module for triggering a wake-up operation according to the wake-up word when the preset voiceprint library includes the voiceprint to be recognized and the confidence value is greater than the first trigger threshold; The trigger module is also used to trigger a wake-up operation according to the wake-up word when the preset voiceprint library does not include the voiceprint to be recognized and the confidence value is greater than the second trigger threshold; The first trigger threshold is less than the second trigger threshold.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the computer program, the voice control method according to any one of claims 1 to 8 is implemented.

11. A computer-readable storage medium, having stored thereon a computer program, wherein, when the computer program is executed by a processor, the voice control method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Audio frequency information display language setting device and method

    CN101101781A

  • Indoor security control method, electronic equipment, storage medium and system

    CN108538035A