A speech recognition method, device, storage medium and electronic device
By integrating voice start detection and voiceprint recognition detection, using deep learning algorithms to judge the characteristic values of speech mediation, the problem of waste of speech recognition resources is solved and more efficient speech recognition is achieved.
Patent Information
- Application Number
- CN202111658530.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-12-31
AI Technical Summary
In the existing voice recognition technology, the independent design of voice control and voiceprint recognition functions leads to the waste of computing and storage resources, reducing the efficiency of voice recognition.
By integrating voice start detection and voiceprint recognition detection, deep learning algorithms are used to determine whether the voice intermediary feature value corresponds to stored registration commands and registered voice features, reducing the need for computing and storage resources.
It improves the efficiency of speech recognition and reduces the use of computing space resources and storage space resources.
Smart Images

Figure CN114267336B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition technology, and in particular, to a speech recognition method, device, storage medium, and electronic device.
Background Art
[0002] The voice control function focuses on voice features that are irrelevant to the user, while the voiceprint recognition function focuses on voice features related to the user. Generally, two functional modules are designed in series, and due to different focused eigenvalue, each also has its own front-end signal processing, including time-frequency conversion and eigenvalue extraction, which increases the computing space resources and storage space resources required for speech recognition and reduces the efficiency of speech recognition.
Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a speech recognition method, device, storage medium, and electronic device to improve the efficiency of speech recognition.
[0004] On the one hand, embodiments of the present invention provide a speech recognition method, including:
[0005] Receiving a voice command input by a user;
[0006] Extracting a voice intermediate eigenvalue of the voice command;
[0007] Judging whether the voice intermediate eigenvalue corresponds to a stored registered command through a deep learning algorithm;
[0008] If it is judged that the voice intermediate eigenvalue corresponds to a stored registered command, judging whether the voice intermediate eigenvalue corresponds to a stored registered voice feature through a deep learning algorithm;
[0009] If it is judged that the voice intermediate eigenvalue corresponds to a stored registered voice feature, recognizing the user's voice.
[0010] Optionally, before receiving the voice command input by the user, it includes:
[0011] Receiving multiple registered commands input by the user;
[0012] Extracting a registered intermediate eigenvalue of each registered command;
[0013] Judging whether the number of the multiple registered commands is equal to a set threshold;
[0014] If it is judged that the number of the multiple registered commands is equal to the set threshold, generating a registered voice feature according to the multiple registered intermediate eigenvalues;
[0015] Storing the registered voice feature.
[0016] Optionally, it further includes:
[0017] If it is determined that the voice mediation eigenvalue does not correspond to the stored registration command, continue to execute the step of receiving the voice command input by the user.
[0018] Optionally, it further includes:
[0019] If it is determined that the voice mediation eigenvalue does not correspond to the stored registered voice feature, continue to execute the step of receiving the voice command input by the user.
[0020] Optionally, it further includes:
[0021] If it is determined that the number of the multiple registration commands is less than a set threshold, store the multiple registration mediation eigenvalues, and continue to execute the step of receiving the multiple registration commands input by the user.
[0022] Optionally, the voice mediation eigenvalue includes a spectral eigenvalue, and the extracting the voice mediation eigenvalue of the voice command includes:
[0023] Extracting the spectral eigenvalue of the voice command;
[0024] Generating a first Mel spectral eigenvalue and a second Mel spectral eigenvalue according to the spectral eigenvalue.
[0025] Optionally, the determining whether the voice mediation eigenvalue corresponds to the stored registration command by a deep learning algorithm includes:
[0026] Determining whether the first Mel spectral eigenvalue corresponds to the stored registration command by a deep learning algorithm;
[0027] The determining whether the voice mediation eigenvalue corresponds to the stored registered voice feature by a deep learning algorithm includes:
[0028] Determining whether the second Mel spectral eigenvalue corresponds to the stored registered voice feature by a deep learning algorithm.
[0029] On the other hand, an embodiment of the present invention provides a voice recognition device, including:
[0030] A receiving module, configured to receive a voice command input by a user;
[0031] A first extraction module, configured to extract the voice mediation eigenvalue of the voice command;
[0032] A first judgment module, configured to determine whether the voice mediation eigenvalue corresponds to the stored registration command by a deep learning algorithm;
[0033] A second judgment module, configured to, if the first judgment module determines that the voice intermediary feature value corresponds to a stored registration command, judge whether the voice intermediary feature value corresponds to a stored registered voice feature through a deep learning algorithm; if it is determined that the voice intermediary feature value corresponds to the stored registered voice feature, the voice of the user is recognized.
[0034] On the other hand, an embodiment of the present invention provides a storage medium, including: the storage medium includes a stored program, wherein, when the program runs, it controls the device where the storage medium is located to execute the above-mentioned voice recognition method.
[0035] On the other hand, an embodiment of the present invention provides an electronic device, including a memory and a processor, the memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions, characterized in that, when the program instructions are loaded and executed by the processor, the steps of the above-mentioned voice recognition method are implemented.
[0036] In the technical solution of the voice recognition method provided by the embodiment of the present invention, a voice command input by a user is received; a voice intermediary feature value of the voice command is extracted; it is judged through a deep learning algorithm whether the voice intermediary feature value corresponds to a stored registration command; if it is determined that the voice intermediary feature value corresponds to the stored registration command, it is judged through a deep learning algorithm whether the voice intermediary feature value corresponds to a stored registered voice feature; if it is determined that the voice intermediary feature value corresponds to the stored registered voice feature, the voice of the user is recognized. In the technical solution provided by the embodiment of the present invention, voice activation detection and voiceprint recognition detection are integrated, reducing the computational space resources and storage space resources required for voice recognition, and improving the efficiency of voice recognition.
Description of the Drawings
[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0038] Figure 1 It is a flowchart of a voice recognition method provided by an embodiment of the present invention;
[0039] Figure 2 It is a flowchart of another voice recognition method provided by an embodiment of the present invention;
[0040] Figure 3 It is a schematic diagram of a voice intermediary feature value;
[0041] Figure 4Schematic structural diagram of a voice recognition device provided by an embodiment of the present invention;
[0042] Figure 5 Schematic diagram of an electronic device provided by an embodiment of the present invention;
[0043] Figure 6 For Figure 5 Schematic diagram of the processor in
Specific embodiments
[0044] For a better understanding of the technical solution of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0045] It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0046] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms of "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0047] It should be understood that the term " / and" used herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0048] The embodiments of the present invention provide a voice recognition method, Figure 1 For the flowchart of a voice recognition method provided by an embodiment of the present invention, as Figure 1 shown, the method includes:
[0049] Step 102, receiving a voice command input by the user.
[0050] In the embodiments of the present invention, each step is executed by an electronic device. For example, the electronic device includes: a mobile phone, a tablet computer, a voice-controlled speaker or a voice-controlled household appliance.
[0051] As an alternative solution, the user inputs a voice command to the electronic device by speaking the voice command into the earpiece of the electronic device, and the electronic device receives the voice command input by the user.
[0052] Step 104, extracting the voice intermediate feature value of the voice command.
[0053] Step 106: Determine whether the voice intermediate eigenvalue corresponds to the stored registration command through a deep learning algorithm.
[0054] In this step, the electronic device performs voice activation detection and determines whether the voice intermediate eigenvalue corresponds to the stored registration command through a deep learning algorithm.
[0055] In an embodiment of the present invention, the electronic device stores the corresponding relationship between the voice intermediate eigenvalue and the registration command.
[0056] In this step, if it is determined that the voice intermediate eigenvalue corresponds to the stored registration command, it indicates that the voice command is a registration voice; if it is determined that the voice intermediate eigenvalue does not correspond to the stored registration command, it indicates that the voice command is not a registration voice.
[0057] Step 108: If it is determined that the voice intermediate eigenvalue corresponds to the stored registration command, determine whether the voice intermediate eigenvalue corresponds to the stored registered voice feature through a deep learning algorithm.
[0058] In this step, the electronic device performs voiceprint recognition detection and determines whether the voice intermediate eigenvalue corresponds to the stored registered voice feature through a deep learning algorithm.
[0059] In an embodiment of the present invention, the electronic device stores the corresponding relationship between the voice intermediate eigenvalue and the registered voice feature.
[0060] In this step, if it is determined that the voice intermediate eigenvalue corresponds to the stored registered voice feature, it indicates that the user corresponding to the voice command is a registered user; if it is determined that the voice intermediate eigenvalue does not correspond to the stored registered voice feature, it indicates that the user corresponding to the voice command is not a registered user.
[0061] Step 110: If it is determined that the voice intermediate eigenvalue corresponds to the stored registered voice feature, recognize the user's voice.
[0062] In this step, if it is determined that the voice intermediate eigenvalue corresponds to the stored registered voice feature, the user's voice is recognized, and the user is a registered user.
[0063] In the technical solution provided by the embodiment of the present invention, a voice command input by a user is received; a voice intermediate feature value of the voice command is extracted; it is determined by a deep learning algorithm whether the voice intermediate feature value corresponds to a stored registered command; if it is determined that the voice intermediate feature value corresponds to the stored registered command, it is determined by the deep learning algorithm whether the voice intermediate feature value corresponds to a stored registered voice feature; if it is determined that the voice intermediate feature value corresponds to the stored registered voice feature, then the user's voice is recognized. In the technical solution provided by the embodiment of the present invention, voice activation detection and voiceprint recognition detection are integrated, reducing the computing space resources and storage space resources required for voice recognition and improving the efficiency of voice recognition.
[0064] The embodiment of the present invention provides another voice recognition method. Figure 2 It is a flowchart of another voice recognition method provided by the embodiment of the present invention. As Figure 2 shown, the method includes:
[0065] Step 202, receive multiple registered commands input by the user.
[0066] In the embodiment of the present invention, each step is executed by an electronic device. For example, the electronic device includes: a mobile phone, a tablet computer, a voice-controlled speaker, or a voice-controlled household appliance.
[0067] As an optional solution, the user inputs a registered command to the electronic device by speaking the registered command into the receiver of the electronic device. The electronic device receives the registered command input by the user. The user is a registered user. To ensure the accuracy of extracting the registered intermediate feature value subsequently, the user needs to input multiple registered commands.
[0068] Step 204, extract the registered intermediate feature value of each registered command.
[0069] In this step, the registered intermediate feature value of each registered command is extracted to integrate the same registered intermediate feature value to identify the registered user.
[0070] Specifically, each registered command is subjected to a frame cutting operation and time-frequency conversion to extract the registered intermediate feature value of each registered command.
[0071] Step 206, determine whether the number of multiple registered commands is equal to a set threshold. If so, execute step 208; if not, execute step 222.
[0072] In the embodiment of the present invention, the set threshold can be set according to the actual situation. For example, the set threshold is 5.
[0073] In an embodiment of the present invention, if it is determined that the number of multiple registration commands is equal to a set threshold, it indicates that sufficient registration commands have been collected, and step 208 is executed; if it is determined that the number of multiple registration commands is less than the set threshold, it indicates that insufficient registration commands have been collected, and step 222 is executed.
[0074] Step 208: Generate a registration voice feature based on multiple registration intermediary feature values.
[0075] In this step, the same registration intermediary feature values are integrated to generate a registration voice feature for identifying the registered user.
[0076] Step 210: Store the registration voice feature.
[0077] Step 212: Receive a voice command input by the user.
[0078] As an optional solution, the user inputs a voice command to the electronic device by speaking the voice command into the earpiece of the electronic device, and the electronic device receives the voice command input by the user.
[0079] Step 214: Extract the voice intermediary feature value of the voice command.
[0080] In an embodiment of the present invention, the voice intermediary feature value includes a spectral feature value.
[0081] Specifically, frame cutting operation is performed on the voice command and time-frequency conversion is performed to extract the voice intermediary feature value of the voice command.
[0082] Specifically, step 214 includes: extracting the spectral feature value of the voice command; generating a first Mel spectral feature value and a second Mel spectral feature value based on the spectral feature value.
[0083] In an embodiment of the present invention, Figure 3 is a schematic diagram of the voice intermediary feature value. As Figure 3 shown, the voice intermediary feature value will generate feature values with different resolutions. Taking the spectral feature value as an example, the spectral feature value contains twice the frequency resolution and time resolution. Through the training of a deep learning algorithm, a first Mel spectral feature value and a second Mel spectral feature value are generated based on the spectral feature value.
[0084] As an optional solution, the first Mel spectral feature value contains twice the frequency resolution and once the time resolution, and the first Mel spectral feature value can be used for voice activation detection in step 216; the second Mel spectral feature value contains once the frequency resolution and twice the time resolution, and the second Mel spectral feature value can be used for speaker recognition detection in step 218.
[0085] Step 216: Determine, through a deep learning algorithm, whether the voice intermediate eigenvalue corresponds to a stored registration command. If so, execute Step 218; if not, execute Step 212.
[0086] Specifically, determine, through a deep learning algorithm, whether the first Mel spectrogram eigenvalue corresponds to a stored registration command.
[0087] In an embodiment of the present invention, the electronic device stores the correspondence between the first Mel spectrogram eigenvalue and the registration command.
[0088] In this step, if it is determined that the first Mel spectrogram eigenvalue corresponds to a stored registration command, it indicates that the voice command is a registration voice, and Step 218 is executed; if it is determined that the first Mel spectrogram eigenvalue does not correspond to a stored registration command, it indicates that the voice command is not a registration voice, and Step 212 is executed.
[0089] Step 218: Determine, through a deep learning algorithm, whether the voice intermediate eigenvalue corresponds to a stored registered voice feature. If so, execute Step 220; if not, execute Step 212.
[0090] Specifically, determine, through a deep learning algorithm, whether the second Mel spectrogram eigenvalue corresponds to a stored registered voice feature.
[0091] In an embodiment of the present invention, the electronic device stores the correspondence between the second Mel spectrogram eigenvalue and the registered voice feature.
[0092] In this step, if it is determined that the second Mel spectrogram eigenvalue corresponds to a stored registered voice feature, it indicates that the user corresponding to the voice command is a registered user, and Step 220 is executed; if it is determined that the second Mel spectrogram eigenvalue does not correspond to a stored registered voice feature, it indicates that the user corresponding to the voice command is not a registered user, and Step 212 is executed.
[0093] Step 220: Recognize the user's voice, and the process ends.
[0094] In this step, the electronic device recognizes the user's voice, and the user is a registered user.
[0095] Step 222: Store multiple registered intermediate eigenvalues, and continue to execute Step 202.
[0096] In the technical solution provided by the embodiment of the present invention, a voice command input by a user is received; a voice intermediate feature value of the voice command is extracted; whether the voice intermediate feature value corresponds to a stored registered command is judged by a deep learning algorithm; if it is judged that the voice intermediate feature value corresponds to the stored registered command, whether the voice intermediate feature value corresponds to a stored registered voice feature is judged by the deep learning algorithm; if it is judged that the voice intermediate feature value corresponds to the stored registered voice feature, the user's voice is recognized. In the technical solution provided by the embodiment of the present invention, voice activation detection and voiceprint recognition detection are integrated, reducing the computational space resources and storage space resources required for voice recognition and improving the efficiency of voice recognition.
[0097] In the technical solution provided by the embodiment of the present invention, on most electronic devices, under the condition of limited computing resources and storage resources, on the premise of ensuring the functional effects of both, the deep learning algorithm module for voice activation detection and the deep learning algorithm module for voiceprint recognition detection are systematically integrated.
[0098] The embodiment of the present invention provides a voice recognition device. Figure 4 As a structural schematic diagram of a voice recognition device provided by the embodiment of the present invention, as Figure 4 shown, the device includes: a receiving module 11, a first extraction module 12, a first judgment module 13, and a second judgment module 14.
[0099] The receiving module 11 is used to receive a voice command input by a user.
[0100] The first extraction module 12 is used to extract the voice intermediate feature value of the voice command.
[0101] The first judgment module 13 is used to judge whether the voice intermediate feature value corresponds to a stored registered command by a deep learning algorithm.
[0102] The second judgment module 14 is used, if the first judgment module 13 judges that the voice intermediate feature value corresponds to a stored registered command, to judge whether the voice intermediate feature value corresponds to a stored registered voice feature by a deep learning algorithm; if it is judged that the voice intermediate feature value corresponds to the stored registered voice feature, the user's voice is recognized.
[0103] In the embodiment of the present invention, the device further includes: a second extraction module 15, a third judgment module 16, a generation module 17, and a first storage module 18.
[0104] The receiving module 11 is further used to receive a plurality of registered commands input by a user.
[0105] The second extraction module 15 is used to extract the registration intermediate feature value of each registered command.
[0106] The third judgment module 16 is used to judge whether the number of the multiple registration commands is equal to a set threshold; if it is judged that the number of the multiple registration commands is equal to the set threshold, the generation module 17 is triggered to generate a registration voice feature according to the multiple registration intermediary feature values.
[0107] The first storage module 18 is used to store the registration voice feature.
[0108] In an embodiment of the present invention, if the first judgment module 13 judges that the voice intermediary feature value does not correspond to a stored registration command, the receiving module 11 is triggered to continue to execute the step of receiving the voice command input by the user.
[0109] In an embodiment of the present invention, if the second judgment module 14 judges that the voice intermediary feature value does not correspond to a stored registration voice feature, the receiving module 11 is triggered to continue to execute the step of receiving the voice command input by the user.
[0110] In an embodiment of the present invention, the device further includes: a second storage module 19.
[0111] If the third judgment module 16 judges that the number of the multiple registration commands is less than the set threshold, the second storage module 19 is triggered to store the multiple registration intermediary feature values, and the receiving module 11 is triggered to continue to execute the step of receiving the multiple registration commands input by the user.
[0112] In an embodiment of the present invention, the voice intermediary feature value includes a spectral feature value, and the first extraction module 12 is specifically used to extract the spectral feature value of the voice command; a first Mel spectral feature value and a second Mel spectral feature value are generated according to the spectral feature value.
[0113] In an embodiment of the present invention, the first judgment module 13 is specifically used to judge whether the first Mel spectral feature value corresponds to a stored registration command through a deep learning algorithm.
[0114] In an embodiment of the present invention, the second judgment module 14 is specifically used to judge whether the second Mel spectral feature value corresponds to a stored registration voice feature through a deep learning algorithm.
[0115] In the technical solution provided by the embodiment of the present invention, a voice command input by a user is received; a voice intermediate feature value of the voice command is extracted; whether the voice intermediate feature value corresponds to a stored registered command is determined through a deep learning algorithm; if it is determined that the voice intermediate feature value corresponds to the stored registered command, whether the voice intermediate feature value corresponds to the stored registered voice feature is determined through the deep learning algorithm; if it is determined that the voice intermediate feature value corresponds to the stored registered voice feature, the user's voice is recognized. In the technical solution provided by the embodiment of the present invention, voice activation detection and voiceprint recognition detection are integrated, reducing the computational space resources and storage space resources required for voice recognition and improving the efficiency of voice recognition.
[0116] The voice recognition device provided in this embodiment can be used to implement the above Figure 1 and Figure 2 voice recognition method. For specific descriptions, reference can be made to the embodiments of the above voice recognition method, which will not be repeated here.
[0117] The embodiment of the present invention provides a storage medium. The storage medium includes a stored program. When the program runs, it controls the device where the storage medium is located to execute the steps of the above embodiment of the voice recognition method. For specific descriptions, reference can be made to the embodiments of the above voice recognition method.
[0118] The embodiment of the present invention provides an electronic device, including a memory and a processor. The memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the program instructions are loaded and executed by the processor, the steps of the above embodiment of the voice recognition method are implemented. For specific descriptions, reference can be made to the embodiments of the above voice recognition method.
[0119] Figure 5 is a schematic diagram of an electronic device provided by an embodiment of the present invention. As Figure 5 shown, the electronic device 20 of this embodiment includes: a processor 21, a memory 22, and a computer program 23 stored in the memory 22 and executable on the processor 21. When the computer program 23 is executed by the processor 21, it implements the application to the voice recognition method in the embodiment. To avoid repetition, it will not be elaborated here one by one. Alternatively, when the computer program is executed by the processor 21, it implements the functions of each model / unit in the voice recognition device in the embodiment. To avoid repetition, it will not be elaborated here one by one.
[0120] The electronic device 20 includes, but is not limited to, a processor 21 and a memory 22. Those skilled in the art can understand that Figure 5 is only an example of the electronic device 20 and does not constitute a limitation on the electronic device 20. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.
[0121] The so-called processor 21 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0122] The memory 22 may be an internal storage unit of the electronic device 20, such as the hard disk or memory of the electronic device 20. The memory 22 may also be an external storage device of the electronic device 20, such as a plug-in hard disk equipped on the electronic device 20, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 22 may also include both the internal storage unit and the external storage device of the electronic device 20. The memory 22 is used to store computer programs and other programs and data required by the electronic device. The memory 22 may also be used to temporarily store data that has been output or is to be output.
[0123] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0124] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other may be through some interfaces, and the indirect couplings or communication connections of the devices or units may be in electrical, mechanical, or other forms.
[0125] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0126] In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing unit, can exist separately as individual physical units, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.
[0127] The above integrated units implemented in the form of software functional units can be stored in a computer-readable storage medium. The above software functional units stored in a storage medium include several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0128] In an embodiment of the present invention, Figure 6 For Figure 5 a schematic diagram of the processor in Figure 6 as shown, the processor 21 includes: a central processing unit (CPU) 211, an analog-to-digital converter (ADC) 212, a digital signal processor (DSP) 213, and a deep algorithm accelerator 214. Among them, the deep algorithm accelerator 214 includes: a graphics processing unit (GPU) or a neural network unit (NN unit).
[0129] In an embodiment of the present invention, when the processor 21 is in the sleep state, it is in the low-frequency power mode. The CPU 211 does not work, and only a small range of hardware modules (ADC 212, DSP 213, or deep algorithm accelerator 214) can work. After the sound signal is received by the microphone and converted by the ADC 212, it enters the DSP 213 for processing. When the speech intermediate feature value of the speech command is extracted, it is transferred to the deep algorithm accelerator 214 for processing. At this time, the deep learning algorithm we adopted can be accelerated by the deep algorithm accelerator 214 to obtain the best performance. Finally, the judgment signal is returned to the DSP 213 or the CPU 211 to perform the above-mentioned processes such as Figure 1 or Figure 2 .
[0130] In an embodiment of the present invention, the ADC 212, DSP 213, deep algorithm accelerator 214, and CPU 211 can be different chip modules or integrated into a single chip.
[0131] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A voice recognition method, characterized in that, including: receiving a voice command input by a user; extracting spectral feature values of the voice command; generating a first Mel spectral feature value and a second Mel spectral feature value according to the spectral feature values; judging whether the first Mel spectral feature value corresponds to a stored registered command through a deep learning algorithm; if it is judged that the first Mel spectral feature value corresponds to a stored registered command, judging whether the second Mel spectral feature value corresponds to a stored registered voice feature through a deep learning algorithm; if it is judged that the second Mel spectral feature value corresponds to a stored registered voice feature, identifying the user's voice.
2. The method according to claim 1, characterized in that, Before the receiving of the voice command input by the user, it includes: receiving a plurality of registered commands input by the user; extracting registered intermediate feature values of each of the registered commands; judging whether the number of the plurality of registered commands is equal to a set threshold; if it is judged that the number of the plurality of registered commands is equal to the set threshold, generating a registered voice feature according to the plurality of registered intermediate feature values; storing the registered voice feature.
3. The method according to claim 1, characterized in that, It further includes: if it is judged that the first Mel spectral feature value does not correspond to a stored registered command, continuing to execute the step of receiving the voice command input by the user.
4. The method according to claim 1, wherein It further includes: if it is judged that the second Mel spectral feature value does not correspond to a stored registered voice feature, continuing to execute the step of receiving the voice command input by the user.
5. The method according to claim 2, characterized in that, It further includes: if it is judged that the number of the plurality of registered commands is less than the set threshold, storing the plurality of registered intermediate feature values and continuing to execute the step of receiving the plurality of registered commands input by the user.
6. A voice recognition device, characterized in that, including: a receiving module, configured to receive a voice command input by a user; a first extraction module, configured to extract spectral feature values of the voice command; generating a first Mel spectral feature value and a second Mel spectral feature value according to the spectral feature values; a first judgment module, configured to judge whether the first Mel spectral feature value corresponds to a stored registered command through a deep learning algorithm; a second judgment module, configured to, if the first judgment module judges that the first Mel spectral feature value corresponds to a stored registered command, judge whether the second Mel spectral feature value corresponds to a stored registered voice feature through a deep learning algorithm; if it is judged that the second Mel spectral feature value corresponds to a stored registered voice feature, identifying the user's voice.
7. A storage medium, characterized in that, including: The storage medium includes a stored program, wherein when the program runs, it controls the device where the storage medium is located to execute a voice recognition method according to any one of claims 1 to 5.
8. An electronic device, comprising a memory and a processor, the memory being used for storing information including program instructions, and the processor being used for controlling the execution of the program instructions, characterized in that, When the program instructions are loaded and executed by a processor, the steps of a voice recognition method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Explicit voiceprint identification method and device
CN107481720A
Voice interaction method, apparatus and device, and computer readable storage medium
CN109272991A