Instruction word recognition method, electronic equipment and storage medium
By performing voice separation and voiceprint comparison after the device wakes up, combined with the judgment of the selection module, the problem of target voice extraction and refusal in the multi-speaker environment is solved, and a more accurate and fault-tolerant voice recognition effect is achieved.
Patent Information
- Application Number
- CN202510439745.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art is difficult to accurately extract the voice of the target object in a multi-speaker environment, and it is impossible to effectively refuse to recognize the voice noise, resulting in the possibility of non-target voice being output in advance.
By awakening the word voiceprint information, preparing the voiceprint model in advance, and then waking up the device to perform voice separation, obtaining the audio characteristics and recognition results of multiple channels, comparing the voiceprint information to obtain the voiceprint score, and determining whether each channel is the channel where the target speaker is located and whether it is necessary to refuse recognition.
It realizes accurate identification of the target speaker in a multi-speaker environment and effectively refuses to recognize human voice noise, improving the accuracy and fault tolerance of speech separation recognition.
Smart Images

Figure CN120126486A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of voice separation, and particularly relates to a method for identifying command words, an electronic device, and a storage medium. Background Art
[0002] In related technologies, the following method is usually adopted for voice separation. First, audio feature vectors in a voice stream are extracted, then the audio feature vectors are reconstructed to obtain a voice wave, and then the voice stream is separated to obtain voice frame vectors. Then, the voice frame vectors are operated on with the audio feature vectors to obtain source feature vectors. After that, the source feature vectors are multiplied by basis functions to obtain waveform signals of each voice frame. Then, target object voices corresponding to the voice waveform and the waveform signals of each voice frame are generated. Finally, the target object voices are recognized to obtain the voice texts of the target objects.
[0003] The inventors found that the related technologies have at least the following technical defects during the implementation of the present application. When there are multiple speakers and the speaker volumes are similar, it is very difficult to extract the target object. When the input audio contains only human voice noise, voice texts will also be output and rejection recognition cannot be performed. When continuously inputting voices, if the target human voice is at the back but there are also other human voices in the front, the voices of non-target people may be output in advance. Summary of the Invention
[0004] Embodiments of the present invention provide a method for identifying command words, an electronic device, and a storage medium, which are used to solve at least one of the above technical problems.
[0005] In a first aspect, embodiments of the present invention provide a method for identifying command words, including: in response to being awakened by a user, extracting wake-up word voiceprint information; performing voice separation on the audio information after awakening to obtain audios of multiple channels; obtaining multiple audio features and multiple recognition results of the audios of the multiple channels, where the multiple audio features at least include multiple signal-to-noise ratios and voiceprint information of multiple speakers; comparing the voiceprint information of the multiple speakers with the wake-up word voiceprint information respectively to obtain multiple voiceprint scores; and sending the multiple audio features, the multiple recognition results, and the multiple voiceprint scores to a selection module, where the selection module is used to output the probability that each channel after separation is the target speaker, and whether rejection recognition is required.
[0006] In a second aspect, embodiments of the present invention further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the method in the first aspect.
[0007] In a third aspect, an embodiment of the present invention further provides a storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0008] In the method of the embodiment of the present application, when the device is awakened by a wake-up word, voiceprint information is obtained by preparing a trained voiceprint model in advance, and then the device is awakened simultaneously. After the device is awakened, surrounding audio information is obtained, and then the obtained audio information is input into a voice separation model for voice separation, separating different audio information into different channels, and then processing the audio of each channel to obtain the audio features and audio recognition results corresponding to the audio information of each channel. Among them, the audio features include voiceprint features, signal-to-noise ratio, volume, etc., which are not limited in this application. Then, the voiceprint features in the audio features corresponding to the audio information of each channel are compared with the voiceprint features of the wake-up word to obtain the voiceprint score corresponding to the audio features of each channel. Then, the voiceprint features, recognition results, and voiceprint scores corresponding to each obtained channel are all sent into a selection module, and the selection module is used to judge the probability that each channel is the target speaker and whether rejection recognition is required, so as to accurately identify the target speaker and filter out aimless speech. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0010] Figure 1 It is a flowchart of a command word recognition method provided by an embodiment of the present invention; Figure 2 It is a flowchart of a specific implementation solution provided by an embodiment of the present invention; Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0011] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0012] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0013] The present invention may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention may also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media including storage devices.
[0014] In the present invention, terms such as "module", "device", "system", etc. refer to related entities applied to a computer, such as hardware, a combination of hardware and software, software, or software in execution. Specifically, for example, a component may be, but is not limited to, a process running on a processor, a processor, an object, an executable component, an execution thread, a program, and / or a computer. Also, an application program or a script program running on a server, and the server may both be components. One or more components may be in a process and / or thread of execution, and the components may be localized on one computer and / or distributed between two or more computers, and may be run by various computer-readable media. The components may also communicate through local and / or remote processes according to a signal having one or more data packets, for example, a signal from data that interacts with another component in a local system, a distributed system, and / or interacts with other systems through a network on the Internet.
[0015] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising" and "including" not only include those elements, but also include other elements not expressly listed, or also include elements inherent to such a process, method, article, or device. Without further limitation, an element defined by the statement "comprising..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.
[0016] An embodiment of the present invention provides a method for identifying instruction words, and this method can be applied to an electronic device. The electronic device may be a computer, a server, or other electronic products, etc., and the present invention does not make any limitation thereto.
[0017] Please refer to Figure 1 , which shows a method for identifying instruction words provided by an embodiment of the present invention.
[0018] As Figure 1 shown, in step 101, in response to being awakened by the user, the voiceprint information of the wake-up word is extracted; In step 102, the audio information after waking up is subjected to voice separation to obtain audio of multiple channels; In step 103, multiple audio features and multiple recognition results of the audio of the multiple channels are obtained, wherein the multiple audio features at least include multiple signal-to-noise ratios and voiceprint information of multiple speakers; In step 104, the voiceprint information of the multiple speakers is respectively compared with the voiceprint information of the wake-up word to obtain multiple voiceprint scores; In step 105, the multiple audio features, the multiple recognition results, and the multiple voiceprint scores are sent to a selection module, wherein the selection module is used to output the probability that each channel after separation is the target speaker, and whether rejection recognition is required.
[0019] In this embodiment, for step 101, in response to being awakened by the user, the voiceprint information of the wake-up word is extracted. For example, when the user wakes up the device, the user needs to say the corresponding wake-up word to wake up the device. After the device detects the wake-up word, the voiceprint information is obtained through a pre-prepared trained voiceprint model, and the device is awakened. Among them, the voiceprint model is a deep learning model, such as convolutional neural network (CNN), recurrent neural network (RNN), LSTM, Transformer model, etc., and this application has no limitation here.
[0020] After that, for step 102, the audio information after waking up is subjected to voice separation to obtain audio of multiple channels. For example, after waking up the device, the device automatically obtains the surrounding audio information, and then inputs the obtained audio information into a voice separation model for voice separation, separating different audio information into different channels, so as to obtain audio of multiple channels. For example, channel 1, channel 2... channel n, and the audio of each channel represents the audio of a possible speaker. Among them, the structure of the voice separation model generally adopts TCN (Temporal Convolutional Network), DPRNN (Dual-Path Recurrent Neural Network), TFPSNet (Time-Frequency Domain Path Scanning Network), etc., and this application has no limitation here.
[0021] Then, for step 103, obtain multiple audio features and multiple recognition results of the audio of the multiple channels, where the multiple audio features at least include multiple signal-to-noise ratios and voiceprint information of multiple speakers. For example, after obtaining the audio information of the multiple channels, process the audio information of each channel to obtain the audio features and audio recognition results corresponding to the audio information of each channel. The audio features include voiceprint features, signal-to-noise ratios, volume, etc., and this application has no limitation here.
[0022] After that, for step 104, compare the voiceprint information of the multiple speakers with the wake-up word voiceprint information respectively to obtain multiple voiceprint scores. For example, compare the voiceprint feature in the audio feature corresponding to each channel obtained in step 103 with the wake-up word voiceprint feature obtained in step 101, calculate the similarity of the two voiceprint features, and thus obtain the voiceprint score of the audio feature of each channel.
[0023] Finally, for step 105, send the multiple audio features, the multiple recognition results, and the multiple voiceprint scores into a selection module, where the selection module is used to output the probability that each channel after separation is the target speaker, and whether rejection is required. For example, send the obtained voiceprint features, recognition results, and voiceprint scores corresponding to each channel into the selection module, and the selection module is used to judge the probability that each channel is the target speaker and whether rejection is required.
[0024] In the method of the embodiment of this application, when the user wakes up the device through the wake-up word, obtain the voiceprint information by preparing a trained voiceprint model in advance, and then wake up the device simultaneously. After the device is woken up, obtain the surrounding audio information, then input the obtained audio information into the voice separation model for voice separation, separate different audio information into different channels, and then process the audio of each channel to obtain the audio features and audio recognition results corresponding to the audio information of each channel. The audio features include voiceprint features, signal-to-noise ratios, volume, etc., and this application has no limitation here. Then compare the voiceprint feature in the audio feature corresponding to the audio information of each channel with the wake-up word voiceprint feature to obtain the voiceprint score corresponding to the audio feature of each channel, and then send the obtained voiceprint features, recognition results, and voiceprint scores corresponding to each channel into the selection module, and the selection module is used to judge the probability that each channel is the target speaker and whether rejection is required, so as to accurately identify the target speaker and filter out the aimless words.
[0025] In some alternative embodiments, the selection module determines whether the current channel is the channel where the target speaker is located and determines whether rejection is required based on the multiple recognition results, the multiple audio features, and the multiple voiceprint scores. Among them, the multiple recognition results include multiple recognition texts and multiple recognition confidences. For example, when the selection module obtains the recognition results, audio features, and voiceprint scores corresponding to each channel, it makes a judgment based on the obtained recognition results, audio features, and voiceprint scores, so as to obtain the probability that each channel is the target speaker and whether rejection is required. Among them, the recognition results include recognition texts, recognition confidences, etc., and the present application has no limitation here. The recognition confidence is the model score of the recognition result, and the recognition confidence represents the reliability of the recognition text.
[0026] In some alternative embodiments, the selection module determines whether the current channel is the channel where the target speaker is located and determines whether rejection is required based on the multiple recognition results and the multiple audio features, including: detecting whether the multiple recognition texts of the multiple channels contain command words. If none of the multiple recognition texts of the multiple channels contain command words, direct rejection is performed. For example, it is determined whether to reject based on the recognition text corresponding to each channel. If none of the recognition texts of all channels contain command words, rejection is performed, so that there is no need to perform the next detection, reducing the consumption of the selection module.
[0027] In some alternative embodiments, the selection module determines whether the current channel is the channel where the target speaker is located and determines whether rejection is required based on the multiple recognition results and the multiple audio features, including: if at least one of the multiple recognition texts of the multiple channels contains a command word, then based on at least one audio feature, at least one voiceprint score, and at least one recognition confidence of the channel corresponding to the at least one recognition text, it is determined whether the current channel is the channel where the target speaker is located; if the current channel is the channel where the target speaker is located, the recognition text of the target speaker is output; if none of the channels corresponding to the at least one recognition text is the channel where the target speaker is located, rejection is performed and the subsequent audio information is continuously processed. For example, if it is detected that the recognition texts of several channels among the recognition texts of multiple channels contain command words, then the corresponding audio features, recognition confidences, and voiceprint scores are obtained from the channels where those recognition texts are located to determine whether the target speaker is in one of the channels. If one of the channels is the channel where the target speaker is located, the recognition text of the speaker of that channel is output. If none of them is the target speaker, rejection is performed, and then the audio information is obtained again for processing, so that the present application can accurately output the channel where the target speaker containing the command word is located.
[0028] In some alternative embodiments, the step of determining whether the current channel is the channel where the target speaker is located based on the multiple recognition results and the multiple audio features includes: calculating, by a selection module, a comprehensive score of the multiple audio features, the multiple voiceprint scores, and the multiple recognition confidences for each channel, where a certain channel with the highest comprehensive score is the channel where the target speaker is located. For example, the audio features, recognition confidences, and voiceprint scores of each channel are obtained, and the selection module calculates these information of the audio features, recognition confidences, and voiceprint scores of each channel to obtain the comprehensive score of each channel, and the channel with the highest comprehensive score is selected as the final target channel, so as to accurately determine the channel where the target speaker is located. The selection module is a neural network model, and the selection module calculates these information of each channel.
[0029] In some alternative embodiments, the step of inputting the multiple audio features, the multiple recognition results, and the multiple voiceprint scores into the selection module includes: determining, according to the multiple audio features, the multiple voiceprint scores, and the multiple recognition confidences of the multiple recognition results, whether the current channel is the channel where the target speaker is located, and determining whether an instruction word is included in the multiple recognition texts corresponding to the multiple channels; if the instruction word is not included in any of the multiple recognition texts, direct rejection is performed; if the instruction word is included in at least one of the recognition texts and the at least one channel corresponding to the at least one recognition text includes the channel where the target speaker is located, the recognition text of the target speaker is output. For example, after obtaining the audio features, voiceprint scores, recognition confidences, and recognition texts of the multiple channel audios, it is detected simultaneously whether the target user's channel and the recognition texts include the instruction word. If the instruction word is not included in any of the recognition texts of the multiple channel audios, direct rejection is performed. If the instruction word is included in the recognition text of a certain channel and the corresponding channel is the target speaker channel, the recognition text of the target speaker channel is output, so that the present application can accurately output the recognition text of the target speaker with the instruction word and the channel where the target speaker is located.
[0030] In some alternative embodiments, before separating the voice in the awakened audio information to obtain audio of multiple channels, the following steps are further included: detecting whether the awakened audio information contains valid human voices; if it contains valid human voices, then performing voice separation; if it does not contain valid human voices, then re-obtaining the awakened audio information for detection. For example, before sending the awakened audio information into a voice separation model, voice detection needs to be performed to determine whether the awakened audio information contains valid human voices through a VAD (Voice Activity Detection) model. If it contains valid human voices, then output it to the voice separation model; if it does not contain valid human voices, then re-obtain the awakened audio information to perform detection again. Through this method, it can be ensured that the audio information recognized in subsequent steps must have valid human voices.
[0031] In some alternative embodiments, the content output by the selection module includes the channel ID of the target speaker, the recognized text, and whether it is a rejection of recognition.
[0032] The inventor found that the above-mentioned related defects are caused by the following: after the voice signal is separated, the channel where the target human voice is located is random to a certain extent, and the time point when the target human voice appears is also random. In the case of multiple people speaking or there is human voice noise in the background, without other auxiliary information, it is very difficult to determine which is the target human voice, and this defect has long existed in the field of voice separation and recognition.
[0033] If these defects are to be solved, those skilled in the art usually add a signal processing module before voice separation or voice recognition. The main function of this module is to enhance the signal of the target speaker and suppress the signal of non-target speakers, so that there is an obvious signal-to-noise ratio difference between the target speaker and non-target speakers; however, this solution has a poor effect for multiple speakers in the same direction.
[0034] The application scenario targeted by this application is that the target person will say some directive words with a purpose, such as: turn on the air conditioner, tell a joke, what's the weather like today, etc.; if the target person says some words without a purpose, such as: one, two, three, four, I'm so happy, the baby is angry, etc., which are words without an obvious purpose, the accuracy of selecting the target person will decrease. When doing some projects where the microphone rotates (such as a sweeping robot, a robot, etc.), after the device is awakened, since the microphone rotates, the awakening direction and the recognition direction will be inconsistent; if signal processing is used to enhance the signal in the awakening direction, there is a high probability that the signal of non-target human voices will be enhanced, resulting in a very poor accuracy of sending the human voice for recognition.
[0035] In this context, an architecture for voice separation and recognition is proposed. The audio received by the microphone is directly fed into the voice separation module, which separates the audio into multiple channels. Each channel contains the audio signals of different speakers in the original audio. These audio signals are simultaneously fed into the voice recognition module. Then, information from multiple dimensions, such as the voiceprint information of the wake-up audio and the separated audio, the volume of the separated audio, the semantic information of the recognized text, and the confidence level of the recognition result, is fed into a selection module to output the channel of the target speaker, thereby obtaining the final recognition result.
[0036] Please refer to Figure 2 , which shows the flowchart of a specific implementation scheme of the present invention; As Figure 2 shown, Step 1: The user wakes up, and the voiceprint information of the wake-up word is extracted. Before extracting the voiceprint information, a trained voiceprint model needs to be prepared in advance. The voiceprint model uses deep learning models such as convolutional neural network (CNN), recurrent neural network (RNN), LSTM, and Transformer model. After the original wake-up word audio is input into the voiceprint model, the corresponding speaker voiceprint embedding vector will be output, and this vector represents the voiceprint information of the waking-up user.
[0037] Step 2: After the user wakes up, voice detection starts, generally using a VAD (Voice Activity Detection) model. The VAD model generally uses model structures such as DNN (Deep Neural Network) and FSMN (Feedforward Sequential Memory Networks) for training. The input is audio, and the output is a binary classification label (silent frame / voice frame) for each acoustic frame, so as to determine the start and end time points of the voice.
[0038] Step 3: After the voice is detected, the voice audio is input into the voice separation model. The model structure of voice separation generally uses TCN (Temporal Convolutional Network), DPRNN (Dual-Path Recurrent Neural Network), TFPSNet (Time-Frequency Domain Path Scanning Network), etc. The number of speakers is preset before training, and the number of output channels represents the preset number of speakers. The audio of each channel represents the audio of a possible speaker.
[0039] Step 4: Calculate the signal-to-noise ratio (SNR), speech recognition, and speaker verification for the multi-channel audio output from voice separation. Calculating the SNR generally involves calculating the root mean square amplitude of the audio and then converting it to decibels. The role of speech recognition is to convert the audio into text. Speech recognition models generally adopt an end-to-end architecture, such as Transformer, Comformer, etc. After converting the speech into recognized text, the model score of the recognition result can also be output as the confidence of the recognition (representing the reliability of the recognized text). The process of speaker verification is similar to that of wake-word speaker verification. The speaker verification model is the same as the model for extracting the wake-word speaker verification, except that the input becomes the audio after voice separation. Compare the speaker embedding vectors of each channel audio with the speaker embedding vector of the wake word and calculate the similarity of the speaker verification as the score of the speaker verification.
[0040] Step 5: Send the SNR, recognized text, recognition confidence, and speaker verification score obtained in Step 4 into the selection module respectively. The selection module can be fine-tuned on some pre-trained models, such as BERT (Bidirectional Encoder Representations from Transformers), ERNIE (Enhanced Representation through Knowledge Integration), etc. The main role of the selection module is to output the probability that each channel after separation is the target speaker and whether rejection is required. Generally, the higher the SNR, recognition confidence, and speaker verification score of the audio, the less likely it is to be rejected. However, if the recognized text is some aimless remarks, it may also be rejected.
[0041] Step 6: Obtain the recognized text of the most likely target speaker and whether rejection is required based on the output of the selection module. If rejection is required, it means that there may be no target speaker with purposeful command words in the current environment, and then it will return to Step 2 to continue monitoring the next segment of human voice. If rejection is not required, the recognized text of the target speaker will be output to end the entire process.
[0042] This application combines each module to make up for each other's deficiencies (voice separation can separate the target speaker, but it cannot know the location of the target speaker; speech recognition can convert speech into text, but it cannot indicate whether the current speech is the target speaker). By combining the information of different modules, more accurate target speaker text is obtained. At the same time, through the rejection function of the selection module, the system has a certain degree of fault tolerance (even if the currently input audio does not contain the target speaker, it can be restarted through rejection).
[0043] The inventor also tried the following technical solutions when implementing this application; 1. Have tried to use voice enhancement instead of voice separation. Voice enhancement only enhances the target human voice and suppresses all non-target human voices.
[0044] Advantages: There is no need to add a selection module. The audio after voice enhancement can be directly output to the recognition module, which is simpler in process and also reduces the errors that may be caused by the selection module.
[0045] Disadvantages: Lack of angle information for waking people. In scenarios with multiple people speaking or human voice noise interference, the enhancement effect of the target human voice is relatively poor, resulting in a relatively poor overall recognition rate.
[0046] 2. Combine the voice separation and voice recognition modules, input the original audio, and directly output the recognition result.
[0047] Advantages: Data preparation is relatively simple. Only one model needs to be optimized, avoiding the error accumulation caused by multiple models, and the joint training is more efficient.
[0048] Disadvantages: It is difficult to directly utilize the latest voice separation and recognition technologies. The integrated solution lacks interpretability and is relatively more difficult to debug and locate the cause of problems.
[0049] The inventors also used the following beta version when implementing this application: In the input of the selection module, semantic information of the recognized text and voiceprint information were not added, and only other information was used for selection; Advantages: The delay will become smaller and the logic is simpler. Disadvantages: When the position of the target human voice is behind the voice end, the risk of noise intrusion is increased.
[0050] The entire voice separation and recognition framework is the key innovation point, and each individual module can be implemented independently. The technical points to be protected are: how each module is combined to obtain a complete solution, what dimensions the information input to the selection module includes, and relatively accurate selection results are obtained.
[0051] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of actions combined. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention. In the above embodiments, each embodiment is described with its own emphasis. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0052] In some embodiments, the embodiments of the present invention provide a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any one of the above instruction word recognition methods of the present invention.
[0053] In some embodiments, the embodiments of the present invention further provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer is enabled to execute any one of the above instruction word recognition methods.
[0054] In some embodiments, the embodiments of the present invention further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the instruction word recognition method.
[0055] Figure 3 is a schematic hardware structure diagram of an electronic device for executing the instruction word recognition method provided by another embodiment of the present application. As Figure 3 shown, the device includes:
[0056] The device for executing the instruction word recognition method may further include: an input device 330 and an output device 340.
[0057] The processor 310, the memory 320, the input device 330, and the output device 340 may be connected by a bus or other means, Figure 3 and taking connection by bus as an example.
[0058] The memory 320, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the instruction word recognition method in the embodiments of the present application. The processor 310 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 320, that is, implements the instruction word recognition method in the above method embodiments.
[0059] The memory 320 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the voiceprint noise reduction device without registration, etc. In addition, the memory 320 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 320 may optionally include a memory remotely provided with respect to the processor 310, and these remote memories may be connected to the voiceprint noise reduction device without registration through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0060] The input device 330 may receive input digital or character information, and generate signals related to user settings and function control of the voiceprint noise reduction device without registration. The output device 340 may include a display device such as a display screen.
[0061] The one or more modules are stored in the memory 320, and when executed by the one or more processors 310, execute the instruction word recognition method in any of the above method embodiments.
[0062] The above product may execute the method provided in the embodiments of the present application, and has function modules and beneficial effects corresponding to the execution of the method. For technical details not described in detail in this embodiment, reference may be made to the method provided in the embodiments of the present application.
[0063] The electronic device in the embodiments of the present application exists in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.
[0064] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc.
[0065] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and intelligent toys and portable vehicle navigation devices.
[0066] (4) Other airborne electronic devices with data interaction functions, such as in-vehicle device installed on a vehicle.
[0067] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0068] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A method for recognizing instruction words, comprising: In response to being awakened by the user, extracting wake-up word voiceprint information; Perform voice separation on the audio information after wake-up to obtain audio of multiple channels; Acquire multiple audio features and multiple recognition results of the audio of the multiple channels, wherein the multiple audio features at least include multiple signal-to-noise ratios and voiceprint information of multiple speakers; Comparing the voiceprint information of the multiple speakers with the wake-up word voiceprint information respectively to obtain multiple voiceprint scores; The multiple audio features, the multiple recognition results and the multiple voiceprint scores are sent to a selection module, wherein the selection module is used to output the probability that each channel after separation is the target speaker, and whether rejection is required.
2. The method according to claim 1, characterized in that The selection module determines whether the current channel is the channel where the target speaker is located and whether rejection is required based on the multiple recognition results, the multiple audio features and the multiple voiceprint scores, wherein the multiple recognition results include multiple recognition texts and multiple recognition confidences.
3. The method according to claim 2, characterized in that The selection module determines whether the current channel is the channel where the target speaker is located and whether rejection is required according to the multiple recognition results and the multiple audio features, including: It is detected whether the multiple recognition texts of the multiple channels contain the instruction word, and if the multiple recognition texts of the multiple channels do not contain the instruction word, the recognition is directly rejected.
4. The method according to claim 2, characterized in that: The selection module determines whether the current channel is the channel where the target speaker is located and whether rejection is required according to the multiple recognition results and the multiple audio features, including: If at least one recognition text among the multiple recognition texts of the multiple channels contains the instruction word, judging whether the current channel is the channel where the target speaker is located according to at least one audio feature, at least one voiceprint score and at least one recognition confidence of the channel corresponding to the at least one recognition text; If the current channel is the channel where the target speaker is located, output the target speaker recognition text; If the channels corresponding to the at least one recognition text are not the channels where the target speaker is located, the recognition is rejected and the subsequent audio information is processed.
5. The method according to claim 2, characterized in that: The determining, according to the multiple recognition results and the multiple audio features, whether the current channel is the channel where the target speaker is located comprises: The selection module calculates a comprehensive score of the multiple audio features, the multiple voiceprint scores, and the multiple recognition confidences of each channel, wherein the channel with the highest comprehensive score is the channel where the target speaker is located.
6. The method according to claim 1, characterized in that The sending the plurality of audio features, the plurality of recognition results, and the plurality of voiceprint scores into a selection module comprises: According to the multiple audio features, the multiple voiceprint scores and the multiple recognition confidences of the multiple recognition results, determining whether the current channel is the channel where the target speaker is located, and determining whether the multiple recognition texts corresponding to the multiple channels contain instruction words; If none of the multiple recognition texts contain the instruction word, the recognition is directly rejected; If there is at least one recognition text containing the instruction word and at least one channel corresponding to the at least one recognition text includes the channel where the target speaker is located, the target speaker recognition text is output.
7. The method according to claim 1, characterized in that Before performing voice separation on the audio information after awakening to obtain audio of multiple channels, the method further includes: Detect whether the audio information after the wake-up contains a valid human voice. If it contains the valid human voice, perform voice separation. If it does not contain the valid human voice, re-acquire the audio information after the wake-up for detection.
8. The method according to claim 3, characterized in that The content output by the selection module includes the channel ID, recognition text and whether to reject the target speaker.
9. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method described in any one of claims 1 to 8.
10. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.