Method, apparatus, storage medium and electronic device for outputting command words

After detecting the target wake-up word, the probability of occurrence of the command word is determined based on subsequent audio data and output it, the problems of low efficiency and high power consumption in the prior art are solved, and fast response and low power consumption are achieved.

CN114242062BActive Publication Date: 2025-07-04ZHEJIANG DAHUA TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111645679.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-07-04
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

In the prior art, command word recognition requires identifying all the words included in the input voice before determining whether the command word is included, resulting in low recognition efficiency and high power consumption.

Method used

After detecting the target wake-up word, the probability of occurrence of the command word is determined based on the subsequent received audio data, and the command word is output immediately when the probability is greater than the threshold value, so as to avoid waiting for all words to be recognized.

Benefits of technology

It improves the recognition efficiency of command words, reduces recognition power consumption, and realizes the rapid response of command words.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114242062B_ABST
    Figure CN114242062B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a method, device, storage medium and electronic device for outputting command words. The method includes: detecting the type of currently received audio data when continuously receiving audio data; in response to detecting that the currently received audio data is target audio data corresponding to a target wake-up word, determining the occurrence probability of audio data corresponding to a command word that appears subsequently based on the audio data received after the target audio data; and outputting the target command word in response to determining that the audio data corresponding to the target command word has an occurrence probability greater than a first probability threshold. Through the present invention, the problem in the related art that it is necessary to recognize all the words included in the input speech before it is possible to determine whether the input speech includes a command word, resulting in low recognition efficiency and high recognition power consumption of the command word, is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] An embodiment of the present invention relates to the field of communications, and in particular, to a method, apparatus, storage medium, and electronic device for outputting command words. Background Art

[0002] Command word recognition is a special scenario of speech recognition and can be applied to intelligent control scenarios such as homes and meetings. Generally, there is only one keyword for wake word detection. For example, "Xiaobai", etc. However, there may be multiple keywords for command word recognition. For example, in a home scenario, there are sentences like "Turn on the air conditioner" and "Turn on the TV", etc., all of which are short imperative sentences with clear operation targets. Since command word recognition is a type of speech recognition, the basic algorithms of speech recognition can also be applied to command word detection algorithms.

[0003] In the related art, traditional command word recognition algorithms can be used alone or as a post - algorithm of a wake word algorithm (common wake word algorithms include conventional speech recognition methods and end - to - end speech recognition methods. Additionally, in the case of deep optimization, the two can achieve similar effects). When used as a post - algorithm, the recognition result of the wake word can be used to further improve the recognition effect of the command word. Without the need to perform the vad (Voice Activity Detection) algorithm, the command word can be output (that is, waiting for the vad algorithm to give a judgment on the end of the input speech, and the command word algorithm outputs all recognition results based on this judgment and determines whether there is a command word), thus being able to more accurately judge the end of the command word. Subsequently, the delay of the command word output result caused by performing the vad algorithm can be reduced. Among them, when multiple keywords need to be recognized for the command word, it is necessary to wait for all the recognition results to come out before determining whether the recognition result is a command word (that is, whether the command word is triggered), which leads to an increase in the power consumption of the command word recognition algorithm, and when the non - command word voice input is too long, there is a problem of an increase in the false alarm rate of the command word recognition algorithm.

[0004] Aiming at the problem in the related art that it is necessary to recognize all the words included in the input speech before it is possible to determine whether the input speech includes a command word, resulting in low recognition efficiency and high recognition power consumption of the command word, no effective solution has been proposed for this problem yet. Summary of the Invention

[0005] An embodiment of the present invention provides a method, apparatus, storage medium, and electronic device for outputting command words to at least solve the problems of low recognition efficiency and high recognition power consumption of command words in the related art.

[0006] According to an embodiment of the present invention, a method for outputting a command word is provided, including: detecting the type of currently received audio data when continuously receiving audio data; in response to detecting that the currently received audio data is target audio data corresponding to a target wake-up word, determining the occurrence probability of audio data corresponding to a command word that appears subsequently based on the audio data received after the target audio data; in response to determining that the occurrence probability of the audio data corresponding to the target command word is greater than a first probability threshold, outputting the target command word.

[0007] In an exemplary embodiment, determining the occurrence probability of audio data corresponding to a command word that appears subsequently based on the audio data received after the target audio data includes: determining a first probability of audio data corresponding to a command word type that appears subsequently and a second probability of audio data corresponding to a non-command word type based on the audio data received after the target audio data; in response to determining that the first probability is greater than a second probability threshold, determining the occurrence probability of the audio data corresponding to each command word based on the subsequently received audio data.

[0008] In an exemplary embodiment, determining a first probability of audio data corresponding to a command word type that appears subsequently and a second probability of audio data corresponding to a non-command word type based on the audio data received after the target audio data includes: adjusting a first weight of audio data corresponding to a command word type that appears subsequently and a second weight of audio data corresponding to a non-command word type in a target decoding graph based on the audio data received after the target audio data; determining the first probability and the second probability based on the first weight and the second weight.

[0009] In an exemplary embodiment, adjusting a first weight of audio data corresponding to a command word type that appears subsequently and a second weight of audio data corresponding to a non-command word type based on the audio data received after the target audio data includes: performing frame-level decoding on the audio data received after the target audio data to obtain a first decoding result; continuously adjusting a first initial weight of a command word path and a second initial weight of a non-command word path included in the target decoding graph based on the first decoding result; determining the adjusted first initial weight as the first weight, and determining the adjusted second initial weight as the second weight.

[0010] In an exemplary embodiment, determining the occurrence probability of the audio data corresponding to each command word based on the subsequently received audio data includes: performing frame-level decoding on the subsequently received audio data to obtain a second decoding result; continuously adjusting the initial weight corresponding to each command word path included in the target decoding graph based on the second decoding result; and determining the adjusted initial weight corresponding to each command word path as the occurrence probability of the audio data corresponding to each command word.

[0011] In an exemplary embodiment, determining the occurrence probability of the subsequent occurrence of the audio data corresponding to a command word based on the audio data received after the target audio data includes: in response to determining that the audio data received after the target audio data includes the audio data corresponding to the prefix word of the command class word, determining the occurrence probability of the subsequent occurrence of the audio data corresponding to the command word based on the audio data received after the target audio data.

[0012] In an exemplary embodiment, the method further includes: in response to determining that the audio data received within a predetermined time period after the target audio data does not include the audio data corresponding to the prefix word of the command class word, terminating the operation of determining the occurrence probability of the subsequent occurrence of the audio data corresponding to the command word based on the audio data received after the target audio data.

[0013] According to another embodiment of the present invention, there is provided an apparatus for outputting a command word, including: a detection module configured to detect the type of the currently received audio data when continuously receiving audio data; a first determination module configured to, in response to detecting that the currently received audio data is the target audio data corresponding to a target wake-up word, determine the occurrence probability of the subsequent occurrence of the audio data corresponding to the command word based on the audio data received after the target audio data; and an output module configured to output the target command word in response to determining the audio data corresponding to the target command word with an occurrence probability greater than a first probability threshold.

[0014] According to still another embodiment of the present invention, there is also provided a computer-readable storage medium having a computer program stored therein, wherein the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0015] According to still another embodiment of the present invention, there is also provided an electronic device including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0016] According to the present invention, in the case of continuously receiving audio data, the type of the currently received audio data can be detected. When it is detected that the currently received audio data type is the audio data corresponding to the wake-up word, the probability of the subsequent appearance of the audio data corresponding to the command word can be determined based on the audio data received after this audio data. Once the audio data corresponding to the target command word with an appearance probability greater than the first probability threshold is determined, the target command word will be immediately output, without waiting until all the words included in the input voice are recognized before determining and outputting the command word. Thus, it effectively solves the problem in the related art that it is necessary to recognize all the words included in the input voice before it is possible to determine whether the input voice includes a command word, resulting in low recognition efficiency of the command word and high recognition power consumption. It realizes the purpose of quickly responding to the recognition of the command word, and achieves the purpose of improving the recognition efficiency of the command word and reducing the recognition power consumption. Description of the Drawings

[0017] Figure 1 is the hardware structure block diagram of a mobile terminal for a method of outputting a command word according to an embodiment of the present invention;

[0018] Figure 2 is the flowchart of a method of outputting a command word according to an embodiment of the present invention;

[0019] Figure 3 is the flowchart of command word recognition according to an embodiment of the present invention;

[0020] Figure 4 is the flowchart of a decoding graph according to an embodiment of the present invention;

[0021] Figure 5 is the flowchart of a decoding optimization algorithm with an output symbol (output symbol) postposed according to an embodiment of the present invention;

[0022] Figure 6 is an example diagram with the prefix word "open" of the command word postposed according to an embodiment of the present invention;

[0023] Figure 7 is the structure block diagram of a device for outputting a command word according to an embodiment of the present invention. Detailed Embodiments

[0024] In the following, embodiments of the present invention will be described in detail with reference to the drawings and in conjunction with the embodiments.

[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence.

[0026] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal for a method of outputting command words according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.

[0027] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the method of outputting command words in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the mobile terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0028] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0029] In this embodiment, a method of outputting command words is provided. Figure 2 is a flowchart of a method of outputting command words according to an embodiment of the present invention. As Figure 2As shown, the method includes the following steps:

[0030] S202, when continuously receiving audio data, detect the type of the currently received audio data;

[0031] S204, in response to detecting that the currently received audio data is target audio data corresponding to a target wake-up word, determine the occurrence probability of audio data corresponding to a command word appearing subsequently based on the audio data received after the target audio data;

[0032] S206, in response to determining that the audio data corresponding to the target command word has an occurrence probability greater than the first probability threshold, output the target command word.

[0033] Among them, the above operations can be performed by a controller, a processor, a voice recognition module, etc., and can also be other processing devices or processing units with similar processing capabilities. Among them, the above devices can be integrated into an intelligent device (for example, an intelligent speaker, an intelligent robot, etc. that can control other intelligent devices), or can be independently set with the intelligent device. Here, taking the controller performing the above operations as an example (only an exemplary description, in actual operations, other devices or modules can also perform the above operations) for illustration:

[0034] In the above embodiment, the controller can continuously receive audio data. Among them, the audio data can include multiple words. The controller can detect the audio data while receiving the audio data. When detecting that the currently received audio data is audio data for the wake-up word, it will determine the occurrence probability of audio data corresponding to the command word appearing subsequently based on the audio data received after the audio data. Among them, the audio data can be the voice issued by the user, or the voice issued by other intelligent devices (for example, an intelligent recorder, etc.), or the voice issued by other intelligent robots, etc., and can also be pre-recorded sounds, etc. For example, in the audio "Xiaobai, open the refrigerator", "Xiaobai" is the wake-up word, "open" is the prefix word of the command word, and "refrigerator" is the suffix word of the command word.

[0035] In the above embodiments, the first probability threshold can be preset. For example, when the first probability threshold is 70% (of course, it can also be other probability values, such as 75%, 80%, 90%, etc.), when the occurrence probability of the subsequent audio data corresponding to the command word received after the target audio data is greater than 70%, the target command word is output. When the occurrence probability of the subsequent audio data corresponding to the command word received after the target audio data is less than or equal to 70%, the recognition operation of the audio data received after the audio data corresponding to the wake-up word can be terminated in advance, thereby reducing the power consumption of the command word recognition algorithm. Of course, the presetting of the first probability threshold can be adjusted according to the actual application situation.

[0036] In the above embodiments, in the case of continuously receiving audio data, the type of the currently received audio data can be detected. When it is detected that the type of the current audio data is the audio data corresponding to the wake-up word, the occurrence probability of the subsequent audio data corresponding to the command word can be determined based on the audio data received after this audio data. Once the audio data corresponding to the target command word with an occurrence probability greater than the first probability threshold is determined, the target command word will be immediately output, without waiting until all the words included in the input speech are recognized before determining and outputting the command word. Thus, it effectively solves the problem in the related art that it is necessary to recognize all the words included in the input speech before it is possible to determine whether the input speech includes the command word, resulting in low recognition efficiency of the command word and high recognition power consumption, achieving the purpose of fast response in the recognition of the command word and reaching the purpose of improving the recognition efficiency of the command word and reducing the recognition power consumption.

[0037] In an exemplary embodiment, determining the occurrence probability of the subsequent audio data corresponding to the command word based on the audio data received after the target audio data includes: determining a first probability of the subsequent occurrence of the audio data corresponding to the command word type and a second probability of the occurrence of the audio data corresponding to the non-command word type based on the audio data received after the target audio data; and in response to determining that the first probability is greater than the second probability threshold, determining the occurrence probability of the audio data corresponding to each command word based on the subsequently received audio data. In this embodiment, before determining the occurrence probability of the audio data corresponding to a specific command word, the occurrence probability of the audio data corresponding to the command word category can be determined first. When it is determined that this probability is greater than a certain threshold, the occurrence probability of the audio data corresponding to the specific command word is then determined, thereby avoiding the increase in useless computational amount caused by detecting the occurrence probability of the audio data corresponding to the specific command word when there may be no audio data corresponding to the command word category, and thus avoiding the problem of excessive consumption of hardware resources. In practical applications, the second probability threshold can be preset, for example, set to 60%, 70%, 85%, etc., so that when it is determined that the first probability is greater than the second probability threshold, the occurrence probability of the audio data corresponding to each command word can be determined based on the subsequently received audio data. In addition, when it is determined that the above first probability is less than or equal to the second probability threshold, the calculation can be terminated in advance.

[0038] In an exemplary embodiment, determining a first probability of subsequent occurrence of audio data corresponding to a command word type and a second probability of occurrence of audio data corresponding to a non-command word type based on the audio data received after the target audio data includes: adjusting a first weight of subsequent occurrence of audio data corresponding to a command word type and a second weight of occurrence of audio data corresponding to a non-command word type in a target decoding graph based on the audio data received after the target audio data; determining the first probability and the second probability based on the first weight and the second weight. In this embodiment, the first weight of subsequent occurrence of audio data corresponding to a command word type and the second weight of occurrence of audio data corresponding to a non-command word type can be adjusted in the target decoding graph based on the audio data received after the audio data corresponding to the wake-up word, where the target decoding graph can be a pre-constructed decoding graph. For example, it can be a static decoding graph constructed based on an HCLG network. During construction, first, the language model, pronunciation dictionary, and acoustic model need to be represented in the corresponding FST format, and then compiled into a large decoding graph through operations such as combination, determinization, and minimization to obtain the target decoding graph. In the above embodiment, the first probability and the second probability can be determined based on the adjusted first weight and second weight. In practical applications, the first weight corresponds to the first probability. If the first weight is large, the first probability is large; if the first weight is small, the first probability is small. Optionally, the ratio of the first weight to the sum of the first weight and the second weight can be determined as the first probability, and the ratio of the second weight to the sum of the first weight and the second weight can be determined as the second probability.

[0039] In an exemplary embodiment, adjusting a first weight for audio data corresponding to a command word type and a second weight for audio data corresponding to a non-command word type that subsequently appear based on the audio data received after the target audio data includes: performing frame-level decoding on the audio data received after the target audio data to obtain a first decoding result; continuously adjusting a first initial weight of a command word path and a second initial weight of a non-command word path included in a target decoding graph based on the first decoding result; determining the adjusted first initial weight as the first weight, and determining the adjusted second initial weight as the second weight. In this embodiment, when decoding audio data, frame-level decoding can be performed, that is, decoding is performed frame by frame of audio data. Among them, the length of one frame of audio data can be preset, generally set to about 10 ms. Of course, in practical applications, the length of one frame of audio data can also be set to other durations (for example, set to 5 ms, 15 ms, 20 ms, etc.). When performing frame-level decoding, it is possible to determine (or predict) whether the audio data of the command word type or the non-command word type may appear in the subsequent frame based on the parsing result of the previous frame, so as to adjust the weights of the branches of the command word type and the non-command word type included in the target decoding graph based on the determination result.

[0040] In an exemplary embodiment, determining the occurrence probability of the audio data corresponding to each command word based on the subsequently received audio data includes: performing frame-level decoding on the subsequently received audio data to obtain a second decoding result; continuously adjusting the initial weights corresponding to each command word path included in the target decoding graph based on the second decoding result; and determining the adjusted initial weights corresponding to each command word path as the occurrence probability of the audio data corresponding to each command word. In this embodiment, the target decoding graph may include two types of weights. One type of weight is the weight of the command word path and the non-command word path, and the other type is the weight of each command word specifically included in the command word path. That is, in this embodiment, when it is determined based on the first weight of the command word path (or the command word path) that the audio data corresponding to the command word is likely to appear, it will continue to determine which specific command word appears. That is, it will perform subsequent adjustment of the weights of each command word specifically included in the command word path. For example, when the prefix of the command word "turn on" is detected, "turn on the light", "turn on the TV", "turn on the air conditioner", and "turn on the humidifier" will be set as candidate command words, and the same weight will be configured for each command word branch. During the subsequent decoding process, when "kong" is detected, the weight of the "turn on the air conditioner" branch will be directly adjusted to the maximum, and the command word "turn on the air conditioner" will be output. In addition, if "dian" is detected, the weights of the "turn on the light" branch and the "turn on the TV" branch will be increased, for example, both adjusted to 4, and the weights of the "turn on the air conditioner" and "turn on the humidifier" branches will be adjusted to the minimum, for example, both adjusted to 1. Then, continue the detection. When "deng" is detected again, the weight of the "turn on the light" branch will be directly adjusted to the maximum, and the command word "turn on the light" will be output. Among them, the adjusted initial weights corresponding to each command word path can be used to determine the occurrence probability of the audio data of each command word. The greater the probability of the command word appearing, the greater the probability of being recognized. And the occurrence probability of the audio data of each command word can be determined by the size of the initial weights corresponding to each command word path adjusted according to actual needs.

[0041] In an exemplary embodiment, determining the occurrence probability of subsequent audio data corresponding to a command word based on the audio data received after the target audio data includes: in response to determining that the audio data received after the target audio data includes audio data corresponding to a prefix word of a command word class, determining the occurrence probability of subsequent audio data corresponding to the command word based on the audio data received after the target audio data. In this embodiment, in various application scenarios, there can be various selections for the prefix words of command word classes. For example, "turn off", "turn on", "start", "turn up", "lower", etc. In practical applications, the occurrence probability of subsequent audio data corresponding to the command word can be determined based on the audio data of such prefix words of command word classes, or frame-level decoding can be performed on subsequent prefix words of command word classes to adjust the weight of the prefix words of command word classes, thereby further adjusting the occurrence probability of subsequent audio data corresponding to the command word.

[0042] In an exemplary embodiment, the above method further includes: in response to determining that the audio data received within a predetermined time period after the target audio data does not include audio data corresponding to a prefix word of a command word class, terminating the operation of determining the occurrence probability of subsequent audio data corresponding to the command word based on the audio data received after the target audio data. In this embodiment, a period of time can be preset in advance, for example, 1s, 1.5s, etc. If it is determined that the audio data received after the target audio data does not include audio data corresponding to a prefix word of a command word class within this period of time, the operation of determining the occurrence probability of subsequent audio data corresponding to the command word can be terminated.

[0043] Obviously, the above-described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0044] The present invention will be specifically described below in conjunction with specific embodiments:

[0045] Figure 3 is a flowchart of command word recognition according to an embodiment of the present invention. As Figure 3 shown, the process includes the following steps:

[0046] S302, Human speech. That is, the input of audio. Among them, the audio is not limited to human language, and other recognizable languages (such as bird calls, dog barks, cat meows, etc.) can also be included;

[0047] S304, perform a first judgment to determine whether the wake-up word algorithm detects a wake-up word;

[0048] S306. In the case where the above first judgment result is yes, command the algorithm to enter the ready state from the sleep state. It should be noted that in the case where the above first judgment result is no, it is necessary to return to step S302.

[0049] S308. Make a second judgment to determine whether the prefix word of the command word is detected.

[0050] S310. In the case where the above second judgment result is no, terminate the algorithm in advance.

[0051] S312. In the case where the above second judgment result is yes, continue to recognize the command word.

[0052] Among them, when the wake-up word algorithm is triggered, the command word detection algorithm will also enter the ready state, waiting for the user's further command word voice input. When the command word is detected, the detection result is returned and the corresponding function is triggered. When the command word is not detected, it enters the sleep state after a period of time.

[0053] Figure 4 It is a flowchart of the decoding graph according to an embodiment of the present invention. As Figure 4 shown, after the wake-up word is triggered, it is detected that the user wants to trigger subsequent operations. At this time, the command word detection algorithm can further optimize the decoding graph or further adjust the weight of the command word language model in the decoding graph, so as to increase the probability of successfully recognizing the command word by voice. Subsequently, the false alarm rate of command word recognition can be reduced. The process includes the following steps:

[0054] S402. Whether the wake-up word is triggered. If it is triggered, it can be considered that the user has the intention to pronounce the command word.

[0055] S404. According to the preset decoding graph enhancement coefficient, adjust the weight of the command word language model, so that the command word recognition result is more reliable and accurate. Among them, the enhancement coefficient can be an empirical range, for example, 1 to 1.5, etc.

[0056] S406. Perform command word recognition on the decoding graph with the adjusted language model weight.

[0057] It should be noted that for constructing the static decoding based on the HCLG network, first, the language model, pronunciation dictionary, and acoustic model need to be represented in the corresponding FST format, and then through operations such as combination, determinization, and minimization, they are compiled into a large decoding graph, namely frame adjustment. During the data stream decoding, it is executed frame by frame, with each frame being about 10 ms. During the frame-level decoding, the weight of the command word that appears in the subsequent frame can be determined. Subsequently, when it is determined that the weight of a certain command word is greater than a certain threshold (corresponding to the above-mentioned first threshold), that command word is output, thereby achieving the purpose of fast command word response. In addition, the decoding graph includes a command word path and a garbage vocabulary path (corresponding to the above-mentioned non-command word path). The command word recognition algorithm is only called when the wake-up word algorithm detects the wake-up word. By using the prior information that the wake-up word has been triggered, the path graph score of a larger command word prefix vocabulary can be further set. Therefore, the probability of command word missed detection is reduced, thereby improving the recognition performance of the command word.

[0058] As can be seen from the above embodiments, the wake-up word algorithm is preposed, and the command word algorithm is adjusted based on the wake-up word recognition result. Subsequently, the weight of the command word prefix is increased, further improving the recognition rate of the command word.

[0059] Figure 5 It is a flowchart of the decoding optimization algorithm with the output symbol (output symbol) postposed according to the embodiment of the present invention. As Figure 5 shown, this algorithm process preposes the wake-up word detection algorithm. By using the prior information that the wake-up word has been triggered, the probability that the user speaks the command word can be greatly increased. Subsequently, it can be determined whether the speech data within about 1 s of the first two words contains the prefix word "open" of the command word. In addition, in various scenarios, the choice of the prefix word can be different. For example, "close", "start", "turn up", "lower", etc. If the speech data does not include the above series of prefix words, the recognition of the command word for this speech can be terminated in advance. This process includes the following steps:

[0060] S502, Detect the speech. Without waiting for the entire speech to be fully input, the word-level result can be output in advance;

[0061] S504, Make a first judgment to determine whether the above word is a prefix word;

[0062] S506, When the result of the above first judgment is negative, only decode the speech within the longest Ts = 1 s, terminate the decoding in advance, and its algorithm no longer runs and consumes hardware resources. At the same time, it avoids the problem of false alarms caused by waiting for a long time without the user speaking the command word;

[0063] S508, When the result of the above first judgment is positive, the decoding operation can continue.

[0064] As can be seen from the above embodiments, the optimization idea of decoding with the output symbol post-positioned is to post-position the result of the decoding graph output in word units (i.e., the output symbol). Subsequently, during decoding, if the decoder generates an output result for a word, it can be determined that the speech corresponding to the word has been input completely. Therefore, when the output result of a command word is recognized, it can be immediately returned. At the same time, if speech is recognized within the threshold time Ts = 1 s, but the decoded result does not contain the prefix word of the command word, under the prerequisite that the wake-up word has been activated, it can be directly returned in advance without further judging the subsequent speech, further reducing the power consumption of the algorithm and also preventing false alarms of command words due to overly long recognized statements. While reducing the power consumption of the algorithm, it can reduce false alarms of command words.

[0065] Figure 6 It is an example diagram in which the prefix word "open" of the command word according to the embodiment of the present invention is post-positioned.

[0066] As can be seen from the foregoing embodiments, the command word under the wake-up word pre-positioned algorithm can terminate invalid command words in advance, effectively reducing the power consumption of the command word algorithm. At the same time, based on the recognition result of the wake-up word, it provides effective prior information for the command word detection algorithm (corresponding to the above prior information). Based on this prior information, it is judged whether to dynamically adjust the weight of the language model of the command word decoding graph, thereby reducing the situations of false alarms and missed detections of command words and further improving the recognition rate of command words. That is, based on the prior decoding result of the wake-up word, the decoding probability of the command word prefix is enhanced, and the recognition rate of the command word algorithm is improved; based on the optimization algorithm of the decoding graph with the output symbol post-positioned, when the prefix of the command word does not appear, the recognition algorithm of the command word is terminated in advance, reducing the consumption of hardware and false alarms of the algorithm command word.

[0067] The present invention utilizes the recognition result of the wake-up word. Among them, the wake-up word is used as a pre-positioned algorithm. In the case of confirming that no valid command word is detected, the recognition algorithm can be terminated in advance, further reducing the energy consumption, and in the case of overly long non-command speech input, reducing the false alarm situation of the command word recognition algorithm. For example, if the wake-up word has been confirmed and the recognized prefix is not a command word, there is no need to further wait for the subsequent command word for recognition. At the same time, when the wake-up word is triggered, the weight of the command word can be increased by adopting the decoding graph design, further reducing the missed detection of the command word.

[0068] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0069] In this embodiment, a device for outputting command words is further provided. This device is used to implement the above embodiments and preferred embodiments, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0070] Figure 7 is a structural block diagram of a device for outputting command words according to an embodiment of the present invention. As Figure 7 shown, the device includes:

[0071] A detection module 72, configured to detect the type of currently received audio data when continuously receiving audio data;

[0072] A first determination module 74, configured to, in response to detecting that the currently received audio data is target audio data corresponding to a target wake-up word, determine the occurrence probability of subsequent audio data corresponding to a command word based on the audio data received after the target audio data;

[0073] An output module 76, configured to output the target command word in response to determining that the occurrence probability of audio data corresponding to a target command word is greater than a first probability threshold.

[0074] In an optional embodiment, the first determination module 74 includes: a first determination sub-module, configured to determine a first probability of subsequent audio data corresponding to a command word type and a second probability of subsequent audio data corresponding to a non-command word type based on the audio data received after the target audio data; a second determination sub-module, configured to, in response to determining that the first probability is greater than a second probability threshold, determine the occurrence probability of audio data corresponding to each command word based on the subsequently received audio data.

[0075] In an optional embodiment, the first determination sub-module includes: a first adjustment unit configured to adjust a first weight of audio data corresponding to a command word type and a second weight of audio data corresponding to a non-command word type that subsequently appear in a target decoding graph based on the audio data received after the target audio data; and a first determination unit configured to determine the first probability and the second probability based on the first weight and the second weight.

[0076] In an optional embodiment, the first adjustment unit includes: a decoding sub-unit configured to perform frame-level decoding on the audio data received after the target audio data to obtain a first decoding result; an adjustment sub-unit configured to continuously adjust a first initial weight of a command word path and a second initial weight of a non-command word path included in the target decoding graph based on the first decoding result; and a determination sub-unit configured to determine the adjusted first initial weight as the first weight and the adjusted second initial weight as the second weight.

[0077] In an optional embodiment, the second determination sub-module includes: a decoding unit configured to perform frame-level decoding on the subsequently received audio data to obtain a second decoding result; a second adjustment unit configured to continuously adjust an initial weight corresponding to each command word path in the command word paths included in the target decoding graph based on the second decoding result; and a second determination unit configured to determine the adjusted initial weight corresponding to each command word path as the appearance probability of the audio data corresponding to each command word.

[0078] In an optional embodiment, the first determination module 74 further includes: a second determination sub-module configured to, in response to determining that the audio data received after the target audio data includes audio data corresponding to a prefix word of a command type word, determine the appearance probability of subsequent audio data corresponding to a command word based on the audio data received after the target audio data.

[0079] In an optional embodiment, the above-mentioned apparatus further includes: a termination module configured to, in response to determining that the audio data received within a predetermined time period after the target audio data does not include audio data corresponding to a prefix word of a command type word, terminate the operation of determining the appearance probability of subsequent audio data corresponding to a command word based on the audio data received after the target audio data.

[0080] It should be noted that the above-mentioned respective modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited thereto: all the above-mentioned modules are located in the same processor; or, the above-mentioned respective modules are located in different processors in any combination form.

[0081] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.

[0082] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM), random access memories (RAM), mobile hard disks, magnetic disks, or optical discs that can store computer programs.

[0083] An embodiment of the present invention further provides an electronic device including a memory and a processor, where the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above method embodiments.

[0084] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0085] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.

[0086] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general-purpose computing device. They can be centralized on a single computing device or distributed over a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. Thus, the present invention is not limited to any specific combination of hardware and software.

[0087] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for outputting command words, characterized in that, Including: Detecting the type of currently received audio data while continuously receiving audio data; In response to detecting that the currently received audio data is target audio data corresponding to a target wake-up word, determining the occurrence probability of subsequent audio data corresponding to a command word based on the audio data received after the target audio data; In response to determining that the audio data corresponding to the target command word has an occurrence probability greater than a first probability threshold, outputting the target command word; Determining the occurrence probability of subsequent audio data corresponding to a command word based on the audio data received after the target audio data includes: determining a first probability of subsequent audio data corresponding to the command word type and a second probability of audio data corresponding to a non-command word type based on the audio data received after the target audio data; in response to determining that the first probability is greater than a second probability threshold, determining the occurrence probability of the audio data corresponding to each command word based on the subsequently received audio data; Determining the occurrence probability of the audio data corresponding to each command word based on the subsequently received audio data includes: performing frame-level decoding on the subsequently received audio data to obtain a second decoding result; continuously adjusting the initial weight corresponding to each command word path in the command word paths included in the target decoding graph based on the second decoding result; determining the adjusted initial weight corresponding to each command word path as the occurrence probability of the audio data corresponding to each command word.

2. The method according to claim 1, wherein Determining a first probability of subsequent audio data corresponding to the command word type and a second probability of audio data corresponding to a non-command word type based on the audio data received after the target audio data includes: Adjusting a first weight of subsequent audio data corresponding to the command word type and a second weight of audio data corresponding to a non-command word type in the target decoding graph based on the audio data received after the target audio data; Determining the first probability and the second probability based on the first weight and the second weight.

3. The method according to claim 2, wherein Adjusting a first weight of subsequent audio data corresponding to the command word type and a second weight of audio data corresponding to a non-command word type based on the audio data received after the target audio data includes: Performing frame-level decoding on the audio data received after the target audio data to obtain a first decoding result; Continuously adjusting the first initial weight of the command word paths and the second initial weight of the non-command word paths included in the target decoding graph based on the first decoding result; Determining the adjusted first initial weight as the first weight, and determining the adjusted second initial weight as the second weight.

4. The method according to claim 1, characterized in that, Determining the occurrence probability of subsequent audio data corresponding to a command word based on the audio data received after the target audio data includes: In response to determining that the audio data received after the target audio data includes audio data corresponding to a prefix word of a command word type, determine the occurrence probability of subsequent audio data corresponding to a command word based on the audio data received after the target audio data.

5. The method according to claim 1, characterized in that, The method further includes: In response to determining that the audio data received within a predetermined time period after the target audio data does not include audio data corresponding to a prefix word of a command word type, terminate the operation of determining the occurrence probability of subsequent audio data corresponding to a command word based on the audio data received after the target audio data.

6. An apparatus for outputting command words, characterized in that, It includes: A detection module, configured to detect the type of the currently received audio data when continuously receiving audio data; A first determination module, configured to, in response to detecting that the currently received audio data is target audio data corresponding to a target wake-up word, determine the occurrence probability of subsequent audio data corresponding to a command word based on the audio data received after the target audio data; An output module, configured to output the target command word in response to determining audio data corresponding to a target command word whose occurrence probability is greater than a first probability threshold; The first determination module includes: a first determination sub-module, configured to determine a first probability of subsequent audio data corresponding to a command word type and a second probability of audio data corresponding to a non-command word type based on the audio data received after the target audio data; a second determination sub-module, configured to, in response to determining that the first probability is greater than a second probability threshold, determine the occurrence probability of audio data corresponding to each command word based on the subsequently received audio data; The second determination sub-module includes: a decoding unit, configured to perform frame-level decoding on the subsequently received audio data to obtain a second decoding result; a second adjustment unit, configured to continuously adjust the initial weight corresponding to each command word path in the command word paths included in the target decoding graph based on the second decoding result; a second determination unit, configured to determine the adjusted initial weight corresponding to each command word path as the occurrence probability of audio data corresponding to each command word.

7. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 5 are implemented.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Speech recognition method and apparatus, terminal, and computer readable storage medium

    CN107644638A

  • Voice recognition method and device

    CN113516967A

  • Contextual Biasing for Speech Recognition

    US20200357387A1

  • Speech control method, electronic device, and storage medium

    US20210319795A1