Audio wake-up method, device, electronic device, storage medium, and computer program

The voice wake-up method improves human-computer interaction by using word and syllable recognition to ensure accurate and efficient wake-up responses, reducing false alarms and power consumption in voice dialogue devices.

JP7811641B2Active Publication Date: 2026-02-05BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024521288
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-07-15
Filing Date
2023-01-17
Publication Date
2026-02-05
Estimated Expiration
2043-01-17

AI Technical Summary

Technical Problem

Existing voice interaction systems face challenges in achieving fast wake-up response, accurate word and syllable recognition, and prompt feedback, which affect the smoothness of human-computer interaction.

Method used

A voice wake-up method that performs word recognition followed by syllable recognition to identify a correct wake-up audio, using a lightweight ProjectedLight-GRU module for word recognition and a Conformer-based model for syllable recognition, ensuring accurate and efficient wake-up even with fewer wake-up words.

Benefits of technology

Enhances wake-up accuracy and reduces false alarms by recognizing both whole words and syllables, improving response speed and reducing power consumption in voice dialogue devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007811641000002
    Figure 0007811641000002
  • Figure 0007811641000003
    Figure 0007811641000003
  • Figure 0007811641000004
    Figure 0007811641000004
Patent Text Reader

Abstract

The present disclosure provides a voice wake-up method, device, electronic device, storage medium and program product, which relate to the field of artificial intelligence technology, particularly to the technical fields of human-computer interaction, deep learning, intelligent voice, etc. A specific implementation manner includes: performing word recognition on the voice to be recognized to obtain a wake-up word recognition result; when it is determined that the wake-up word recognition result indicates that the voice to be recognized contains a predetermined wake-up word, performing syllable recognition on the voice to be recognized to obtain a wake-up syllable recognition result; and when it is determined that the wake-up syllable recognition result indicates that the voice to be recognized contains a predetermined syllable, determining that the voice to be recognized is a correct wake-up voice.
Need to check novelty before this filing date? Find Prior Art

Description

cross reference

[0001] This application claims priority to a Chinese patent application bearing application number 202210838284.6, filed on July 15, 2022, the entire contents of which are incorporated herein by reference. [Technical Field]

[0002] The present disclosure relates to the field of artificial intelligence technology, particularly to the fields of human-computer interaction, deep learning, intelligent voice, etc. Specifically, the present disclosure relates to a voice wake-up method, device, electronic device, storage medium, and computer Program Mu Regarding. [Background technology]

[0003] Voice interaction is a natural human interaction method. With the development of artificial intelligence technology, devices are now able to listen to human speech, understand the underlying meaning of speech, and provide corresponding feedback. In these operations, the wake-up response speed, wake-up difficulty, accurate understanding of word meaning, and promptness of feedback are all factors that affect the smoothness of voice interaction. Summary of the Invention

[0004] The present disclosure relates to a voice wake-up method, a device, an electronic device, a storage medium, and computer Program M provide.

[0005] According to one aspect of the present disclosure, there is provided an audio wake-up method including: performing word recognition on audio to be recognized and obtaining a wake-up word recognition result; if the wake-up word recognition result indicates that the audio to be recognized contains a predetermined wake-up word, performing syllable recognition on the audio to be recognized and obtaining a wake-up syllable recognition result; and if the wake-up syllable recognition result indicates that the audio to be recognized contains a predetermined syllable, identifying the audio to be recognized as a correct wake-up audio.

[0006] According to another aspect of the present disclosure, there is provided an audio wake-up device including: a word recognition module that performs word recognition on audio to be recognized and obtains a wake-up word recognition result; a syllable recognition module that performs syllable recognition on the audio to be recognized and obtains a wake-up syllable recognition result when it is determined that the wake-up word recognition result indicates that the audio to be recognized contains a predetermined wake-up word; and a first identification module that identifies the audio to be recognized as a correct wake-up audio when it is determined that the wake-up syllable recognition result indicates that the audio to be recognized contains a predetermined syllable.

[0007] According to another aspect of the present disclosure, there is provided an electronic device including at least one processor and a memory communicatively coupled to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method of the present disclosure.

[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium having computer instructions stored thereon, the computer instructions causing the computer to perform a method of the present disclosure, is provided.

[0009] According to another aspect of the present disclosure, there is provided a computer program product that, when executed by a processor, implements the method of the present disclosure. M provide.

[0010] It should be understood that the contents described in this section are not intended to identify key features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily apparent from the following specification. [Brief explanation of the drawings]

[0011] The drawings are for a better understanding of the invention and are not intended to limit the disclosure. [Figure 1] FIG. 1 schematically illustrates an exemplary system architecture to which the voice wake-up method and apparatus according to the embodiments of the present disclosure can be applied. [Figure 2] FIG. 2 is a schematic diagram illustrating a flowchart of a voice wake-up method according to an embodiment of the present disclosure. [Figure 3] FIG. 3 is a schematic diagram illustrating a network configuration of a wake-up word recognition model according to an embodiment of the present disclosure. [Figure 4] FIG. 4 is a schematic diagram illustrating a network configuration of a wake-up syllable recognition model according to an embodiment of the present disclosure. [Figure 5] FIG. 5 is a schematic flow chart of a voice wake-up method according to another embodiment of the present disclosure. [Figure 6] FIG. 6 is a schematic diagram illustrating an application diagram of a voice wake-up method according to another embodiment of the present disclosure. [Figure 7] FIG. 7 is a schematic block diagram of a voice wake-up device according to an embodiment of the present disclosure. [Figure 8] FIG. 8 is a schematic block diagram of an electronic device suitable for implementing the audio wake-up method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0012]

[0023] The following description of exemplary embodiments of the present disclosure will be made with reference to the drawings. For ease of understanding, various details of the embodiments of the present disclosure are included, and it should be understood that these details are merely exemplary. Therefore, it should be understood that those skilled in the art can make various changes and modifications to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, the following description will omit descriptions of known functions and structures.

[0013] The present disclosure relates to a voice wake-up method, a device, an electronic device, a storage medium, and computer Program M provide.

[0014] According to one aspect of the present disclosure, there is provided a voice wake-up method, comprising: performing word recognition on a voice to be recognized and obtaining a wake-up word recognition result; if the wake-up word recognition result indicates that the voice to be recognized contains a predetermined wake-up word, performing syllable recognition on the voice to be recognized and obtaining a wake-up syllable recognition result; and if the wake-up syllable recognition result indicates that the voice to be recognized contains a predetermined syllable, identifying the voice to be recognized as a correct wake-up voice.

[0015] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, application, etc. of such user personal information shall all comply with the provisions of relevant laws and regulations, take necessary security measures, and not violate public order and morals.

[0016] In the technical solution disclosed herein, the user's approval or consent is obtained before obtaining or collecting the user's personal information.

[0017] FIG. 1 is a schematic diagram illustrating an application scenario of a voice wake-up method and device according to an embodiment of the present disclosure.

[0018] Note that FIG. 1 merely illustrates an example of an application scene of the embodiments of the present disclosure to help those skilled in the art understand the technical content of the present disclosure, and is not intended to mean that the embodiments of the present disclosure are not applicable to other devices, systems, environments, or scenes.

[0019] As shown in FIG. 1, a user 102 can transmit a voice to be recognized to the voice dialogue device 101, and the voice dialogue device 101 can determine whether the voice to be recognized is a correct wake-up voice. If the voice dialogue device 101 determines that the voice to be recognized is a correct wake-up voice, it acquires an instruction voice containing the user's intention information, performs the intended operation in the instruction voice, and realizes human-computer interaction between the user 102 and the voice dialogue device 101.

[0020] The voice interaction device 101 may have installed thereon various communication client applications such as (by way of example only) a knowledge browsing application, a web browser application, a search application, an instant messaging tool, a mailbox client and / or social platform software.

[0021] The voice dialogue device 101 may include a sound collector, such as a microphone, for collecting the voice of the user 102 to be recognized and an instruction voice including intention information. The voice dialogue device 101 may further include a voice player, such as a speaker, for playing back the voice transmitted from the voice dialogue device.

[0022] The voice interaction device 101 may be any electronic device capable of interacting via voice signals, including, but not limited to, a smartphone, a tablet computer, a laptop portable computer, a smart home appliance, a smart speaker, an in-car speaker, a smart home appliance, or a smart robot.

[0023] It should be noted that the syllable recognition model and the keyword recognition model according to the embodiments of the present disclosure may be installed in the voice dialogue device 101, and the voice processing method may generally be executed by the voice dialogue device 101. Accordingly, the voice processing device according to the embodiments of the present disclosure may be provided in the voice dialogue device 101. The terminal device can realize the voice wake-up method and device according to the embodiments of the present disclosure without needing to interact with a server.

[0024] Without being limited to this, in other embodiments of the present disclosure, the voice dialogue device may transmit the voice to be recognized to a server via a network, and the server may process the voice to be recognized and determine whether the voice to be recognized is a correct wake-up voice.

[0025] It should be noted that the sequence numbers of each operation in the following methods are merely described as a representation of the operations and should not be considered as indicating the order in which the operations are performed. Unless otherwise specified, the methods do not have to be performed in the exact order shown.

[0026] FIG. 2 is a schematic diagram illustrating a flowchart of a voice wake-up method according to an embodiment of the present disclosure.

[0027] As shown in FIG. 2, the method includes operations S210-S230.

[0028] In operation S210, word recognition is performed on the speech to be recognized, and a wake-up word recognition result is obtained.

[0029] In operation S220, if it is determined that the wake-up word recognition result indicates that the speech to be recognized contains a predetermined wake-up word, syllable recognition is performed on the speech to be recognized to obtain a wake-up syllable recognition result.

[0030] In operation S230, if it is determined that the wake-up syllable recognition result indicates that the speech to be recognized contains a predetermined syllable, it is determined that the speech to be recognized is a correct wake-up speech.

[0031] According to an embodiment of the present disclosure, the speech to be recognized may be a wake-up speech, which may be a speech signal received before the voice interaction function wakes up, such as a speech containing a wake-up word or a speech containing a non-wake-up word.

[0032] According to an embodiment of the present disclosure, the correct wake-up voice may be a voice containing a wake-up word or a voice that can wake up a voice interaction function. If the voice to be recognized is determined to be the correct wake-up voice, the voice interaction function of the voice interaction device may be triggered. If the voice to be recognized is determined to be an incorrect wake-up voice, the operation may be stopped and no response may be made to the user.

[0033] According to an embodiment of the present disclosure, the voice interaction function may be a function that can receive an interaction voice from a user and output an audio feedback result corresponding to the interaction voice to the user.

[0034] According to an embodiment of the present disclosure, performing word recognition on the speech to be recognized may be performing wake-up word recognition on the speech to be recognized. The word recognition is performed on the speech to be recognized, and a wake-up word recognition result is obtained. The wake-up word recognition result may indicate whether the speech to be recognized includes a predetermined wake-up word.

[0035] According to an embodiment of the present disclosure, word recognition is performed on the speech to be recognized, and the speech to be recognized is recognized globally or comprehensively to obtain a wake-up word recognition result. For example, if the predetermined wake-up word is "small D", the speech to be recognized is "Hello small D", and word recognition is performed on the speech to be recognized to obtain a wake-up word recognition result indicating that the speech to be recognized contains the predetermined wake-up word.

[0036] According to another embodiment of the present disclosure, it is possible to determine whether the voice to be recognized is a correct wake-up voice based on the wake-up word recognition result. For example, if it is determined that the wake-up word recognition result indicates that the voice to be recognized contains a predetermined wake-up word, it is possible to determine that the voice to be recognized is a correct wake-up voice. The man-machine interaction function can be turned on. If it is determined that the wake-up word recognition result indicates that the voice to be recognized does not contain the predetermined wake-up word, it is possible to determine that the voice to be recognized is an incorrect wake-up voice. There is no need to respond.

[0037] According to an embodiment of the present disclosure, when it is determined that the wake-up word recognition result indicates that the speech to be recognized contains a predetermined wake-up word, syllable recognition can be performed on the speech to be recognized to obtain a wake-up syllable recognition result.

[0038] According to an embodiment of the present disclosure, performing syllable recognition on the speech to be recognized may be performing syllable recognition corresponding to a wake-up word on the speech to be recognized to obtain a wake-up syllable recognition result. The wake-up syllable recognition result indicates whether the speech to be recognized includes a predetermined syllable. The predetermined syllable may refer to a syllable corresponding to a predetermined wake-up word.

[0039] According to an embodiment of the present disclosure, syllable recognition is performed on the speech to be recognized, and the speech to be recognized is recognized by a local or slave byte unit. For example, if the predetermined syllables corresponding to the predetermined wake-up word "small D" are the syllables "small" and "D", the speech to be recognized is "small D hello", and by performing syllable recognition on the speech to be recognized, a wake-up message indicating that the predetermined syllable is included in the speech to be recognized is generated. syllable The recognition result can be obtained.

[0040] According to an embodiment of the present disclosure, if it is determined that the wake-up syllable recognition result indicates that the speech to be recognized contains a predetermined syllable, the speech to be recognized is determined to be a correct wake-up speech. If it is determined that the wake-up syllable recognition result indicates that the speech to be recognized does not contain a predetermined syllable, the speech to be recognized is determined to be an incorrect wake-up speech.

[0041] According to another embodiment of the present disclosure, it is possible to obtain a wake-up syllable recognition result by performing only syllable recognition on the speech to be recognized without performing word recognition on the speech to be recognized. Whether the speech to be recognized is a correct wake-up speech is determined based on the wake-up syllable recognition result. For example, if it is determined that the wake-up syllable recognition result indicates that the speech to be recognized contains a predetermined syllable, it is determined that the speech to be recognized is a correct wake-up speech. The man-machine interaction function can be turned on. If it is determined that the wake-up syllable recognition result indicates that the speech to be recognized does not contain a predetermined syllable, it is determined that the speech to be recognized is an incorrect wake-up speech. No response is required.

[0042] According to an embodiment of the present disclosure, compared to a method of determining whether the voice to be recognized is a correct wake-up voice based only on the wake-up word recognition result, or determining whether the voice to be recognized is a correct wake-up voice based only on the wake-up syllable recognition result, the method provided in the present disclosure, "if the wake-up word recognition result determines that the voice to be recognized contains a predetermined wake-up word, perform syllable recognition on the voice to be recognized to obtain the wake-up syllable recognition result, and determine whether the voice to be recognized is a correct wake-up voice based on the wake-up syllable recognition result," can use word recognition operation to perform whole-word unit recognition of the wake-up word on the voice to be recognized, and at the same time, use syllable recognition operation to perform word unit recognition of the wake-up word on the voice to be recognized, so that the voice to be recognized can be recognized both globally and locally, thereby ensuring wake-up accuracy and avoiding false wake-up alarms when the number of wake-up words is four or less, for example, three or two.

[0043] According to another embodiment of the present disclosure, in operation S210 shown in FIG. 2 , performing word recognition on the speech to be recognized and obtaining a wake-up word recognition result may further include performing a convolution operation on the speech to be recognized to obtain a first-stage feature vector sequence; performing a gate recurrent operation on the first-stage feature vector sequence to obtain a second-stage feature vector sequence; and performing a classification operation on the second-stage feature vector sequence to obtain a wake-up word recognition result.

[0044] According to an embodiment of the present disclosure, the speech to be recognized may include a sequence of speech frames, and the first-stage feature vector sequence has a one-to-one correspondence with the speech frame sequence.

[0045] According to an embodiment of the present disclosure, word recognition can be performed on speech to be recognized using a wake-up word recognition model to obtain a wake-up word recognition result. However, this is not limited to this. Word recognition can be performed on speech to be recognized using other methods, and any word recognition method that can obtain a wake-up word recognition result is acceptable.

[0046] FIG. 3 is a schematic diagram illustrating a network configuration of a wake-up word recognition model according to an embodiment of the present disclosure.

[0047] As shown in FIG. 3, the wake-up word recognition model includes a convolution module 310, a gated recurrent unit 320, and a wake-up word classification module 330 in this order.

[0048] 3, the speech to be recognized 340 is input to a convolution module 310 to obtain a first-stage feature vector sequence. The first-stage feature vector sequence is input to a gated recurrent unit 320 to obtain a second-stage feature vector sequence. The second-stage feature vector sequence is input to a wake-up word classification module 330 to obtain a wake-up word recognition result 350.

[0049] According to an embodiment of the present disclosure, the number of convolutional modules in the wake-up word recognition model is not limited to one, and may include multiple stacked convolutional modules. Similarly, the wake-up word recognition model may include multiple stacked gate recurrent units.

[0050] According to an embodiment of the present disclosure, the convolution module may include one or more combinations of CNNs (Convolutional Neural Networks), RNNs (Recurrent Neural Networks), LSTMs (Long Short-Term Memory Networks), and the like.

[0051] According to an embodiment of the present disclosure, the wake-up word classification module may include a fully connected layer and an activation function. The activation function may be, but is not limited to, a Softmax activation function, and may be a Sigmoid activation function. The number of layers in the fully connected layer is not limited, and may be, for example, one layer or multiple layers.

[0052] According to an embodiment of the present disclosure, the gate recurrent unit may refer to a GRU (GateRecurrentUnit), but is not limited to this, and may be, for example, a GRU induction module after performing a lightening process on the GRU.

[0053] According to an embodiment of the present disclosure, a GRU guidance module, also known as a ProjectedLight-GRU module, is used to install a wake-up word recognition model in a terminal device, such as a voice dialogue device, which is advantageous for lightweight placement on the terminal side and ensures real-time word recognition for the speech to be recognized.

[0054] According to another embodiment of the present disclosure, performing a gate recurrent operation on a first-stage feature vector sequence to obtain a second-stage feature vector sequence includes repeatedly identifying a current time update gate and current time candidate hidden layer information based on a previous time output vector and a current time input vector that is a first-stage feature vector for the current time in the first-stage feature vector sequence, identifying current time hidden layer information based on the current time candidate hidden layer information, previous time hidden layer information, and current time update gate, and identifying a current time output vector that is a second-stage feature vector for the current time in the second-stage feature vector sequence based on the current time hidden layer information and a predetermined parameter.

[0055] According to an embodiment of the present disclosure, the predetermined parameters, also referred to as projection parameters, are identified based on a threshold value for the number of lightening parameters.

[0056] According to an embodiment of the present disclosure, the threshold value of the number of lightening parameters refers to a parameter setting standard, for example, a threshold value of the number of predetermined parameters, and the magnitude of the predetermined parameters is less than or equal to the threshold value of the number of lightening parameters, thereby reducing the amount of data processing of the wake-up word recognition model.

[0057] According to an embodiment of the present disclosure, the ProjectedLight-GRU module can be expressed by the following equations (1)-(4).

[0058]

number

[0059] According to an embodiment of the present disclosure, compared to a standard GRU, the ProjectedLight-GRU module according to the embodiment of the present disclosure eliminates the reset gate and introduces a predetermined parameter, thereby reducing the computational complexity of the wake-up word recognition model. The wake-up word recognition model with the ProjectedLight-GRU module is applied to a voice dialogue device to achieve high performance while reducing resource overhead. The wake-up word recognition model installed in the voice dialogue device realizes all-weather operation conditions and improves the wake-up response speed of the voice dialogue device.

[0060] According to another embodiment of the present disclosure, in operation S220 as shown in Figure 2, if it is determined that the wake-up word recognition result indicates that the speech to be recognized contains a predetermined wake-up word, performing syllable recognition on the speech to be recognized and obtaining a wake-up syllable recognition result may further include performing syllable feature extraction on the speech to be recognized and obtaining a syllable feature matrix, and performing a classification operation on the syllable feature matrix to obtain the wake-up syllable recognition result.

[0061] According to an embodiment of the present disclosure, syllable recognition can be performed on speech to be recognized using a syllable recognition model to obtain a wake-up syllable recognition result. However, this is not limiting. Syllable recognition can be performed on speech to be recognized using other methods, and any syllable recognition method that can obtain a wake-up syllable recognition result may be used.

[0062] FIG. 4 is a schematic diagram illustrating a network configuration of a wake-up syllable recognition model according to an embodiment of the present disclosure.

[0063] As shown in FIG. 4, the wake-up syllable recognition model includes a feature extraction and encoding module 410 and a syllable classification module 420 in that order.

[0064] As shown in Figure 4, speech to be recognized 430 is input to a feature extraction and encoding module 410, which performs syllable feature extraction and outputs a syllable feature matrix. The syllable feature matrix is ​​input to a syllable classification module 420, which performs classification and outputs a wake-up syllable recognition result 440.

[0065] According to an embodiment of the present disclosure, the syllable classification module may include a fully connected layer and an activation function. The activation function may be, but is not limited to, a Softmax activation function, and may also be a Sigmoid activation function. The number of layers in the fully connected layer is not limited, and may be, for example, one layer or multiple layers.

[0066] According to an embodiment of the present disclosure, the feature extraction encoding module may be constructed by a network structure in a Conformer model (an encoder based on convolutional enhancement), but is not limited to this, and may also adopt a Conformer module in a Conformer model, or the Conformer model or Conformer module may be a network structure obtained through a lightweight process such as pruning.

[0067] According to an embodiment of the present disclosure, performing syllable feature extraction on the speech to be recognized and obtaining a syllable feature matrix may further include performing feature extraction on the speech to be recognized and obtaining a feature matrix; performing dimension reduction on the feature matrix to obtain a dimension-reduced feature matrix; and performing a multi-stage speech enhancement encoding process on the dimension-reduced feature matrix to obtain a syllable feature matrix.

[0068] According to an embodiment of the present disclosure, the feature extraction encoding module may sequentially include a feature extraction layer, a dimensionality reduction layer, and an encoding layer. The feature extraction layer may be used to perform feature extraction on the speech to be recognized to obtain a feature matrix. The dimensionality reduction layer may be used to perform dimensionality reduction on the feature matrix to obtain a reduced-dimensionality feature matrix. The encoding layer may be used to perform multiple speech-accurate encoding processes on the reduced-dimensionality feature matrix to obtain a syllable feature matrix.

[0069] According to an embodiment of the present disclosure, the feature extraction layer may include at least one of at least one relative sinusoidal position coding layer, at least one convolution layer, and at least one feed forward layer.

[0070] According to an embodiment of the present disclosure, the coding layer may include a conformer module, such as at least one of a plurality of feedforward layers, at least one multi-headed self-attention module, and at least one convolutional layer.

[0071] According to an embodiment of the present disclosure, the dimensionality reduction layer may include, but is not limited to, a mapping function, and may include, for example, a layer structure that reduces the dimension of a high-dimensional matrix to obtain a low-dimensional matrix.

[0072] According to an embodiment of the present disclosure, the amount of data input to the coding layer can be reduced by using a dimension reduction layer, thereby reducing the amount of calculation for the syllable recognition model. The number of stacked coding layers can also be reduced, for example, by specifying the number of stacked coding layers as one to four based on a threshold value for the number of weight reduction parameters.

[0073] According to an embodiment of the present disclosure, by designing a dimension reduction layer in a wake-up syllable recognition model and controlling the number of stacked coding layers, it is possible to reduce the weight and size of the wake-up syllable recognition model while ensuring recognition accuracy, thereby further improving recognition efficiency, and when the wake-up syllable recognition model is applied to a terminal device, it is possible to reduce the power consumption of the processor of the terminal device.

[0074] FIG. 5 is a schematic flow chart of a voice wake-up method according to another embodiment of the present disclosure.

[0075] As shown in FIG. 5 , speech 510 to be recognized is input to a wake-up word recognition model 520, and a wake-up word recognition result 530 is obtained. If the wake-up word recognition result 530 is determined to indicate that the speech 510 to be recognized contains a predetermined wake-up word, the speech 510 to be recognized is input to a wake-up syllable recognition model 540, and a wake-up syllable recognition result 550 is obtained. If the wake-up syllable recognition result 550 is determined to indicate that the speech to be recognized contains a predetermined syllable, the speech to be recognized is determined to be correct wake-up speech. The speech dialogue device is woken up, and subsequent human-computer interaction can be performed. If the wake-up word recognition result is determined to indicate that the speech to be recognized does not contain the predetermined wake-up word, the speech to be recognized is determined to be incorrect wake-up speech, and operation is stopped. If it is determined that the wake-up syllable recognition result indicates that the speech to be recognized does not contain a predetermined syllable, the speech to be recognized is determined to be an incorrect wake-up speech, and the speech dialogue device is not woken up.

[0076] According to another embodiment of the present disclosure, the speech to be recognized may be input to a wake-up syllable recognition model to obtain a wake-up syllable recognition result. If the wake-up syllable recognition result is determined to indicate that the speech to be recognized contains a predetermined syllable, the speech to be recognized is input to a wake-up word recognition model to obtain a wake-up word recognition result. If the wake-up word recognition result is determined to indicate that the speech to be recognized contains a predetermined wake-up word, the speech to be recognized is determined to be correct wake-up speech. The voice dialogue device is woken up, and subsequent human-computer interaction can be performed. If the wake-up syllable recognition result is determined to indicate that the speech to be recognized does not contain the predetermined syllable, the speech to be recognized is determined to be incorrect wake-up speech, and operation is stopped. If the wake-up word recognition result is determined to indicate that the speech to be recognized contains a predetermined wake-up word, the speech to be recognized is determined to be incorrect wake-up speech, and the voice dialogue device is not woken up.

[0077] According to another embodiment of the present disclosure, the speech to be recognized may be input to a wake-up word recognition model to obtain a wake-up word recognition result. The speech to be recognized may be input to a wake-up syllable recognition model to obtain a wake-up syllable recognition result. If the wake-up word recognition result indicates that the speech to be recognized contains a predetermined wake-up word and the syllable recognition result indicates that the speech to be recognized contains a predetermined syllable, the speech to be recognized is determined to be correct wake-up speech. If the wake-up word recognition result indicates that the speech to be recognized does not contain the predetermined wake-up word or if the syllable recognition result indicates that the speech to be recognized does not contain a predetermined syllable, the speech to be recognized is determined to be incorrect wake-up speech.

[0078] According to an embodiment of the present disclosure, by processing the speech to be recognized using the above-mentioned wake-up word recognition model and wake-up syllable recognition model, it can be applied to scenes where the number of wake-up words is reduced, and when the number of wake-up words is one, two, or three, it is possible to reduce the false alarm rate while maintaining recognition accuracy.

[0079] According to an embodiment of the present disclosure, compared to a method of "first performing syllable recognition on the speech to be recognized using a wake-up syllable model" or a method of "performing syllable recognition on the speech to be recognized using a wake-up syllable model while performing word recognition on the speech to be recognized using a wake-up word recognition model," the method of "first performing word recognition on the speech to be recognized using a wake-up word recognition model" has the characteristics of a simpler network structure of the wake-up word recognition model and a smaller amount of calculation, and therefore, when the terminal device is in a real-time active state, it is possible to reduce power consumption when the voice dialogue device is used as the terminal device while ensuring recognition accuracy.

[0080] FIG. 6 is a schematic diagram illustrating an application diagram of a voice wake-up method according to another embodiment of the present disclosure.

[0081] As shown in Figure 6, a user 610 transmits a speech to be recognized to a speech dialogue device 620. The speech dialogue device uses a wake-up word recognition model and a wake-up syllable recognition model installed in the speech dialogue device 610 to operate the speech wake-up method on the speech to be recognized and determine whether the speech to be recognized is a correct wake-up speech. If the speech to be recognized is determined to be a correct wake-up speech, the speech dialogue device 620 presents a target object 630 on a display interface 621 of the speech dialogue device 620, and simultaneously outputs a feedback speech. This allows for flexible and vivid expression of human-computer interaction.

[0082] FIG. 7 shows a schematic block diagram of a voice wake-up device according to an embodiment of the present disclosure.

[0083] As shown in FIG. 7, the voice wake-up device 700 includes a word recognition module 710, a syllable recognition module 720, and a first identification module 730.

[0084] The word recognition module 710 performs word recognition on the speech to be recognized and obtains a wake-up word recognition result.

[0085] If the syllable recognition module 720 determines that the wake-up word recognition result indicates that the speech to be recognized contains a predetermined wake-up word, it performs syllable recognition on the speech to be recognized and obtains a wake-up syllable recognition result.

[0086] If the first identification module 730 determines that the wake-up syllable recognition result indicates that the voice to be recognized contains a predetermined syllable, it identifies the voice to be recognized as the correct wake-up voice.

[0087] According to an embodiment of the present disclosure, the word recognition module includes a convolution means, a gate unit, and a word classification means.

[0088] The convolution means performs a convolution operation on the speech to be recognized to obtain a first-stage feature vector sequence, where the speech to be recognized includes a speech frame sequence, and the first-stage feature vector sequence corresponds one-to-one with the speech frame sequence.

[0089] The gate unit performs a gate recurrent operation on the first-stage feature vector sequence to obtain a second-stage feature vector sequence.

[0090] The word classification means performs a classification operation on the second-stage feature vector sequence to obtain a wake-up word recognition result.

[0091] According to an embodiment of the present disclosure, the gate unit includes the following overlapping sub-means:

[0092] The first identification sub-means identifies a current time update gate and current time candidate hidden layer information based on the previous time output vector and the current time input vector, which is the first-stage feature vector for the current time in the first-stage feature vector sequence.

[0093] The second specifying sub-means specifies the current time hidden layer information based on the current time candidate hidden layer information, the previous time hidden layer information, and the current time update gate.

[0094] The third specifying sub-means specifies a current time output vector, which is a second-stage feature vector at the current time in the second-stage feature vector sequence, based on the current time hidden layer information and predetermined parameters.

[0095] According to an embodiment of the present disclosure, the syllable recognition module includes an extraction means and a syllable classification means.

[0096] The extraction means extracts syllable features from the speech to be recognized and obtains a syllable feature matrix.

[0097] The syllable classification means performs a classification operation on the syllable feature matrix to obtain a wake-up syllable recognition result.

[0098] According to an embodiment of the present disclosure, the extracting means includes an extracting sub-means, a dimensionality reducing sub-means, and an encoding sub-means.

[0099] The extraction sub-means extracts features from the speech to be recognized and obtains a feature matrix.

[0100] The dimension reduction sub-means reduces the dimension of the feature matrix to obtain a dimension-reduced feature matrix.

[0101] The encoding sub-means performs a multi-stage speech enhancement encoding process on the dimension-reduced feature matrix to obtain a syllable feature matrix.

[0102] According to an embodiment of the present disclosure, the voice wake-up device further includes a second identification module.

[0103] The second identifying module identifies the voice to be recognized as an incorrect wake-up voice when it determines that the wake-up word recognition result indicates that the voice to be recognized does not contain the predetermined wake-up word.

[0104] According to an embodiment of the present disclosure, the predetermined parameters are identified based on a threshold value of the number of weight-reducing parameters.

[0105] According to an embodiment of the present disclosure, the audio wake-up device further includes a display module and a feedback module.

[0106] If the display module identifies the sound to be recognized as the correct wake-up sound, it displays the target object on the display interface.

[0107] The feedback module outputs a feedback sound.

[0108] According to an embodiment of the present disclosure, the present disclosure provides an electronic device, a readable storage medium, and a computer program. M More to offer.

[0109] According to an embodiment of the present disclosure, an electronic device includes at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions can be executed by the at least one processor to perform a method of an embodiment of the present disclosure.

[0110] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium having computer instructions stored thereon is provided, the computer instructions being used to cause a computer to perform a method of an embodiment of the present disclosure.

[0111] According to an embodiment of the present disclosure , puThe present disclosure also includes a computer program that, when executed by a processor, implements the method of the disclosed embodiment.

[0112] 8 is a schematic block diagram illustrating an example of an electronic device 800 in which embodiments of the present disclosure can be implemented. The electronic device may represent various types of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The electronic device may also represent various types of mobile devices, such as personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely exemplary and do not limit the implementation of the present disclosure as described and / or claimed herein.

[0113] 8, electronic device 800 may include a computing means 801 that performs various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 802 or loaded from a storage means 808 into a random access memory (RAM) 803. RAM 803 may further store various programs and data necessary for the operation of electronic device 800. The computing means 801, ROM 802, and RAM 803 are interconnected by a bus 804. An input / output interface 805 is also connected to the bus 804.

[0114] Multiple components in electronic device 800 are connected to an I / O interface 805, which includes input means 806 such as a keyboard, a mouse, etc., output means 807 such as various types of displays, speakers, etc., storage means 808 such as a magnetic disk, an optical disk, etc., and communication means 809 such as a network card, a modem, a wireless communication transceiver, etc. The communication means 809 enables electronic device 800 to exchange information / data with other devices via a computer network such as the Internet or various types of telecommunications networks.

[0115] The computing means 801 may be a general-purpose and / or dedicated processing module having various processing and computing capabilities. Examples of the computing means 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, computing means for executing various machine learning model algorithms, a digital signal processor (DSP), any suitable processor, controller, microcontroller, etc. The computing means 801 performs the methods and processes described above, such as the voice wake-up method. For example, in one embodiment, the voice wake-up method is implemented as a computer software program, which is temporarily contained in a machine-readable medium, such as the storage means 808. In one embodiment, part or all of the computer program is loaded and / or installed into the electronic device 800 via the ROM 802 and / or the communication means 809. When the computer program is loaded into the RAM 803 and executed by the computing means 801, it may perform one or more steps of the voice wake-up method described above. Alternatively, in other embodiments, the computing means 801 is configured to implement the audio wake-up method in any other suitable manner (eg, firmware).

[0116] Various embodiments of the systems and techniques described herein may be realized in digital electronic circuitry systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may be embodied in one or more computer programs that may be executed and / or interpreted by a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor, and that may receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0117] Program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, so that when the program code is executed by the processor or controller, the functions and operations specified in the flowcharts and / or block diagrams are performed. The program code may be executed entirely on a device, partially on a device, partially on a device as a separate software package, and partially on a remote device, or entirely on a remote device or server.

[0118] In the context of the present disclosure, a machine-readable medium may be a tangible medium, and may contain or store a program for use in or in connection with an instruction execution system, device, or electronic device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or electronic device, or any suitable combination of the above. More specific examples of machine-readable storage media include an electrical connection of one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0119] To provide for user interaction, a computer may implement the systems and techniques described herein and include a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) through which a user may provide input to the computer. Other types of devices may also provide for user interaction; for example, the feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or tactile feedback) and may receive input from the user in any form (including voice input, audio input, or tactile input).

[0120] The systems and techniques described herein may be implemented in a computing system including background components (e.g., a data server), or middleware components (e.g., an application server), or front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or any combination of such background, middleware, or front-end components. The components of the system may be connected to each other by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include, by way of example, a local area network (LAN), a wide area network (WAN), and the Internet.

[0121] The computer system may include a client and a server. The client and server are generally remote and typically interact through a communication network. The relationship between the client and the server is created by computer programs running on the corresponding computers and having a client-server relationship. The server may be a cloud server, a server in a distributed system, or a server in a blockchain combination.

[0122] It should be understood that various types of flows shown above may be used, and operations may be rearranged, added, or deleted. For example, the operations described in this disclosure may be performed in parallel, sequentially, or in a different order, as long as the desired results of the invention of this disclosure are achieved, and this specification is not limited thereto.

[0123] The above specific embodiments do not limit the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications, combinations, subcombinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present disclosure should be included within the scope of protection of the present disclosure.

Claims

1. 1. A voice wake-up method, comprising: performing word recognition on the speech to be recognized and obtaining a wake-up word recognition result; If the wake-up word recognition result indicates that the speech to be recognized includes a predetermined wake-up word, performing syllable recognition on the speech to be recognized to obtain a wake-up syllable recognition result; If the wake-up syllable recognition result indicates that the speech to be recognized contains a predetermined syllable, determining that the speech to be recognized is a correct wake-up speech; performing word recognition on the speech to be recognized and obtaining a wake-up word recognition result; performing a convolution operation on the speech to be recognized, the convolution operation including a sequence of speech frames, to obtain a sequence of first-stage feature vectors that correspond one-to-one to the sequence of speech frames; performing a gated recurrent operation on the first-stage feature vector sequence to obtain a second-stage feature vector sequence; performing a classification operation on the second-stage feature vector sequence to obtain the wake-up word recognition result; performing a gate recurrent operation on the first-stage feature vector sequence to obtain the second-stage feature vector sequence; Identifying a current time update gate and current time candidate hidden layer information based on the previous time output vector and the current time input vector, which is the first-stage feature vector for the current time in the first-stage feature vector sequence; Identifying current time hidden layer information based on the current time candidate hidden layer information, the previous time hidden layer information, and the current time update gate; Identifying a current time output vector, which is a second-stage feature vector for the current time in the second-stage feature vector sequence, based on the current time hidden layer information and predetermined parameters; This includes repeating the operation The predetermined parameter is identified based on a threshold value of the number of weight-reducing parameters. Voice wake-up method.

2. When the wake-up word recognition result indicates that the speech to be recognized includes a predetermined wake-up word, performing syllable recognition on the speech to be recognized to obtain a wake-up syllable recognition result includes: extracting syllable features from the speech to be recognized and obtaining a syllable feature matrix; performing a classification operation on the syllable feature matrix to obtain the wake-up syllable recognition result. The method of claim 1.

3. extracting syllable features from the speech to be recognized and obtaining a syllable feature matrix, extracting features from the speech to be recognized to obtain a feature matrix; performing dimension reduction on the feature matrix to obtain a dimension-reduced feature matrix; and performing a multi-stage speech enhancement encoding process on the dimension-reduced feature matrix to obtain the syllable feature matrix. The method of claim 2.

4. If the wake-up word recognition result indicates that the speech to be recognized does not contain a predetermined wake-up word, the method further includes determining that the speech to be recognized is an incorrect wake-up speech. The method of claim 1.

5. If the sound to be recognized is determined to be a correct wake-up sound, displaying a target object on a display interface; and outputting a feedback sound. The method of claim 1.

6. A voice wake-up device, comprising: a word recognition module that performs word recognition on the speech to be recognized and obtains a wake-up word recognition result; a syllable recognition module that, when the wake-up word recognition result indicates that the speech to be recognized contains a predetermined wake-up word, performs syllable recognition on the speech to be recognized to obtain a wake-up syllable recognition result; a first identification module that identifies the speech to be recognized as a correct wake-up speech when the wake-up syllable recognition result indicates that the speech to be recognized contains a predetermined syllable; The word recognition module convolution means for performing a convolution operation on the speech to be recognized, which includes a speech frame sequence, to obtain a first-stage feature vector sequence that corresponds one-to-one with the speech frame sequence; a gate unit for performing a gate recurrent operation on the first-stage feature vector sequence to obtain a second-stage feature vector sequence; a word classification means for performing a classification operation on the second-stage feature vector sequence to obtain the wake-up word recognition result; The gate unit includes: a first specifying sub-means for specifying a current time update gate and current time candidate hidden layer information based on the previous time output vector and the current time input vector, which is the first stage feature vector at the current time in the first stage feature vector sequence; A second specifying sub-means for specifying current time hidden layer information based on the current time candidate hidden layer information, the previous time hidden layer information, and the current time update gate; a third specifying sub-means for specifying a current time output vector, which is a second-stage feature vector at the current time in the second-stage feature vector sequence, based on the current-time hidden layer information and a predetermined parameter; The predetermined parameter is identified based on a threshold value of the number of weight-reducing parameters. Voice wake-up device.

7. The syllable recognition module an extraction means for extracting syllable features from the speech to be recognized and obtaining a syllable feature matrix; a syllable classification means for performing a classification operation on the syllable feature matrix to obtain the wake-up syllable recognition result.

7. The apparatus of claim 6.

8. The extraction means an extraction sub-means for extracting features from the speech to be recognized and obtaining a feature matrix; a dimension reduction sub-means for performing dimension reduction on the feature matrix and obtaining a dimension-reduced feature matrix; and an encoding sub-means for performing a multi-stage speech enhancement encoding process on the feature matrix after dimension reduction to obtain the syllable feature matrix.

8. The apparatus of claim 7.

9. The system further includes a second identifying module that identifies the voice to be recognized as an incorrect wake-up voice when the wake-up word recognition result indicates that the voice to be recognized does not contain a predetermined wake-up word. An apparatus according to any one of claims 6 to 8.

10. a display module for displaying a target object on a display interface when the sound to be recognized is determined to be a correct wake-up sound; a feedback module that outputs a feedback sound.

7. The apparatus of claim 6.

11. at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor can perform the method of any one of claims 1 to 5. electronic equipment.

12. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: The computer instructions cause a computer to perform the method of any one of claims 1 to 5. storage medium.

13. A computer program comprising: A computer program product which, when executed by a processor, implements the method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Speech recognition method and device and electronic equipment

    CN112420050A

  • Intermediate information analyzer, optimizing device and feature visualizing device of neural network

    JP2019046453A

  • Dynamic threshold for always listening speech trigger

    JP2019091472A

  • System and program for detecting abnormality

    JP2020144626A

  • Server-side hotwording

    JP2020507815A