Voice wake-up method, device, storage medium and electronic device
By extracting cached M-frame status information and decoding, the problem of long voice wake-up delay time is solved, and faster voice wake-up is achieved without losing wake-up performance.
Patent Information
- Application Number
- CN202210705805.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-21
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-06-21
AI Technical Summary
The voice wake-up delay time is long, which leads to an increase in user waiting time and affects user experience.
By determining the character probability vector corresponding to the continuous multi-frame speech, and using the preset pre-wake word as the decoded character path, the cached M-frame status information is extracted and decoded to shorten the wake-up delay time.
There is no need to wait for the input of subsequent frames of voice, save waiting time due to right-looking, shorten wake-up delay time, speed up voice wake-up speed, and maintain wake-up performance.
Smart Images

Figure CN114913853B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of voice wake-up technology, and in particular to a voice wake-up method, device, storage medium and electronic device. Background Art
[0002] Voice wake-up is the first step for users to interact with smart voice devices. When the user says the wake-up word for the smart voice device, the smart voice device wakes up and enters the working state. The delay of voice wake-up determines the experience of human-machine voice interaction between the user and the device. The time from when the user says the wake-up word to when the device responds is the wake-up delay. The lower the wake-up delay, the shorter the user's waiting time, which is conducive to improving the user experience. Generally speaking, when the voice wake-up model recognizes the wake-up word from the user's voice, it generally looks at a few frames of voice to ensure the wake-up effect, which will increase the wake-up delay. In order to reduce the delay, you can choose to reduce the number of frames to look at, but this will cause a certain degree of damage to the wake-up effect. Summary of the invention
[0003] This section is provided to introduce the concepts in a brief form, which will be described in detail in the detailed implementation section below. This section is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to be used to limit the scope of the technical solution claimed for protection.
[0004] In a first aspect, the present disclosure provides a voice wake-up method, comprising:
[0005] Determine the character probability vectors corresponding to the continuous multiple frames of speech; wherein the character probability vector corresponding to each frame of speech is obtained according to the state information of the current frame of speech and the state information of the right M frames of speech;
[0006] In the character probability vector corresponding to the continuous multi-frame speech, a preset pre-wake-up word is used as a decoding character path, and a first path decoding score corresponding to the pre-wake-up word is determined; wherein the pre-wake-up word is composed of the first N characters of the preset complete wake-up word, and M and N are both positive integers;
[0007] In response to the first path decoding score being greater than a first threshold, sequentially extracting M frames of state information currently cached, and obtaining a corresponding character probability vector according to each extracted state information;
[0008] Determine a second path decoding score corresponding to the complete wake-up word according to the path decoding information of the pre-wake-up word and the M character probability vectors corresponding to the M frame state information;
[0009] In response to the second path decoding score being greater than a second threshold, it is determined to wake up the device.
[0010] In a second aspect, the present disclosure provides a voice wake-up device, comprising:
[0011] The first character determination module is used to determine the character probability vector corresponding to the continuous multiple frames of speech; wherein the character probability vector corresponding to each frame of speech is obtained according to the state information of the current frame of speech and the state information of the right M frames of speech;
[0012] A pre-wake-up detection module, configured to use a preset pre-wake-up word as a decoding character path in a character probability vector corresponding to the continuous multi-frame speech, and determine a first path decoding score corresponding to the pre-wake-up word; wherein the pre-wake-up word is composed of the first N characters of the preset complete wake-up word, and M and N are both positive integers;
[0013] A second character determination module is used for extracting the M frames of state information currently cached in sequence in response to the first path decoding score being greater than a first threshold, and obtaining a corresponding character probability vector according to each extracted state information;
[0014] A complete wake-up detection module, used to determine a second path decoding score corresponding to the complete wake-up word according to the path decoding information of the pre-wake-up word and the M character probability vectors corresponding to the M frame state information;
[0015] The wake-up confirmation module is configured to determine to wake up a device in response to the second path decoding score being greater than a second threshold.
[0016] In a third aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.
[0017] In a fourth aspect, the present disclosure provides an electronic device, including:
[0018] a storage device having at least one computer program stored thereon;
[0019] At least one processing device, used to execute the at least one computer program in the storage device to implement the steps of the method described in the first aspect.
[0020] In the above scheme, the pre-wake-up condition is set in advance, that is, the first path decoding score corresponding to the pre-wake-up word needs to be greater than the first threshold. When it is determined that the first path decoding score is greater than the first threshold, there is no need to wait for the input of the next frame of voice. It is only necessary to extract the cached M frame status information, obtain the character probability vectors corresponding to the M frame status information, and then decode based on the character probability vectors corresponding to the M frame status information to determine the second path decoding score corresponding to the complete wake-up word. It is then possible to determine whether to wake up the device based on the second path decoding score. Therefore, this scheme does not need to wait for the input of subsequent frames of voice, saves the waiting time caused by right viewing, and thus shortens the wake-up delay time and speeds up the voice wake-up speed. In addition, since the right viewing of the model in the pre-wake-up stage is also guaranteed, the wake-up performance will not be lost.
[0021] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale. In the drawings:
[0023] Figure 1 A schematic diagram of a target model in an exemplary embodiment is shown;
[0024] Figure 2 A flow chart of a voice wake-up method provided by an exemplary embodiment is shown;
[0025] Figure 3 A flowchart showing the training process of the target model is shown;
[0026] Figure 4 Another flow chart showing the training process of the target model;
[0027] Figure 5 A block diagram of a voice wake-up device provided by an exemplary embodiment is shown;
[0028] Figure 6 A block diagram of an electronic device provided by an exemplary embodiment is shown. DETAILED DESCRIPTION
[0029] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0030] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0031] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0032] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0033] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0034] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0035] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0036] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0037] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0038] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0039] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0040] All actions of acquiring signals, information or data in the present disclosure are carried out in compliance with the relevant data protection laws and policies of the country where they are located, and with the authorization given by the owner of the corresponding device.
[0041] The wake-up delay is partly due to the delay caused by the voice wake-up model (referred to as the target model in this article). Figure 1 A schematic diagram of a target model in an exemplary embodiment of the present disclosure is shown, and the target model includes a first processing unit and a second processing unit, wherein the first processing unit is used to receive an input frame of speech and process the input speech to obtain corresponding status information, and the second processing unit is used to output a corresponding character probability vector based on the one-frame status information and multiple-frame status information viewed to the right of the frame status information.
[0042] For example, taking the target model looking at M frames from the right as an example, the first processing unit processes the input k-th frame of speech to obtain the intermediate state information of the model corresponding to the k-th frame of speech, recorded as state information Sk, and caches the state information Sk, processes the input k+1-th frame of speech to obtain and cache the state information Sk+1 corresponding to the k+1-th frame of speech, processes the input k+2-th frame of speech to obtain and cache the state information Sk+2 corresponding to the k+2-th frame of speech, and so on, processes the input k+M-th frame of speech to obtain and cache the state information Sk+M corresponding to the k+M-th frame of speech, and the second processing unit is used to output a character probability vector Pk corresponding to the k-th frame of speech according to the state information Sk~Sk+M, where Pk={p1,p2,…,pn}, and p1, p2,…,pn respectively represent the probability of each character corresponding to the k-th frame of speech.
[0043] As an example, when the device is started, the voice around the device is collected by the collection device in the device, and the voice collected by the collection device is framed. Assuming that every 10 milliseconds is recorded as a frame of voice, and taking 4 frames on the right as an example, when the first frame of voice is obtained, the first frame of voice is input into the target model, and the first processing unit in the target model processes the first frame of voice to obtain state information S1 corresponding to the first frame of voice, and caches the state information S1. When the second frame of voice is obtained, the second frame of voice is input into the target model, and the first processing unit in the target model processes the second frame of voice to obtain state information S2 corresponding to the second frame of voice, and caches the state information S2. When the third frame of voice is obtained, the corresponding state information S3 is obtained and cached. When the fourth frame of voice is obtained, the corresponding state information S4 is obtained and cached.
[0044] At this time, 4 frames of state information have been cached. When the 5th frame of speech is obtained, the 5th frame of speech is input into the target model, the first processing unit in the target model processes the 5th frame of speech, obtains and caches the state information S5 corresponding to the 5th frame of speech, and also obtains a complete input S1~S5 with a length of 5 frames according to the state information S1 corresponding to the 1st frame of speech and the cached state information S2~S5 corresponding to the 2nd to 5th frames of speech. The second processing unit in the target model outputs the character probability vector corresponding to the 1st frame of speech according to the complete input S1~S5. The character probability vector is obtained according to the state information corresponding to the 1st frame of speech and the cached state information corresponding to the 4 frames of speech viewed to the right of the 1st frame of speech. Therefore, the character probability vector can accurately represent the probability of each character corresponding to the 1st frame of speech.
[0045] Similarly, when the 6th frame of speech is acquired, the corresponding state information S6 is obtained and cached. At the same time, based on the state information S2 corresponding to the 2nd frame of speech and the cached state information S3~S6 corresponding to the 3rd to 6th frames of speech, a complete input S2~S6 with a length of 5 frames is obtained. The second processing unit in the target model outputs the character probability vector corresponding to the 2nd frame of speech based on the complete input S2~S6. The character probability vector is obtained based on the state information corresponding to the 2nd frame of speech and the cached state information corresponding to the 4 frames of speech viewed to the right of the 2nd frame of speech.
[0046] As new voice frames are continuously acquired, the above process will be repeated. After that, decoding will be performed based on multiple consecutive character probability vectors to identify whether they contain the preset wake-up word. If they contain the preset wake-up word, the device will be awakened, otherwise the recognition will continue. The wake-up word refers to the voice command for the user to wake up the device, which can be set by the user or the system.
[0047] It is worth noting that from the above process, it can be seen that the target model will output the first character probability vector only after the fifth frame of speech is input to the target model, that is, when the target model looks right by M frames, the target model will always delay the output by M frames. Due to the right-looking mechanism of the target model, it is necessary to continue to wait for the input of subsequent frames of speech during the speech wake-up process, and this waiting time causes wake-up delay.
[0048] Therefore, the embodiment of the present disclosure provides a voice wake-up method. Figure 2 A flow chart of a voice wake-up method provided by an exemplary embodiment of the present disclosure is shown. Figure 2 As shown, the method includes:
[0049] S110, determining character probability vectors corresponding to a plurality of consecutive speech frames; wherein the character probability vector corresponding to each speech frame is obtained based on the state information of the current speech frame and the state information of the M speech frames to the right.
[0050] In one embodiment, the character probability vector corresponding to each frame of speech can be determined in the following manner: after each frame of speech is acquired, the frame of speech is processed to obtain and cache the state information corresponding to the frame of speech, and for any frame of speech, such as the kth frame of speech, the character probability vector corresponding to the kth frame of speech is obtained based on the state information corresponding to the kth frame of speech and the cached state information corresponding to the k+1th frame to the k+Mth frame of speech; wherein k is a positive integer. In this way, the character probability vectors corresponding to multiple consecutive frames of speech can be obtained.
[0051] S120, in the character probability vector corresponding to the continuous multiple frames of speech, using a preset pre-wake-up word as a decoding character path, and determining a first path decoding score corresponding to the pre-wake-up word.
[0052] The pre-wake-up word consists of the first N characters of the preset complete wake-up word. In the present disclosure, M and N are both positive integers, and M is the number of frames to be viewed right, for example, if 4 frames are viewed right, then M is equal to 4, and N is a positive integer less than the number of characters of the complete wake-up word, for example, if the complete wake-up word consists of four characters, then N can be set to 3. The complete wake-up word refers to the voice command for the user to wake up the device, which can be set by the user or the system.
[0053] S130 , in response to the first path decoding score being greater than the first threshold, sequentially extracting M frames of state information currently cached, and obtaining a corresponding character probability vector according to each extracted state information.
[0054] In response to the first path decoding score corresponding to the pre-wake-up word being greater than the first threshold, it is considered that the pre-wake-up word is detected from a plurality of consecutive character probability vectors, and the pre-wake-up state is entered. After entering the pre-wake-up state, the M frames of state information currently cached are extracted in sequence, and the corresponding character probability vector is obtained according to the state information extracted each time.
[0055] S140, determining a second path decoding score corresponding to the complete wake-up word according to the path decoding information of the pre-wake-up word and the M character probability vectors corresponding to the M frame state information.
[0056] After entering the pre-wake-up state, based on the path decoding information of the pre-wake-up word, according to the M character probability vectors corresponding to the M frame state information, the complete wake-up word is used as the decoding character path to determine the second path decoding score corresponding to the complete wake-up word.
[0057] S150: In response to the second path decoding score being greater than a second threshold, determine to wake up the device.
[0058] In the above scheme, the corresponding pre-wake-up condition is set in advance for entering the pre-wake-up state, that is, the first path decoding score corresponding to the pre-wake-up word needs to be greater than the first threshold. When it is determined that the first path decoding score is greater than the first threshold, the pre-wake-up state is entered. After entering the pre-wake-up state, there is no need to wait for the input of the next frame of voice, but only needs to extract the cached M frame state information to obtain the character probability vectors corresponding to the M frame state information, and then decode based on the character probability vectors corresponding to the M frame state information to determine the second path decoding score corresponding to the complete wake-up word. It can be determined whether to wake up the device based on the second path decoding score. Therefore, this scheme does not need to wait for the input of subsequent frames of voice, saves the waiting time caused by right viewing, and thus shortens the wake-up delay time and speeds up the voice wake-up speed. In addition, since the right viewing of the model in the pre-wake-up stage is also guaranteed, the wake-up performance will not be lost.
[0059] In an exemplary embodiment, in step S110, the speech around the device is collected by a collection device in the device, and the speech collected by the collection device is framed, for example, every 10 milliseconds is recorded as a frame of speech. After each frame of speech is obtained, the frame of speech is input into the first processing unit of the target model, and the input speech is processed by the first processing unit to obtain the state information corresponding to the frame of speech, and the state information is cached. Repeat the above steps to cache the state information corresponding to multiple consecutive frames of speech.
[0060] For any frame of speech, assuming it is the k-th frame of speech, the state information corresponding to the k-th frame of speech and the state information corresponding to the (k + 1)-th to (k + M)-th frames of speech cached are used as the complete input to the second processing unit of the target model. The second processing unit performs character decoding on this complete input and outputs a character probability vector corresponding to the k-th frame of speech. Here, k is a positive integer.
[0061] Repeat the above process to obtain character probability vectors corresponding to multiple consecutive frames of speech.
[0062] Exemplarily, taking looking right at 4 frames as an example, after inputting the 5-th frame of speech to the first processing unit of the target model, a complete input S1 to S5 with a length of 5 frames can be obtained based on the state information S1 corresponding to the 1-st frame of speech and the state information S2 to S5 corresponding to the cached 2-nd to 5-th frames of speech. The second processing unit of the target model performs character decoding on this complete input S1 to S5 and outputs a character probability vector corresponding to the 1-st frame of speech.
[0063] In an exemplary embodiment, in step S120, among the character probability vectors corresponding to multiple consecutive frames of speech, using the pre-awakening word as the decoding character path, the first path decoding score corresponding to the pre-awakening word is determined.
[0064] Exemplarily, taking the Dali table lamp device as an example, the preset complete awakening word is "Dali Dali", and the first three characters of the complete awakening word are set as the pre-awakening word, that is, the pre-awakening word is "Dali Da". Among the character probability vectors corresponding to multiple consecutive frames of speech, using "Da → Li → Da" as the decoding character path, the first path decoding score corresponding to "Dali Da" in the consecutive multiple character probability vectors is determined.
[0065] Optionally, the above decoding process can be implemented using the Viterbi decoding algorithm. The Viterbi decoding algorithm is a dynamic programming algorithm that can perform dynamic programming among consecutive multiple character probability vectors to find the optimal decoding character path of the pre-awakening word, and obtain the first path decoding score corresponding to the pre-awakening word according to the sum of the decoding scores of each decoding character in the optimal decoding character path.
[0066] For example, after dynamic programming, an optimal decoded character path is obtained, which is composed of the character "大" in the probability vectors of the 100th to 120th characters, the character "力" in the probability vectors of the 121st to 140th characters, and the character "大" in the probability vectors of the 141st to 160th characters connected in sequence. According to the above optimal decoded character path, the probabilities of the character "大" in the probability vectors of the 100th to 120th characters are added up as the decoding score of the first character "大" in the pre-awakening word, the probabilities of the character "力" in the probability vectors of the 121st to 140th characters are added up as the decoding score of the second character "力" in the pre-awakening word, and the probabilities of the character "大" in the probability vectors of the 141st to 160th characters are added up as the decoding score of the third character "大" in the pre-awakening word. According to the sum of the decoding scores of each character in the pre-awakening word, the first path decoding score corresponding to the pre-awakening word can be obtained.
[0067] In an exemplary embodiment, in step S130, in response to the first path decoding score corresponding to the pre-awakening word being greater than the first threshold, it is considered that the pre-awakening word is detected from a series of character probability vectors, and then the pre-awakening state is entered. After entering the pre-awakening state, the M-frame state information currently cached is sequentially extracted, and the corresponding character probability vector is obtained according to the state information extracted each time.
[0068] It can be understood that after entering the pre-awakening state, in some embodiments, the target model can output the corresponding character probability vector only based on one frame of state information extracted this time. In other embodiments, in order to match the right-looking processing process of the target model, when processing the M-frame state information currently cached, the right-looking is also performed on multiple frames to ensure the accuracy of the output result. Since the embodiments of the present disclosure do not wait for the input of subsequent frame voices and cannot meet the requirement of right-looking at M frames only based on the M-frame state information currently cached, for each frame of state information, the insufficient part of the right-looking is filled with zeros.
[0069] In specific implementation, the M-frame state information currently cached is sequentially extracted, and for each frame of state information extracted, according to this frame of state information, the state information of the remaining frames in the M-frame state information, and the zero state of the corresponding number of frames, a complete state information with a length of M + 1 frames is obtained. According to this complete state information, the character probability vector corresponding to this frame of state information extracted this time is obtained.
[0070] As an example, assuming that the character probability vector corresponding to the 300th frame of speech output by the target model determines that the first path decoding score corresponding to the pre-wake-up word in multiple consecutive character probability vectors is greater than the first threshold, at this time, the currently cached M frame state information includes state information S301 corresponding to the 301st frame of speech, state information S302 corresponding to the 302nd frame of speech, state information S303 corresponding to the 303rd frame of speech, and state information S304 corresponding to the 304th frame of speech. In response to the first path decoding score corresponding to the pre-wake-up word being greater than the first threshold, the currently cached M frame state information, i.e., state information S301, S302, S303, and S304, is sequentially extracted in chronological order.
[0071] First, extract the state information S301. Since the cached state information S302~S304 is three frames of state information viewed to the right relative to the state information S301, the state information S301, the state information S302~S304 and a frame of zero state will be taken as a complete input and input into the second processing unit of the target model. The second processing unit of the target model outputs the character probability vector corresponding to the 301st frame of speech based on the above input.
[0072] Then, the state information S302 is extracted. Since the cached state information S303~S304 is two frames of state information viewed to the right relative to the state information S302, the state information S302, the state information S303~S304 and the two frames of zero state will be taken as a complete input and input into the second processing unit of the target model. The second processing unit of the target model outputs the character probability vector corresponding to the 302th frame of speech based on the above input.
[0073] Then, the state information S303 is extracted. Since the cached state information S304 is a frame of state information viewed to the right relative to the state information S303, the state information S303, the state information S304 and the three-frame zero state will be taken as a complete input and input into the second processing unit of the target model. The second processing unit of the target model outputs the character probability vector corresponding to the 303th frame of speech based on the above input.
[0074] Then, the state information S304 is extracted. Since the state information of the currently cached M frames does not contain state information that is viewed to the right of the state information S304, the state information S304 and the four-frame zero state are taken as a complete input and input into the second processing unit of the target model. The second processing unit of the target model outputs the character probability vector corresponding to the 304th frame of speech based on the above input.
[0075] Through the above method, one frame of status information can be extracted from the M frames of status information currently cached each time, and a corresponding character probability vector can be output based on the extracted one frame of status information, the multiple frames of status information viewed to the right, and the zero state of the corresponding number of frames, and then one frame of status information can be further extracted and a corresponding character probability vector can be output.
[0076] In this way, for status information S301, there is no need to wait for status information corresponding to the 305th frame of speech, for status information S302, there is no need to wait for status information corresponding to the 305th to 306th frames of speech, for status information S303, there is no need to wait for status information corresponding to the 305th to 307th frames of speech, and for status information S304, there is no need to wait for status information corresponding to the 305th to 308th frames of speech, thereby saving the time for waiting for the 305th to 308th frames of speech input, reducing the wake-up delay, and at the same time ensuring the right viewing of the model to the greatest extent, reducing the loss of wake-up performance.
[0077] In an exemplary embodiment, in step S140, after each corresponding character probability vector is obtained, the remaining characters in the complete wake-up word except the pre-wake-up word are used as decoding characters according to the path decoding information of the pre-wake-up word and the corresponding character probability vector that has been obtained, and the decoding scores corresponding to the remaining characters are determined. The second path decoding score corresponding to the complete wake-up word is determined according to the first path decoding score corresponding to the pre-wake-up word and the decoding scores corresponding to the remaining characters.
[0078] It can be understood that in the process of decoding based on multiple consecutive character probability vectors and taking the pre-wake-up word as the decoding character path, the path decoding information of the pre-wake-up word will be synchronously recorded. Therefore, in the above steps, it is only necessary to decode the remaining characters in the complete wake-up word based on the path decoding information of the pre-wake-up word. The second path decoding score is the sum of the first path decoding score and the decoding scores of the remaining characters.
[0079] In an exemplary embodiment, in step S150, in response to the second path decoding score being greater than the second threshold, it is considered that a complete wake-up word is detected in the voice, and the device can be woken up.
[0080] Furthermore, in some cases, the user did not actually say the pre-wake-up word, but due to reasons such as inaccurate output results of the target model, the result of the first path decoding score being greater than the first threshold was erroneously obtained when the pre-wake-up word was not actually included, and the pre-wake-up state was erroneously entered, resulting in a misjudgment.
[0081] In order to reduce the occurrence of the above-mentioned misjudgment, in step S130, in response to the first path decoding score being greater than the first threshold, the pronunciation duration corresponding to the pre-wake-up word is determined according to the path decoding information of the pre-wake-up word, and it is determined whether the pronunciation duration is within a preset duration range. The preset minimum value and maximum value within the preset duration range respectively represent the shortest time and the longest time required to pronounce the pre-wake-up word. When the pronunciation duration is within the preset duration range, the steps of sequentially extracting the currently cached M frame status information and obtaining the corresponding character probability vector according to the status information extracted each time are performed. If the pronunciation duration is not within the preset duration range, the above steps are not performed, and it is necessary to continue to obtain the speech of subsequent frames and continue to perform the above steps S110 to S120.
[0082] It can be understood that in the process of decoding according to multiple consecutive character probability vectors and taking the pre-wake-up word as the decoding character path, the path decoding information of the pre-wake-up word will be synchronously recorded. The path decoding information includes the boundary time points corresponding to each character in the pre-wake-up word. According to the starting boundary time point of the first character in the pre-wake-up word and the ending boundary time point of the last character, the pronunciation duration corresponding to the pre-wake-up word can be obtained.
[0083] In the above scheme, a preset duration range is pre-set, and the minimum and maximum values of the preset duration range respectively represent the shortest time and the longest time required to pronounce the pre-wake-up word, and the minimum and maximum values can be empirical values preset according to actual tests. According to actual tests, assuming that the time required for the user to pronounce the pre-wake-up word "Da Li Da" is generally about 0.3 seconds at the fastest and about 2 seconds at the slowest, then the preset duration range can be set to [0.3, 2], so that the pronunciation duration of the pre-wake-up word is limited to this range. When the first path decoding score corresponding to the pre-wake-up word is greater than the first threshold, it is necessary to further determine whether the pronunciation duration of the pre-wake-up word is within the range of [0.3, 2] based on the path decoding information of the pre-wake-up word. If it is not within the range of [0.3, 2], it indicates that the result that the first path decoding score obtained at this time is greater than the first threshold may be caused by the erroneous output of the target model, so the current cached M frame state information is not extracted. If it is within the range of [0.3, 2], the current cached M frame state information is extracted.
[0084] The present disclosure can avoid mistakenly entering the pre-wake-up state by setting a limit range for the pronunciation duration of the pre-wake-up word, thereby avoiding false wake-up to a certain extent.
[0085] Furthermore, in some cases, the user did not say the complete wake-up word, but due to reasons such as inaccurate output results of the target model, the second path decoding score was mistakenly obtained to be greater than the second threshold when the complete wake-up word was not actually included, and the device was mistakenly woken up, resulting in a misjudgment.
[0086] In order to reduce the occurrence of the above-mentioned misjudgment, in the above-mentioned step S150, in response to the second path decoding score being greater than the second threshold, it is determined whether the decoding scores corresponding to the remaining characters in the second path decoding score are greater than the third threshold, and in response to the decoding scores corresponding to the remaining characters being greater than the third threshold, it is determined to wake up the device.
[0087] That is to say, even if the second path decoding score corresponding to the complete wake-up word is greater than the second threshold, the device can only be woken up if the decoding scores corresponding to the remaining characters are also greater than the third threshold. If the decoding scores corresponding to the remaining characters are not greater than the third threshold, the device still cannot be woken up, thereby effectively avoiding false wake-ups caused by misjudgment of the target model.
[0088] Further, it can be seen from the above step S130 that in the process of extracting the M frame status information currently cached, for each frame status information extracted, the frame status information, the status information of the remaining frames in the cached M frame status information, and the zero state of the corresponding number of frames need to be used as a complete input of the second processing unit of the target model. In this complete input, the number of frames that are effectively viewed to the right relative to the frame status information does not actually reach M frames, and the zero state of the corresponding number of frames is included. For example, in the aforementioned example, for the status information S301 extracted from the M frame status information, 3 frames are effectively viewed to the right, for the status information S302 extracted from the M frame status information, 2 frames are effectively viewed to the right, for the status information S303 extracted from the M frame status information, 1 frame is effectively viewed to the right, and for the status information S304 extracted from the M frame status information, 0 frames are effectively viewed to the right.
[0089] That is to say, in the voice wake-up method disclosed in the present invention, the target model actually dynamically looks right between frames 0 and 4. In order to match this processing process of the target model, the target model also needs to learn this dynamic right-looking processing method when training it, so that the trained target model can accurately output results when looking right at different frame numbers.
[0090] Figure 3 A training process provided by an exemplary embodiment is shown, by performing the following steps on an original model: Figure 3 As shown in the training, to obtain the target model. Figure 3 As shown, this training process includes:
[0091] S210, obtaining the original model and training corpus, and dividing the training corpus into multiple frames of training speech.
[0092] S220, for each frame of training speech in the multiple frames of training speech, input the frame of training speech into the original model, process the input training speech through the original model, and obtain and cache state information corresponding to the frame of training speech.
[0093] S230, obtaining complete state information of a length of M+1 frames according to the state information corresponding to the i-th frame of training speech in the multiple frames of training speech and the state information corresponding to the i+1-th frame to the i+M-th frame of training speech in the buffer.
[0094] Wherein, i is a positive integer.
[0095] S240, randomly cover frames 0 to M in the status information of the last M frames of the complete status information to obtain covered status information.
[0096] S250, outputting a character probability vector corresponding to the i-th frame of training speech according to the covering state information through the original model.
[0097] S260, updating the parameters of the original model according to the character probability vector corresponding to each frame of training speech and the label information corresponding to the training corpus;
[0098] S270, using the trained original model as the target model.
[0099] In the above scheme, the covered state information is obtained by randomly masking frames 0 to M in the state information of the last M frames of the complete state information, and then the original model outputs the corresponding character probability vector according to the covered state information. In this way, it is equivalent to simulating the actual voice wake-up process, in which the target model dynamically looks at different numbers of frames according to a frame of state information extracted from the currently cached M frames of state information. The model is trained based on such input, so that the trained original model, that is, the target model, can output accurate results when looking at 0 frames, 1 frame, ..., M frames.
[0100] Figure 4 A training process provided by an exemplary embodiment is shown, by performing the following steps on an original model: Figure 4 As shown in the training, to obtain the target model. Figure 4 As shown, this training process includes:
[0101] S310, obtaining the original model and the training corpus containing the complete wake-up word, truncating the pronunciation of the remaining characters in the training corpus by random length, and dividing the truncated training corpus into multiple frames of training speech.
[0102] The random length can be between "0 and all", 0 means that the pronunciation of the remaining characters is not truncated, and all means that the pronunciation of the remaining characters is truncated.
[0103] S320, for each frame of training speech in the multiple frames of training speech, input the frame of training speech into the original model, process the input training speech through the original model, and obtain and cache state information corresponding to the frame of training speech.
[0104] S330, obtaining complete state information of a length of M+1 frames according to the state information corresponding to the j-th frame of training speech in the multi-frame training speech and the state information corresponding to the j+1-th frame to the j+M-th frame of training speech in the cache.
[0105] Wherein, j is a positive integer.
[0106] S340, outputting the character probability vector corresponding to the j-th frame of training speech according to the complete state information through the original model.
[0107] S350, updating the parameters of the original model according to the character probability vector corresponding to each frame of training speech and the label information corresponding to the training corpus;
[0108] S360, takes the trained original model as the target model.
[0109] In the above scheme, the pronunciation of the remaining characters in the complete wake-up word in the training corpus is truncated with random length. After truncation, the model can only be trained based on the remaining part of the training corpus that has not been truncated. When processing the speech frames of the remaining characters, since the pronunciation of the remaining characters has been randomly truncated, it is impossible to fully view the pronunciation of the remaining characters for M frames. For those that are less than M frames, they can only be supplemented with zeros.
[0110] Compared with the previous exemplary embodiment, this embodiment simulates the process of dynamically looking right on the speech frames of the remaining characters by forcibly truncating the training corpus itself.
[0111] It is understandable that in Figure 3 and Figure 4 In the training process shown, the various processing performed by the original model can refer to the description of the target model in the aforementioned embodiment, and will not be repeated here.
[0112] Figure 5 A block diagram of a voice wake-up device provided by an exemplary embodiment is shown. Figure 5 , the device 400 comprises:
[0113] The first character determination module 410 is used to determine the character probability vector corresponding to the continuous multiple frames of speech; wherein the character probability vector corresponding to each frame of speech is obtained according to the state information of the current frame of speech and the state information of the right M frames of speech;
[0114] A pre-wake-up detection module 420 is used to determine a first path decoding score corresponding to a pre-wake-up word in a character probability vector corresponding to the continuous multi-frame speech, using the pre-wake-up word as a decoding character path; wherein the pre-wake-up word is composed of the first N characters of the pre-set complete wake-up word, and M and N are both positive integers;
[0115] A second character determination module 430 is configured to extract the M frames of state information currently cached in sequence in response to the first path decoding score being greater than a first threshold, and obtain a corresponding character probability vector according to each extracted state information;
[0116] A complete wake-up detection module 440 is used to determine a second path decoding score corresponding to the complete wake-up word according to the path decoding information of the pre-wake-up word and the M character probability vectors corresponding to the M frame state information;
[0117] The wake-up confirmation module 450 is configured to determine to wake up a device in response to the second path decoding score being greater than a second threshold.
[0118] Optionally, the device 400 further includes a character recognition module, which is used to:
[0119] After each frame of speech is acquired, the frame of speech is processed to obtain and cache the state information corresponding to the frame of speech;
[0120] According to the state information corresponding to the kth frame of speech and the cached state information corresponding to the k+1th frame to the k+Mth frame of speech, a character probability vector corresponding to the kth frame of speech is obtained; wherein k is a positive integer.
[0121] Optionally, the second character determination module 430 includes:
[0122] The information extraction subunit is used to extract the M frames of status information currently cached in sequence, and for each frame of status information extracted, obtain complete status information with a length of M+1 frames according to the frame status information, the status information of the remaining frames in the M frame status information, and the zero state of the corresponding frame number;
[0123] The character recognition subunit is used to obtain a character probability vector corresponding to the frame state information extracted this time according to the complete state information.
[0124] Optionally, the second character determination module 430 includes:
[0125] a duration determination subunit, configured to determine, in response to the first path decoding score being greater than a first threshold, a pronunciation duration corresponding to the pre-wake-up word according to the path decoding information of the pre-wake-up word;
[0126] A range detection subunit, used to determine whether the pronunciation duration is within a preset duration range, wherein a preset minimum value and a preset maximum value within the preset duration range respectively represent the shortest time and the longest time required to pronounce the pre-wake-up word;
[0127] The result jump subunit is used to control the execution of the steps of sequentially extracting the M-frame state information of the current cache and obtaining the corresponding character probability vector according to the state information extracted each time when the pronunciation duration is within the preset duration range.
[0128] Optionally, the complete wake-up detection module 440 includes:
[0129] a decoding score determination subunit, configured to, after each acquisition of a corresponding character probability vector, determine the decoding scores corresponding to the remaining characters in the complete wake-up word except the pre-wake-up word according to the path decoding information of the pre-wake-up word and the acquired corresponding character probability vector, and use the remaining characters in the complete wake-up word as decoding characters;
[0130] The path score determination subunit is used to determine the second path decoding score corresponding to the complete wake-up word according to the first path decoding score corresponding to the pre-wake-up word and the decoding scores corresponding to the remaining characters.
[0131] Optionally, the wake-up confirmation module 450 includes:
[0132] a decoding score determination subunit, configured to determine, in response to the second path decoding score being greater than a second threshold, whether the decoding scores corresponding to the remaining characters in the second path decoding score are greater than a third threshold;
[0133] The wake-up confirmation subunit is used to determine to wake up the device in response to the decoding score corresponding to the remaining characters being greater than a third threshold.
[0134] Optionally, the status information and the character probability vector are obtained by processing a target model, and the target model is used to receive a frame of speech and process the frame of speech, output the status information of the frame of speech, and also to output a corresponding character probability vector based on a frame of status information and M frames of status information viewed to the right of the frame of status information.
[0135] It should be noted that the above modules in the device 400 are configured in an electronic device capable of voice wake-up.
[0136] Optionally, the device 400 further includes a training module for training an original model to obtain a target model, wherein the training module may be configured in another electronic device.
[0137] Optionally, the training module is used to perform the following training process:
[0138] Obtaining an original model and training corpus, and dividing the training corpus into multiple frames of training speech;
[0139] For each frame of training speech in the multiple frames of training speech, the frame of training speech is input into the original model, the input training speech is processed by the original model, and state information corresponding to the frame of training speech is obtained and cached;
[0140] According to the state information corresponding to the i-th frame of training speech in the multi-frame training speech and the state information corresponding to the i+1-th frame to the i+M-th frame of training speech in the cache, complete state information with a length of M+1 frames is obtained; wherein i is a positive integer;
[0141] Randomly masking frames 0 to M in the state information of the last M frames of the complete state information to obtain masked state information;
[0142] Outputting a character probability vector corresponding to the i-th frame of training speech according to the covering state information through the original model;
[0143] Update the parameters of the original model according to the character probability vector corresponding to each frame of training speech and the label information corresponding to the training corpus;
[0144] The trained original model is used as the target model.
[0145] Optionally, the training module is used to perform the following training process:
[0146] Obtaining an original model and a training corpus containing the complete wake-up word;
[0147] Performing random length truncations on the pronunciations of the remaining characters in the training corpus, and dividing the truncated training corpus into multiple frames of training speech;
[0148] For each frame of training speech in the multiple frames of training speech, the frame of training speech is input into the original model, the input training speech is processed by the original model, and state information corresponding to the frame of training speech is obtained and cached;
[0149] According to the state information corresponding to the j-th frame of training speech in the multi-frame training speech and the state information corresponding to the j+1-th frame to the j+M-th frame of training speech in the cache, complete state information with a length of M+1 frames is obtained; wherein j is a positive integer;
[0150] Outputting a character probability vector corresponding to the j-th frame of training speech according to the complete state information through the original model;
[0151] Update the parameters of the original model according to the character probability vector corresponding to each frame of training speech and the label information corresponding to the training corpus;
[0152] The trained original model is used as the target model.
[0153] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0154] Reference below Figure 6 , which shows a schematic diagram of the structure of an electronic device 500 suitable for implementing the embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include but is not limited to a mobile phone, a laptop computer, a digital broadcast receiver, a PDA (personal digital assistant), a PAD (tablet computer), a PMP (portable multimedia player), a vehicle terminal, a digital TV, a desktop computer, a desk lamp device, a smart speaker, etc. Through this technical solution, the electronic device 500 can be quickly awakened by the user's voice. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0155] like Figure 6 As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0156] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 500 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0157] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0158] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0159] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0160] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: determines the character probability vector corresponding to multiple consecutive frames of speech; wherein the character probability vector corresponding to each frame of speech is obtained based on the status information of the current frame of speech and the status information of the M frames of speech viewed to the right; in the character probability vector corresponding to the multiple consecutive frames of speech, a preset pre-wake-up word is used as a decoding character path to determine the first path decoding score corresponding to the pre-wake-up word; wherein the pre-wake-up word is composed of the first N characters of the preset complete wake-up word, and M and N are both positive integers; in response to the first path decoding score being greater than a first threshold, the currently cached M frames of status information are extracted in sequence, and the corresponding character probability vector is obtained according to each extracted status information; according to the path decoding information of the pre-wake-up word and the M character probability vectors corresponding to the M frame status information, the second path decoding score corresponding to the complete wake-up word is determined; in response to the second path decoding score being greater than a second threshold, the device is determined to be woken up.
[0161] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0162] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0163] The modules involved in the embodiments described in the present disclosure may be implemented by software or hardware. The name of the module does not limit the module itself in some cases. For example, the first character determination module may also be described as a "module for determining the character probability vector corresponding to multiple consecutive frames of speech."
[0164] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0165] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0166] According to one or more embodiments of the present disclosure, Example 1 provides a voice wake-up method, including:
[0167] Determine the character probability vectors corresponding to the continuous multiple frames of speech; wherein the character probability vector corresponding to each frame of speech is obtained according to the state information of the current frame of speech and the state information of the right M frames of speech;
[0168] In the character probability vector corresponding to the continuous multi-frame speech, a preset pre-wake-up word is used as a decoding character path, and a first path decoding score corresponding to the pre-wake-up word is determined; wherein the pre-wake-up word is composed of the first N characters of the preset complete wake-up word, and M and N are both positive integers;
[0169] In response to the first path decoding score being greater than a first threshold, sequentially extracting M frames of state information currently cached, and obtaining a corresponding character probability vector according to each extracted state information;
[0170] Determine a second path decoding score corresponding to the complete wake-up word according to the path decoding information of the pre-wake-up word and the M character probability vectors corresponding to the M frame state information;
[0171] In response to the second path decoding score being greater than a second threshold, it is determined to wake up the device.
[0172] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, wherein the method further includes:
[0173] After each frame of speech is acquired, the frame of speech is processed to obtain and cache the state information corresponding to the frame of speech;
[0174] According to the state information corresponding to the kth frame of speech and the cached state information corresponding to the k+1th frame to the k+Mth frame of speech, a character probability vector corresponding to the kth frame of speech is obtained; wherein k is a positive integer.
[0175] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 1, wherein sequentially extracting the M-frame state information of the current cache, and obtaining the corresponding character probability vector according to the state information extracted each time, includes:
[0176] Extracting the M frames of status information currently cached in sequence, and for each frame of status information extracted, obtaining complete status information with a length of M+1 frames according to the frame status information, the status information of the remaining frames in the M frames of status information, and the zero state of the corresponding frame number;
[0177] According to the complete state information, a character probability vector corresponding to the frame state information extracted this time is obtained.
[0178] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 1, wherein in response to the first path decoding score being greater than a first threshold, sequentially extracting the M-frame state information of the current cache, and obtaining the corresponding character probability vector according to each extracted state information, including:
[0179] In response to the first path decoding score being greater than a first threshold, determining a pronunciation duration corresponding to the pre-wake-up word according to the path decoding information of the pre-wake-up word;
[0180] Determine whether the pronunciation duration is within a preset duration range, wherein a preset minimum value and a preset maximum value within the preset duration range respectively represent the shortest time and the longest time required to pronounce the pre-wake-up word;
[0181] If the pronunciation duration is within the preset duration range, the step of sequentially extracting the M frames of state information currently cached and obtaining the corresponding character probability vector according to the state information extracted each time is performed.
[0182] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 1, wherein determining the second path decoding score corresponding to the complete wake-up word according to the path decoding information of the pre-wake-up word and the M character probability vectors corresponding to the M frame state information includes:
[0183] After each corresponding character probability vector is obtained, according to the path decoding information of the pre-wake-up word and the corresponding character probability vector that has been obtained, the remaining characters in the complete wake-up word except the pre-wake-up word are used as decoding characters, and the decoding scores corresponding to the remaining characters are determined;
[0184] A second path decoding score corresponding to the complete wake-up word is determined according to the first path decoding score corresponding to the pre-wake-up word and the decoding scores corresponding to the remaining characters.
[0185] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 5, wherein in response to the second path decoding score being greater than a second threshold, determining to wake up the device includes:
[0186] In response to the second path decoding score being greater than a second threshold, determining whether the decoding scores corresponding to the remaining characters in the second path decoding score are greater than a third threshold;
[0187] In response to the decoding scores corresponding to the remaining characters being greater than a third threshold, it is determined to wake up the device.
[0188] According to one or more embodiments of the present disclosure, Example 7 provides the method of Example 3, wherein the state information and the character probability vector are obtained by processing a target model, the target model is used to receive a frame of speech and process the frame of speech, output the state information of the frame of speech, and also to output a corresponding character probability vector based on a frame of state information and M frames of state information viewed to the right of the frame of state information.
[0189] According to one or more embodiments of the present disclosure, Example 8 provides the method of Example 7, wherein the target model is trained by the following training process:
[0190] Obtaining an original model and training corpus, and dividing the training corpus into multiple frames of training speech;
[0191] For each frame of training speech in the multiple frames of training speech, the frame of training speech is input into the original model, the input training speech is processed by the original model, and state information corresponding to the frame of training speech is obtained and cached;
[0192] According to the state information corresponding to the i-th frame of training speech in the multi-frame training speech and the state information corresponding to the i+1-th frame to the i+M-th frame of training speech in the cache, complete state information with a length of M+1 frames is obtained; wherein i is a positive integer;
[0193] Randomly masking frames 0 to M in the state information of the last M frames of the complete state information to obtain masked state information;
[0194] Outputting a character probability vector corresponding to the i-th frame of training speech according to the covering state information through the original model;
[0195] Update the parameters of the original model according to the character probability vector corresponding to each frame of training speech and the label information corresponding to the training corpus;
[0196] The trained original model is used as the target model.
[0197] According to one or more embodiments of the present disclosure, Example 9 provides the method of Example 7, wherein the target model is trained by the following training process:
[0198] Obtaining an original model and a training corpus containing the complete wake-up word;
[0199] Performing random length truncations on the pronunciations of the remaining characters in the training corpus, and dividing the truncated training corpus into multiple frames of training speech;
[0200] For each frame of training speech in the multiple frames of training speech, the frame of training speech is input into the original model, the input training speech is processed by the original model, and state information corresponding to the frame of training speech is obtained and cached;
[0201] According to the state information corresponding to the j-th frame of training speech in the multi-frame training speech and the state information corresponding to the j+1-th frame to the j+M-th frame of training speech in the cache, complete state information with a length of M+1 frames is obtained; wherein j is a positive integer;
[0202] Outputting a character probability vector corresponding to the j-th frame of training speech according to the complete state information through the original model;
[0203] Update the parameters of the original model according to the character probability vector corresponding to each frame of training speech and the label information corresponding to the training corpus;
[0204] The trained original model is used as the target model.
[0205] According to one or more embodiments of the present disclosure, Example 10 provides a voice wake-up device, including:
[0206] The first character determination module is used to determine the character probability vector corresponding to the continuous multiple frames of speech; wherein the character probability vector corresponding to each frame of speech is obtained according to the state information of the current frame of speech and the state information of the right M frames of speech;
[0207] A pre-wake-up detection module, configured to use a preset pre-wake-up word as a decoding character path in a character probability vector corresponding to the continuous multi-frame speech, and determine a first path decoding score corresponding to the pre-wake-up word; wherein the pre-wake-up word is composed of the first N characters of the preset complete wake-up word, and M and N are both positive integers;
[0208] A second character determination module is used for extracting the M frames of state information currently cached in sequence in response to the first path decoding score being greater than a first threshold, and obtaining a corresponding character probability vector according to each extracted state information;
[0209] A complete wake-up detection module, used to determine a second path decoding score corresponding to the complete wake-up word according to the path decoding information of the pre-wake-up word and the M character probability vectors corresponding to the M frame state information;
[0210] The wake-up confirmation module is configured to determine to wake up a device in response to the second path decoding score being greater than a second threshold.
[0211] According to one or more embodiments of the present disclosure, Example 11 provides the device of Example 10, wherein the device further includes a character recognition module, configured to:
[0212] After each frame of speech is acquired, the frame of speech is processed to obtain and cache the state information corresponding to the frame of speech;
[0213] According to the state information corresponding to the kth frame of speech and the cached state information corresponding to the k+1th frame to the k+Mth frame of speech, a character probability vector corresponding to the kth frame of speech is obtained; wherein k is a positive integer.
[0214] According to one or more embodiments of the present disclosure, Example 12 provides the apparatus of Example 10, wherein the second character determination module includes:
[0215] The information extraction subunit is used to extract the M frames of status information currently cached in sequence, and for each frame of status information extracted, obtain complete status information with a length of M+1 frames according to the frame status information, the status information of the remaining frames in the M frame status information, and the zero state of the corresponding frame number;
[0216] The character recognition subunit is used to obtain a character probability vector corresponding to the frame state information extracted this time according to the complete state information.
[0217] According to one or more embodiments of the present disclosure, Example 13 provides the apparatus of Example 10, wherein the second character determination module includes:
[0218] a duration determination subunit, configured to determine, in response to the first path decoding score being greater than a first threshold, a pronunciation duration corresponding to the pre-wake-up word according to the path decoding information of the pre-wake-up word;
[0219] A range detection subunit, used to determine whether the pronunciation duration is within a preset duration range, wherein a preset minimum value and a preset maximum value within the preset duration range respectively represent the shortest time and the longest time required to pronounce the pre-wake-up word;
[0220] The result jump subunit is used to control the execution of the steps of sequentially extracting the M-frame state information of the current cache and obtaining the corresponding character probability vector according to the state information extracted each time when the pronunciation duration is within the preset duration range.
[0221] According to one or more embodiments of the present disclosure, Example 14 provides the apparatus of Example 10, wherein the complete wake-up detection module includes:
[0222] a decoding score determination subunit, configured to, after each acquisition of a corresponding character probability vector, determine the decoding scores corresponding to the remaining characters in the complete wake-up word except the pre-wake-up word according to the path decoding information of the pre-wake-up word and the acquired corresponding character probability vector, and use the remaining characters in the complete wake-up word as decoding characters;
[0223] The path score determination subunit is used to determine the second path decoding score corresponding to the complete wake-up word according to the first path decoding score corresponding to the pre-wake-up word and the decoding scores corresponding to the remaining characters.
[0224] According to one or more embodiments of the present disclosure, Example 15 provides the apparatus of Example 14, and the wake-up confirmation module includes:
[0225] a decoding score determination subunit, configured to determine, in response to the second path decoding score being greater than a second threshold, whether the decoding scores corresponding to the remaining characters in the second path decoding score are greater than a third threshold;
[0226] The wake-up confirmation subunit is used to determine to wake up the device in response to the decoding score corresponding to the remaining characters being greater than a third threshold.
[0227] According to one or more embodiments of the present disclosure, Example 16 provides the device of Example 12, wherein the state information and the character probability vector are obtained by processing a target model, the target model is used to receive a frame of speech and process the frame of speech, output the state information of the frame of speech, and also to output a corresponding character probability vector based on a frame of state information and M frames of state information viewed to the right of the frame of state information.
[0228] According to one or more embodiments of the present disclosure, Example 17 provides the apparatus of Example 16, wherein the apparatus further includes a training module for training an original model to obtain a target model, and the training module is used to perform the following training process:
[0229] Obtaining an original model and training corpus, and dividing the training corpus into multiple frames of training speech;
[0230] For each frame of training speech in the multiple frames of training speech, the frame of training speech is input into the original model, the input training speech is processed by the original model, and state information corresponding to the frame of training speech is obtained and cached;
[0231] According to the state information corresponding to the i-th frame of training speech in the multi-frame training speech and the state information corresponding to the i+1-th frame to the i+M-th frame of training speech in the cache, complete state information with a length of M+1 frames is obtained; wherein i is a positive integer;
[0232] Randomly masking frames 0 to M in the state information of the last M frames of the complete state information to obtain masked state information;
[0233] Outputting a character probability vector corresponding to the i-th frame of training speech according to the covering state information through the original model;
[0234] Update the parameters of the original model according to the character probability vector corresponding to each frame of training speech and the label information corresponding to the training corpus;
[0235] The trained original model is used as the target model.
[0236] According to one or more embodiments of the present disclosure, Example 18 provides the apparatus of Example 16, wherein the apparatus further includes a training module for training an original model to obtain a target model, and the training module is used to perform the following training process:
[0237] Obtaining an original model and a training corpus containing the complete wake-up word;
[0238] Performing random length truncations on the pronunciations of the remaining characters in the training corpus, and dividing the truncated training corpus into multiple frames of training speech;
[0239] For each frame of training speech in the multiple frames of training speech, the frame of training speech is input into the original model, the input training speech is processed by the original model, and state information corresponding to the frame of training speech is obtained and cached;
[0240] According to the state information corresponding to the j-th frame of training speech in the multi-frame training speech and the state information corresponding to the j+1-th frame to the j+M-th frame of training speech in the cache, complete state information with a length of M+1 frames is obtained; wherein j is a positive integer;
[0241] Outputting a character probability vector corresponding to the j-th frame of training speech according to the complete state information through the original model;
[0242] Update the parameters of the original model according to the character probability vector corresponding to each frame of training speech and the label information corresponding to the training corpus;
[0243] The trained original model is used as the target model.
[0244] According to one or more embodiments of the present disclosure, Example 19 provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the method described in any one of Examples 1-9 when executed by a processing device.
[0245] According to one or more embodiments of the present disclosure, Example 20 provides an electronic device, including:
[0246] a storage device having at least one computer program stored thereon;
[0247] At least one processing device is used to execute the at least one computer program in the storage device to implement the steps of the method described in any one of Examples 1-9.
[0248] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.
[0249] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0250] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.
Claims
1. A voice wake-up method, characterized in that: include: Determine the character probability vectors corresponding to the continuous multiple frames of speech; wherein the character probability vector corresponding to each frame of speech is obtained according to the state information of the current frame of speech and the state information of the right M frames of speech; In the character probability vector corresponding to the continuous multi-frame speech, a preset pre-wake-up word is used as a decoding character path to determine a first path decoding score corresponding to the pre-wake-up word; wherein the pre-wake-up word is composed of the first N characters of the preset complete wake-up word, and M and N are both positive integers; In response to the first path decoding score being greater than a first threshold, sequentially extracting M frames of state information currently cached, and obtaining a corresponding character probability vector according to each extracted state information; Determine a second path decoding score corresponding to the complete wake-up word according to the path decoding information of the pre-wake-up word and the M character probability vectors corresponding to the M frame state information; In response to the second path decoding score being greater than a second threshold, determining to wake up a device; The determining, according to the path decoding information of the pre-wake-up word and the M character probability vectors corresponding to the M frame state information, a second path decoding score corresponding to the complete wake-up word includes: After each corresponding character probability vector is obtained, according to the path decoding information of the pre-wake-up word and the corresponding character probability vector that has been obtained, the remaining characters in the complete wake-up word except the pre-wake-up word are used as decoding characters, and the decoding scores corresponding to the remaining characters are determined; A second path decoding score corresponding to the complete wake-up word is determined according to the first path decoding score corresponding to the pre-wake-up word and the decoding scores corresponding to the remaining characters.
2. The method according to claim 1, characterized in that The method further comprises: After each frame of speech is acquired, the frame of speech is processed to obtain and cache the state information corresponding to the frame of speech; According to the state information corresponding to the kth frame of speech and the cached state information corresponding to the k+1th frame to the k+Mth frame of speech, a character probability vector corresponding to the kth frame of speech is obtained; wherein k is a positive integer.
3. The method according to claim 1, characterized in that The step of sequentially extracting the M frames of state information currently cached and obtaining the corresponding character probability vector according to the state information extracted each time includes: Extracting the M frames of status information currently cached in sequence, and for each frame of status information extracted, obtaining complete status information with a length of M+1 frames according to the frame status information, the status information of the remaining frames in the M frames of status information, and the zero state of the corresponding frame number; According to the complete state information, a character probability vector corresponding to the frame state information extracted this time is obtained.
4. The method according to claim 1, characterized in that: In response to the first path decoding score being greater than a first threshold, sequentially extracting M frames of state information currently cached, and obtaining a corresponding character probability vector according to each extracted state information, including: In response to the first path decoding score being greater than a first threshold, determining a pronunciation duration corresponding to the pre-wake-up word according to the path decoding information of the pre-wake-up word; Determine whether the pronunciation duration is within a preset duration range, wherein a preset minimum value and a preset maximum value within the preset duration range respectively represent the shortest time and the longest time required to pronounce the pre-wake-up word; If the pronunciation duration is within the preset duration range, the step of sequentially extracting the M frames of state information currently cached and obtaining the corresponding character probability vector according to the state information extracted each time is performed.
5. The method according to claim 1, characterized in that In response to the second path decoding score being greater than a second threshold, determining to wake up the device includes: In response to the second path decoding score being greater than a second threshold, determining whether the decoding scores corresponding to the remaining characters in the second path decoding score are greater than a third threshold; In response to the decoding scores corresponding to the remaining characters being greater than a third threshold, it is determined to wake up the device.
6. The method according to claim 3, characterized in that The state information and the character probability vector are obtained by processing a target model, and the target model is used to receive a frame of speech and process the frame of speech, output the state information of the frame of speech, and also to output a corresponding character probability vector based on a frame of state information and M frames of state information viewed to the right of the frame of state information.
7. The method according to claim 6, characterized in that The target model is trained through the following training process: Obtaining an original model and training corpus, and dividing the training corpus into multiple frames of training speech; For each frame of training speech in the multiple frames of training speech, the frame of training speech is input into the original model, the input training speech is processed by the original model, and state information corresponding to the frame of training speech is obtained and cached; According to the state information corresponding to the i-th frame of training speech in the multi-frame training speech and the state information corresponding to the i+1-th frame to the i+M-th frame of training speech in the cache, complete state information with a length of M+1 frames is obtained; wherein i is a positive integer; Randomly masking frames 0 to M in the state information of the last M frames of the complete state information to obtain masked state information; Outputting a character probability vector corresponding to the i-th frame of training speech according to the covering state information through the original model; Update the parameters of the original model according to the character probability vector corresponding to each frame of training speech and the label information corresponding to the training corpus; The trained original model is used as the target model.
8. The method according to claim 6, characterized in that The target model is trained through the following training process: Obtaining the original model and the training corpus containing the complete wake-up word; The pronunciation of the remaining characters in the training corpus is truncated at random lengths, and the truncated training corpus is divided into a plurality of frames of training speech; For each frame of training speech in the multiple frames of training speech, the frame of training speech is input into the original model, the input training speech is processed by the original model, and state information corresponding to the frame of training speech is obtained and cached; According to the state information corresponding to the j-th frame of training speech in the multi-frame training speech and the state information corresponding to the j+1-th frame to the j+M-th frame of training speech in the cache, complete state information with a length of M+1 frames is obtained; wherein j is a positive integer; Outputting a character probability vector corresponding to the j-th frame of training speech according to the complete state information through the original model; Update the parameters of the original model according to the character probability vector corresponding to each frame of training speech and the label information corresponding to the training corpus; The trained original model is used as the target model.
9. A voice wake-up device, characterized in that: include: The first character determination module is used to determine the character probability vector corresponding to the continuous multiple frames of speech; wherein the character probability vector corresponding to each frame of speech is obtained according to the state information of the current frame of speech and the state information of the right M frames of speech; A pre-wake-up detection module, configured to use a preset pre-wake-up word as a decoding character path in a character probability vector corresponding to the continuous multi-frame speech, and determine a first path decoding score corresponding to the pre-wake-up word; wherein the pre-wake-up word is composed of the first N characters of the preset complete wake-up word, and M and N are both positive integers; A second character determination module is used for extracting the M frames of state information currently cached in sequence in response to the first path decoding score being greater than a first threshold, and obtaining a corresponding character probability vector according to each extracted state information; A complete wake-up detection module, used to determine a second path decoding score corresponding to the complete wake-up word according to the path decoding information of the pre-wake-up word and the M character probability vectors corresponding to the M frame state information; a wake-up confirmation module, configured to determine to wake up a device in response to the second path decoding score being greater than a second threshold; Among them, the complete wake-up detection module includes: a decoding score determination subunit, configured to, after each acquisition of a corresponding character probability vector, determine the decoding scores corresponding to the remaining characters in the complete wake-up word except the pre-wake-up word according to the path decoding information of the pre-wake-up word and the acquired corresponding character probability vector, and use the remaining characters in the complete wake-up word as decoding characters; The path score determination subunit is used to determine the second path decoding score corresponding to the complete wake-up word according to the first path decoding score corresponding to the pre-wake-up word and the decoding scores corresponding to the remaining characters.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processing device, the steps of the method described in any one of claims 1 to 8 are implemented.
11. An electronic device, characterized in that: include: a storage device having at least one computer program stored thereon; At least one processing device, configured to execute the at least one computer program in the storage device to implement the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Voice wakeup method and voice interaction device
CN106448663A
Wakeup word detection method, device and equipment based on artificial intelligence, and medium
CN110838289A