Voice signal recognition method, device, electronic device, storage medium and product
By using the decoding parameters and node information in the first decoding diagram and the second decoding diagram in the voice wake-up technique, the recognition results of the target voice signal are identified, and the problem of high false wake-up rate is solved, thereby achieving higher recognition accuracy and lower false wake-up rate.
Patent Information
- Application Number
- CN202111539867.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2041-12-15
AI Technical Summary
In the existing voice wake-up technology, the false wake-up rate is high, mainly due to improper setting of the preset threshold, which leads to words similar to those pronunciation of wake-up words being misidentified as wake-up words.
By receiving the target speech signal, a plurality of speech frames it contains are determined, and corresponding decoding parameters are determined in the first decoding diagram and the second decoding diagram. Based on these decoding parameters and node information in the path, the recognition result of the target voice signal is determined to determine whether to wake up the electronic device.
The accuracy of voice signal recognition is improved, the false wake-up rate is reduced, and the similarity relationship between the voice signal and the wake-up signal and the decoding path information corresponding to the voice signal is comprehensively considered.
Smart Images

Figure CN114299945B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice interaction technologies, and in particular, to a method, apparatus, electronic device, storage medium, and product for recognizing voice signals. Background Art
[0002] Currently, the voice wake-up function has been widely applied in voice interaction technologies. Before a user conducts voice interaction with an electronic device, the electronic device needs to be woken up by a wake-up word.
[0003] In related technologies, an electronic device recognizes a received voice signal to determine the similarity between the recognition result and the wake-up word; if the similarity is greater than a preset threshold, the electronic device is woken up, and if the similarity is not greater than the preset threshold, the electronic device is not woken up.
[0004] However, in the above-mentioned related technologies, in order to meet the needs of different users, the preset threshold is generally set to a moderate fixed value. However, there may be words with pronunciations similar to the wake-up word (for example, the word "rabbit" with a pronunciation similar to the wake-up word "Xiaodu"), and the similarity between the recognition result and the wake-up word will be greater than the preset threshold, resulting in the situation of mis-waking up the electronic device. Therefore, the mis-wake-up rate of the above method is relatively high. Summary of the Invention
[0005] Embodiments of this application provide a method, apparatus, electronic device, storage medium, and product for recognizing voice signals, which can reduce the mis-wake-up rate of voice wake-up. The technical solution is as follows:
[0006] On the one hand, a method for recognizing a voice signal is provided. The method includes:
[0007] Receiving a target voice signal and determining a plurality of voice frames included in the target voice signal;
[0008] Determining first decoding parameters of a first path of the plurality of voice frames in a first decoding graph and second decoding parameters of a second path of the plurality of voice frames in a second decoding graph. The first decoding graph includes decoding paths corresponding to a plurality of basic voice signals, and the second decoding graph includes decoding paths corresponding to a plurality of wake-up voice signals;
[0009] When the difference between the first decoding parameters and the second decoding parameters is not greater than a preset difference, determining a plurality of first nodes included in the first path and decoding parameters of each first node;
[0010] Based on the first decoding parameters, the plurality of first nodes, and the decoding parameters of each first node, determining a recognition result of the target voice signal, where the recognition result is used to indicate whether to wake up the electronic device.
[0011] In one possible implementation, determining the recognition result of the target speech signal based on the first decoding parameter, the multiple first nodes, and the decoding parameter of each first node includes:
[0012] Inputting the first decoding parameter, the multiple first nodes, and the decoding parameter of each first node into a speech recognition model to obtain the recognition result of the target speech signal, where the speech recognition model is used to obtain the recognition result based on the decoding parameter of the path, the multiple nodes included in the path, and the decoding parameter of each node.
[0013] In another possible implementation, the process of training the speech recognition model includes:
[0014] Obtaining a sample speech signal, where the sample speech signal includes a first speech signal and a second speech signal, the first speech signal is the speech signal corresponding to the wake-up word, and the second speech signal is the speech signal corresponding to the non-wake-up word;
[0015] Training an initial recognition model based on the first speech signal and the second speech signal until the accuracy rate of the initial recognition model reaches a preset threshold to obtain the speech recognition model.
[0016] In another possible implementation, training the initial recognition model based on the first speech signal and the second speech signal includes:
[0017] Determining a third path of the multiple speech frames included in the first speech signal in the first decoding graph, and a fourth path of the multiple speech frames included in the second speech signal in the first decoding graph;
[0018] Determining first path information and second path information, where the first path information includes the decoding parameter of the third path, the multiple third nodes included in the third path, and the decoding parameter of each third node, and the second path information includes the decoding parameter of the fourth path, the multiple fourth nodes included in the fourth path, and the decoding parameter of each fourth node;
[0019] Training the initial recognition model based on the first path information and the second path information.
[0020] In another possible implementation, obtaining the sample speech signal includes:
[0021] Receiving the speech signal corresponding to the wake-up word and the speech signal corresponding to the non-wake-up word;
[0022] Perform noise addition processing on the voice signal corresponding to the wake-up word to obtain a first voice signal, and perform noise addition processing on the voice signal corresponding to the non-wake-up word to obtain a second voice signal.
[0023] In another possible implementation, the determining of the first decoding parameter of the first path of the multiple voice frames in the first decoding graph includes:
[0024] Determine the decoding parameters of the multiple decoding paths of the multiple voice frames in the first decoding graph;
[0025] From the decoding parameters of the multiple decoding paths, determine the decoding parameter with the largest value as the first decoding parameter of the first path.
[0026] In another possible implementation, the determining of the decoding parameters of the multiple decoding paths of the multiple voice frames in the first decoding graph includes:
[0027] For each decoding path in the first decoding graph, determine the basic voice signal corresponding to the decoding path; determine the first language decoding parameter and the first acoustic decoding parameter of the multiple voice frames under the decoding path, where the first language decoding parameter is used to represent the matching probability between the multiple voice frames and the word sequence corresponding to the basic voice signal, and the first acoustic decoding parameter is used to represent the matching probability between the multiple voice frames and the first phoneme sequence, and the first phoneme sequence is obtained by decomposing the word sequence;
[0028] Determine the product of the first language decoding parameter and the first acoustic decoding parameter to obtain the decoding parameter of the multiple voice frames under the decoding path.
[0029] In another possible implementation, the determining of the second decoding parameter of the second path of the multiple voice frames in the second decoding graph includes:
[0030] Determine the decoding parameters of the multiple decoding paths of the multiple voice frames in the second decoding graph;
[0031] From the decoding parameters of the multiple decoding paths, determine the decoding parameter with the largest value as the second decoding parameter of the second path.
[0032] In another possible implementation, the determining of the decoding parameters of the multiple decoding paths of the multiple voice frames in the second decoding graph includes:
[0033] For each decoding path in the second decoding graph, determine the wake-up voice signal corresponding to the decoding path; determine the second language decoding parameters and the second acoustic decoding parameters of the multiple speech frames under the decoding path, where the second language decoding parameters are used to represent the matching probability between the multiple speech frames and the wake-up word sequence corresponding to the wake-up voice signal, and the second acoustic decoding parameters are used to represent the matching probability between the multiple speech frames and the second phoneme sequence, and the second phoneme sequence is obtained by decomposing the wake-up word sequence;
[0034] Determine the product of the second language decoding parameters and the second acoustic decoding parameters to obtain the decoding parameters of the multiple speech frames under the decoding path.
[0035] In another possible implementation, each speech frame includes a voice signal of a first preset duration;
[0036] The determining the multiple speech frames included in the target voice signal includes:
[0037] Divide the target voice signal according to a preset period to obtain the multiple speech frames included in the target voice signal.
[0038] In another possible implementation, the determining the multiple first nodes included in the first path and the decoding parameters of each first node includes:
[0039] Determine the jump order of the multiple first nodes included in the first path;
[0040] According to the jump order, determine the speech frame corresponding to each first node;
[0041] Determine the probability value that the phoneme corresponding to each first node is consistent with the phoneme corresponding to the speech frame, and use the probability value as the decoding parameter of each first node.
[0042] On the other hand, a voice signal recognition device is provided, and the device includes:
[0043] A receiving module, configured to receive a target voice signal and determine the multiple speech frames included in the target voice signal;
[0044] A first determination module, configured to determine the first decoding parameters of the first path in the first decoding graph of the multiple speech frames and determine the second decoding parameters of the second path in the second decoding graph of the multiple speech frames, where the first decoding graph includes decoding paths corresponding to multiple basic voice signals, and the second decoding graph includes decoding paths corresponding to multiple wake-up voice signals;
[0045] A second determination module, configured to determine a plurality of first nodes included in the first path and the decoding parameter of each first node when a difference between the first decoding parameter and the second decoding parameter is not greater than a preset difference.
[0046] A third determination module, configured to determine an identification result of the target voice signal based on the first decoding parameter, the plurality of first nodes, and the decoding parameter of each first node, where the identification result is used to indicate whether to wake up the electronic device.
[0047] In a possible implementation manner, the third determination module is configured to input the first decoding parameter, the plurality of first nodes, and the decoding parameter of each first node into a voice recognition model to obtain an identification result of the target voice signal, where the voice recognition model is used to obtain an identification result based on the decoding parameter of a path, a plurality of nodes included in the path, and the decoding parameter of each node.
[0048] In another possible implementation manner, the apparatus further includes a training module, and the training module includes:
[0049] An acquisition unit, configured to acquire a sample voice signal, where the sample voice signal includes a first voice signal and a second voice signal, the first voice signal is a voice signal corresponding to a wake-up word, and the second voice signal is a voice signal corresponding to a non-wake-up word;
[0050] A training unit, configured to train an initial recognition model based on the first voice signal and the second voice signal until an accuracy rate of the initial recognition model reaches a preset threshold to obtain the voice recognition model.
[0051] In another possible implementation manner, the training unit is configured to determine a third path of a plurality of voice frames included in the first voice signal in the first decoding graph, and a fourth path of a plurality of voice frames included in the second voice signal in the first decoding graph; determine first path information and second path information, where the first path information includes the decoding parameter of the third path, a plurality of third nodes included in the third path, and the decoding parameter of each third node, and the second path information includes the decoding parameter of the fourth path, a plurality of fourth nodes included in the fourth path, and the decoding parameter of each fourth node; and train the initial recognition model based on the first path information and the second path information.
[0052] In another possible implementation manner, the acquisition unit is configured to receive a voice signal corresponding to a wake-up word and a voice signal corresponding to a non-wake-up word; perform noise addition processing on the voice signal corresponding to the wake-up word to obtain a first voice signal, and perform noise addition processing on the voice signal corresponding to the non-wake-up word to obtain a second voice signal.
[0053] In another possible implementation, the first determination module is configured to determine the decoding parameters of multiple decoding paths of the multiple speech frames in the first decoding graph; and determine the decoding parameter with the largest value among the decoding parameters of the multiple decoding paths as the first decoding parameter of the first path.
[0054] In another possible implementation, the first determination module is configured to, for each decoding path in the first decoding graph, determine the base speech signal corresponding to the decoding path; determine the first language decoding parameter and the first acoustic decoding parameter of the multiple speech frames under the decoding path, where the first language decoding parameter is used to represent the matching probability between the multiple speech frames and the word sequence corresponding to the base speech signal, and the first acoustic decoding parameter is used to represent the matching probability between the multiple speech frames and the first phoneme sequence, and the first phoneme sequence is obtained by decomposing the word sequence; and determine the product of the first language decoding parameter and the first acoustic decoding parameter to obtain the decoding parameter of the multiple speech frames under the decoding path.
[0055] In another possible implementation, the first determination module is further configured to determine the decoding parameters of multiple decoding paths of the multiple speech frames in the second decoding graph; and determine the decoding parameter with the largest value among the decoding parameters of the multiple decoding paths as the second decoding parameter of the second path.
[0056] In another possible implementation, the first determination module is further configured to, for each decoding path in the second decoding graph, determine the wake-up speech signal corresponding to the decoding path; determine the second language decoding parameter and the second acoustic decoding parameter of the multiple speech frames under the decoding path, where the second language decoding parameter is used to represent the matching probability between the multiple speech frames and the wake-up word sequence corresponding to the wake-up speech signal, and the second acoustic decoding parameter is used to represent the matching probability between the multiple speech frames and the second phoneme sequence, and the second phoneme sequence is obtained by decomposing the wake-up word sequence; and determine the product of the second language decoding parameter and the second acoustic decoding parameter to obtain the decoding parameter of the multiple speech frames under the decoding path.
[0057] In another possible implementation, each speech frame includes a speech signal with a first preset duration;
[0058] The receiving module is configured to divide the target speech signal according to a preset period to obtain multiple speech frames included in the target speech signal.
[0059] In another possible implementation, the second determination module is configured to determine the jump order of a plurality of first nodes included in the first path; determine the speech frames corresponding to each first node according to the jump order; determine the probability value that the phoneme corresponding to each first node is consistent with the phoneme corresponding to the speech frame, and use the probability value as the decoding parameter of each first node.
[0060] On the other hand, an electronic device is provided. The electronic device includes one or more processors and one or more memories. At least one program code is stored in the one or more memories. The at least one program code is loaded and executed by the one or more processors to implement the speech signal recognition method according to any of the above implementations.
[0061] On the other hand, a computer-readable storage medium is provided. At least one program code is stored in the computer-readable storage medium. The at least one program code is loaded and executed by a processor to implement the speech signal recognition method according to any of the above implementations.
[0062] On the other hand, a computer program product is provided. The computer program product includes at least one program code. The at least one program code is loaded and executed by a processor to implement the speech signal recognition method according to any of the above implementations.
[0063] The beneficial effects of the technical solutions provided in the embodiments of the present application at least include:
[0064] The embodiments of the present application provide a method for recognizing a speech signal. Since the difference between the first decoding parameter and the second decoding parameter is considered to take into account the similarity relationship between the speech signal and the wake-up signal, and multiple parameters such as the decoding parameter corresponding to the first path, a plurality of first nodes on the first path, and the decoding parameter of each first node are considered to take into account the decoding path information corresponding to the speech signal. In this way, both the similarity relationship between the speech signal and the wake-up signal and the decoding path information corresponding to the speech signal are considered, so the accuracy of the recognition result is improved and the false wake-up rate is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0066] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0067] Figure 2 It is a flowchart of a method for recognizing a voice signal provided by an embodiment of the present application;
[0068] Figure 3 It is a flowchart of a method for recognizing a voice signal provided by an embodiment of the present application;
[0069] Figure 4 It is a block diagram of a device for recognizing a voice signal provided by an embodiment of the present application;
[0070] Figure 5 It is a block diagram of a device for recognizing a voice signal provided by an embodiment of the present application;
[0071] Figure 6 It is a block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0072] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0073] Terms such as "first", "second", "third", and "fourth" in the specification, claims, and accompanying drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes unlisted steps or units, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0074] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present application. Refer to Figure 1 , this implementation environment includes an electronic device 101 and a server 102. A client for the service provided by the server 102 is installed on the electronic device 101. The user corresponding to the electronic device 101 can implement functions such as data transmission and voice interaction between the client and the server 102 through the client. The client at least has the function of recognizing a voice signal, that is, recognizing whether the voice signal wakes up the electronic device 101. The client can also have functions such as voice control. Among them, the client can be a voice assistant or a voice control application, etc.
[0075] In a possible implementation, after the electronic device 101 recognizes a voice signal and determines that the voice signal is used to wake up the electronic device 101, the electronic device 101 is woken up, and then the voice signal is collected again. The control instruction corresponding to the voice signal collected again is recognized, and the operation corresponding to the control instruction is executed. Among them, the electronic device 101 recognizes the control instruction corresponding to the voice signal collected again, or sends the voice signal collected again to the server 102. The server 102 recognizes the control instruction corresponding to the voice signal collected again and returns it to the electronic device 101.
[0076] The electronic device 101 can be a computer, a mobile phone, a speaker, an air conditioner, a TV, or other electronic devices. The server 102 can be a single server, or a server cluster composed of several servers, or a cloud computing service center.
[0077] Figure 2 is a flowchart of a method for recognizing a voice signal provided by an embodiment of the present application. Refer to Figure 2 , the method includes:
[0078] 201. Receive a target voice signal and determine a plurality of voice frames included in the target voice signal.
[0079] 202. Determine first decoding parameters of a first path in a first decoding graph for the plurality of voice frames, and determine second decoding parameters of a second path in a second decoding graph for the plurality of voice frames. The first decoding graph includes decoding paths corresponding to a plurality of basic voice signals, and the second decoding graph includes decoding paths corresponding to a plurality of wake-up voice signals.
[0080] 203. When the difference between the first decoding parameters and the second decoding parameters is not greater than a preset difference, determine a plurality of first nodes included in the first path and decoding parameters of each first node.
[0081] 204. Based on the first decoding parameters, the plurality of first nodes, and the decoding parameters of each first node, determine an identification result of the target voice signal, where the identification result is used to indicate whether to wake up the electronic device.
[0082] In a possible implementation, based on the first decoding parameters, the plurality of first nodes, and the decoding parameters of each first node, determining the identification result of the target voice signal includes:
[0083] Input the first decoding parameters, the plurality of first nodes, and the decoding parameters of each first node into a voice recognition model to obtain an identification result of the target voice signal. The voice recognition model is used to obtain an identification result based on the decoding parameters of the path, the plurality of nodes included in the path, and the decoding parameters of each node.
[0084] In another possible implementation, the process of training a speech recognition model includes:
[0085] Obtain sample speech signals, where the sample speech signals include a first speech signal and a second speech signal. The first speech signal is the speech signal corresponding to the wake-up word, and the second speech signal is the speech signal corresponding to the non-wake-up word.
[0086] Based on the first speech signal and the second speech signal, train an initial recognition model until the accuracy of the initial recognition model reaches a preset threshold to obtain a speech recognition model.
[0087] In another possible implementation, training the initial recognition model based on the first speech signal and the second speech signal includes:
[0088] Determine the third path of multiple speech frames included in the first speech signal in the first decoding graph, and the fourth path of multiple speech frames included in the second speech signal in the first decoding graph.
[0089] Determine first path information and second path information. The first path information includes the decoding parameters of the third path, multiple third nodes included in the third path, and the decoding parameters of each third node. The second path information includes the decoding parameters of the fourth path, multiple fourth nodes included in the fourth path, and the decoding parameters of each fourth node.
[0090] Based on the first path information and the second path information, train the initial recognition model.
[0091] In another possible implementation, obtaining the sample speech signals includes:
[0092] Receive the speech signal corresponding to the wake-up word and the speech signal corresponding to the non-wake-up word.
[0093] Perform noise addition processing on the speech signal corresponding to the wake-up word to obtain the first speech signal, and perform noise addition processing on the speech signal corresponding to the non-wake-up word to obtain the second speech signal.
[0094] In another possible implementation, determining the first decoding parameter of the first path of multiple speech frames in the first decoding graph includes:
[0095] Determine the decoding parameters of multiple decoding paths of multiple speech frames in the first decoding graph.
[0096] From the decoding parameters of multiple decoding paths, determine the decoding parameter with the largest value as the first decoding parameter of the first path.
[0097] In another possible implementation, determining the decoding parameters of multiple decoding paths of multiple speech frames in the first decoding graph includes:
[0098] For each decoding path in the first decoding graph, determine the corresponding base speech signal; determine the first language decoding parameters and the first acoustic decoding parameters of multiple speech frames under the decoding path, where the first language decoding parameters are used to represent the matching probability between the multiple speech frames and the word sequence corresponding to the base speech signal, and the first acoustic decoding parameters are used to represent the matching probability between the multiple speech frames and the first phoneme sequence, and the first phoneme sequence is obtained by decomposing the word sequence;
[0099] Determine the product of the first language decoding parameters and the first acoustic decoding parameters to obtain the decoding parameters of the multiple speech frames under the decoding path.
[0100] In another possible implementation, determining the second decoding parameters of the second path of the multiple speech frames in the second decoding graph includes:
[0101] Determine the decoding parameters of multiple decoding paths of the multiple speech frames in the second decoding graph;
[0102] From the decoding parameters of the multiple decoding paths, determine the decoding parameter with the largest value as the second decoding parameter of the second path.
[0103] In another possible implementation, one wake-up speech signal corresponds to one wake-up word sequence;
[0104] Determine the decoding parameters of multiple decoding paths of the multiple speech frames in the second decoding graph, including:
[0105] For each decoding path in the second decoding graph, determine the corresponding wake-up speech signal;
[0106] Determine the second language decoding parameters and the second acoustic decoding parameters of the multiple speech frames under the decoding path, where the second language decoding parameters are used to represent the matching probability between the multiple speech frames and the wake-up word sequence corresponding to the wake-up speech signal, and the second acoustic decoding parameters are used to represent the matching probability between the multiple speech frames and the second phoneme sequence, and the second phoneme sequence is obtained by decomposing the wake-up word sequence;
[0107] Determine the product of the second language decoding parameters and the second acoustic decoding parameters to obtain the decoding parameters of the multiple speech frames under the decoding path.
[0108] In another possible implementation, each speech frame includes a speech signal of a first preset duration;
[0109] Determining the multiple speech frames included in the target speech signal includes:
[0110] Divide the target speech signal according to a preset period to obtain multiple speech frames included in the target speech signal.
[0111] In another possible implementation, determining the multiple first nodes included in the first path and the decoding parameters of each first node includes:
[0112] Determining the jump order of the multiple first nodes included in the first path;
[0113] According to the jump order, determining the voice frames corresponding to each first node;
[0114] Determining the probability value that the phoneme corresponding to each first node is consistent with the phoneme corresponding to the voice frame, and using the probability value as the decoding parameter of each first node.
[0115] An embodiment of the present application provides a method for identifying a voice signal. Since the difference between the first decoding parameter and the second decoding parameter is considered, the similarity relationship between the voice signal and the wake-up signal is considered. By multiple parameters such as the decoding parameter corresponding to the first path, the multiple first nodes on the first path, and the decoding parameter of each first node, the decoding path information corresponding to the voice signal is considered. In this way, both the similarity relationship between the voice signal and the wake-up signal and the decoding path information corresponding to the voice signal are considered, so the accuracy of the recognition result is improved and the false wake-up rate is reduced.
[0116] Figure 3 It is a flowchart of a method for identifying a voice signal provided by an embodiment of the present application, which is executed by an electronic device. Refer to Figure 3 , and the method includes:
[0117] 301. The electronic device receives a target voice signal and determines the multiple voice frames included in the target voice signal.
[0118] The electronic device includes a sleep state and a wake-up state. When the electronic device is in the sleep state, the electronic device is awakened by a voice signal, and the electronic device switches from the sleep state to the wake-up state. In one possible implementation, the target voice signal is any voice signal received when the electronic device is in the sleep state. Optionally, the voice signal is a voice signal corresponding to a wake-up word issued by the user.
[0119] In one possible implementation, each voice frame includes a voice signal of a first preset duration. Correspondingly, the step for the electronic device to determine the multiple voice frames included in the target voice signal is: the electronic device divides the target voice signal according to a preset period to obtain the multiple voice frames included in the target voice signal. Optionally, the preset period is the first preset duration, and the electronic device divides once every first preset duration. In the embodiment of the present application, the value of the first preset duration is not specifically limited and can be set and modified as needed. Optionally, the first preset duration is any value between 0.01 s and 0.1 s, for example: the first preset duration is 0.01 s, 0.05 s, 0.1 s, etc.
[0120] In a possible implementation, the target voice signal is a voice signal whose signal duration is greater than a second preset duration. Correspondingly, the step of the electronic device receiving the target voice signal is as follows: The electronic device receives a voice signal, determines the signal duration of the voice signal, and if the signal duration is greater than the second preset duration, determines that the voice signal is the target voice signal. In the embodiments of the present application, the value of the second preset duration is not specifically limited and can be set and modified as needed. Optionally, the second preset duration is any value between 0.5 s and 5 s. For example, the second preset duration is 0.5 s, 1 s, 1.5 s, etc.
[0121] In the embodiments of the present application, since the voice signal is determined to be the target voice signal only when the signal duration of the voice signal meets the preset duration, invalid voice signals with too short signal durations can be filtered out, the effectiveness of the target voice signal is improved, and thus the accuracy of the recognition method is improved.
[0122] 302. The electronic device determines the first decoding parameter of the first path of multiple voice frames in the first decoding graph, and the first decoding graph includes decoding paths corresponding to multiple basic voice signals.
[0123] In a possible implementation, the first decoding graph is a basic decoding graph in a WFST (Weighted Finite State Transducers) decoding graph. The basic decoding graph includes decoding paths corresponding to multiple basic voice signals, where the multiple basic voice signals include voice signals corresponding to wake-up words and voice signals corresponding to non-wake-up words. The decoding parameter of a path represents the path score of the path in the first decoding graph. When decoding a voice signal through the first decoding graph, the optimal path with the highest path score in the basic decoding graph is determined as the decoding path corresponding to the voice signal.
[0124] In a possible implementation, when decoding the target voice signal through the first decoding graph, the value of the first decoding parameter of the first path in the first decoding graph is the largest; where the first decoding parameter is the path score of the first path, that is to say, the first path is the optimal path with the highest score of the target voice signal in the first decoding graph. Correspondingly, the step of the electronic device determining the first decoding parameter of the first path of multiple voice frames in the first decoding graph is as follows: The electronic device determines the decoding parameters of multiple decoding paths of multiple voice frames in the first decoding graph; and determines the decoding parameter with the largest value among the decoding parameters of the multiple decoding paths as the first decoding parameter of the first path.
[0125] In a possible implementation, the electronic device determines the decoding parameters of multiple speech frames under each decoding path of the first decoding graph through acoustic decoding parameters and language decoding parameters. Correspondingly, the steps for the electronic device to determine the decoding parameters of multiple speech frames in multiple decoding paths of the first decoding graph are as follows: For each decoding path in the first decoding graph, the electronic device determines the basic speech signal corresponding to the decoding path; determines the first language decoding parameter and the first acoustic decoding parameter of multiple speech frames under the decoding path, where the first language decoding parameter is used to represent the matching probability between multiple speech frames and the word sequence corresponding to the basic speech signal, and the first acoustic decoding parameter is used to represent the matching probability between multiple speech frames and the first phoneme sequence, and the first phoneme sequence is obtained by decomposing the word sequence corresponding to the basic speech signal; determines the product of the first language decoding parameter and the first acoustic decoding parameter to obtain the decoding parameter of multiple speech frames under the decoding path.
[0126] Optionally, in the first decoding graph, the decoding parameter of multiple speech frames under the decoding path is used to represent the path score for decoding multiple speech frames through the decoding path. In a possible implementation, the electronic device determines the matching probability between multiple speech frames and the word sequence corresponding to the basic speech signal through the linguistic model in the chain model, and determines this matching probability as the first language decoding parameter of multiple speech frames under the decoding path. The electronic device determines the matching probability between multiple speech frames and the first phoneme sequence through the acoustic model in the chain model, and determines this matching probability as the first acoustic decoding parameter of multiple speech frames under the decoding path.
[0127] In the embodiments of the present application, since the electronic device determines the decoding parameters of the decoding path through the linguistic model and the acoustic model in the chain model, thus comprehensively referring to the language decoding parameter and the acoustic decoding parameter, the accuracy of the determined decoding parameter is improved.
[0128] 303. The electronic device determines the second decoding parameter of the second path in the second decoding graph, and the second decoding graph includes decoding paths corresponding to multiple wake-up speech signals.
[0129] In a possible implementation, the WFST decoding graph can decode the wake-up speech signal to obtain the decoding path corresponding to the wake-up speech signal. The second decoding graph includes decoding paths corresponding to multiple wake-up speech signals. The wake-up speech signal can be the speech signal corresponding to the wake-up word stored in the electronic device. The wake-up word stored in the electronic device can be any wake-up word. For example: If the wake-up word stored in the electronic device is "Hello", then the wake-up speech signal is the speech signal corresponding to the wake-up word "Hello".
[0130] In a possible implementation, when decoding the target voice signal in the second decoding graph, the value of the second decoding parameter of the second path in the second decoding graph is the largest; wherein, the second decoding parameter is the path score of the second path, that is to say, the second path is the optimal path with the highest score for the target voice signal in the second decoding graph. Correspondingly, the steps for the electronic device to determine the second decoding parameter of the second path of multiple speech frames in the second decoding graph are as follows: The electronic device determines the decoding parameters of multiple decoding paths of multiple speech frames in the second decoding graph; from the decoding parameters of multiple decoding paths, it determines the decoding parameter with the largest value as the second decoding parameter of the second path.
[0131] In a possible implementation, the electronic device determines the decoding parameters of multiple speech frames under each decoding path in the second decoding graph through acoustic decoding parameters and language decoding parameters. Correspondingly, the steps for the electronic device to determine the decoding parameters of multiple decoding paths of multiple speech frames in the second decoding graph are as follows: For each decoding path in the second decoding graph, the electronic device determines the wake-up voice signal corresponding to the decoding path; it determines the second language decoding parameter and the second acoustic decoding parameter of multiple speech frames under the decoding path. The second language decoding parameter is used to represent the matching probability between multiple speech frames and the wake-up word sequence corresponding to the wake-up voice signal, and the second acoustic decoding parameter is used to represent the matching probability between multiple speech frames and the second phoneme sequence, and the second phoneme sequence is obtained by decomposing the wake-up word sequence; it determines the product of the second language decoding parameter and the second acoustic decoding parameter to obtain the decoding parameter of multiple speech frames under the decoding path.
[0132] Optionally, in the second decoding graph, the decoding parameter of multiple speech frames under the decoding path is used to represent the path score for decoding multiple speech frames through the decoding path. In a possible implementation, the electronic device determines the matching probability between multiple speech frames and the word sequence corresponding to the wake-up voice signal through the linguistic model in the chain model, and determines this matching probability as the second language decoding parameter of multiple speech frames under the decoding path. The electronic device determines the matching probability between multiple speech frames and the second phoneme sequence through the acoustic model in the chain model, and determines this matching probability as the second acoustic decoding parameter of multiple speech frames under the decoding path.
[0133] In the embodiments of the present application, since the electronic device determines the decoding parameters of the decoding path through the linguistic model and the acoustic model in the chain model, thus comprehensively referring to the language decoding parameter and the acoustic decoding parameter, the accuracy of the determined decoding parameter is improved.
[0134] It should be noted that there is no necessary order between step 302 and step 303. The electronic device may first execute step 302 and then execute step 303; it may also first execute step 303 and then execute step 302, or may execute step 302 and step 303 simultaneously.
[0135] 304. When the difference between the first decoding parameter and the second decoding parameter of the electronic device is not greater than a preset difference, the electronic device determines a plurality of first nodes included in the first path and the decoding parameter of each first node.
[0136] In a possible implementation manner, if the target voice signal is a voice signal corresponding to a wake-up word, the first decoding graph and the second decoding graph are used to decode the target voice signal, and the obtained first decoding parameter and the second decoding parameter are relatively close; if the target voice signal is a voice signal corresponding to a non-wake-up word, the first decoding graph and the second decoding graph are used to decode the target voice signal, and the obtained first decoding parameter and the second decoding parameter are relatively different. Before the electronic device determines a plurality of first nodes included in the first path and the decoding parameter of each first node, it is necessary to first determine whether the difference between the first decoding parameter and the second decoding parameter is greater than a preset difference. When the difference between the first decoding parameter and the second decoding parameter of the electronic device is not greater than the preset difference, the step of determining a plurality of first nodes included in the first path and the decoding parameter of each first node is executed; when the difference between the first decoding parameter and the second decoding parameter is greater than the preset difference, it is determined that the target voice signal is invalid and the electronic device is not awakened. In this step, the value of the preset difference is not specifically limited and can be set and modified as needed. Optionally, the preset difference is any value between 0.001 and 0.1. For example, the preset difference is 0.005, 0.05, 0.1, etc.
[0137] In the embodiment of the present application, since the electronic device makes a primary judgment on the target voice signal through the difference between the first decoding parameter and the second decoding parameter, that is, when the target voice signal meets the conditions of the wake-up voice, the recognition result of the target voice signal is further determined according to the path information feature, thus effectively avoiding the interference of invalid voice signals, so the recognition efficiency of the recognition method is improved.
[0138] In a possible implementation manner, the step for the electronic device to determine a plurality of first nodes included in the first path and the decoding parameter of each first node is as follows: the electronic device determines the jump order of a plurality of first nodes included in the first path; according to the jump order, determines the voice frame corresponding to each first node; determines the probability value that the phoneme corresponding to each first node is consistent with the phoneme corresponding to the voice frame, and uses the probability value as the decoding parameter of each first node. Optionally, the phoneme is the smallest voice unit. For example, vowels, consonants, etc. in English; or initials, finals, etc. in Chinese.
[0139] In a possible implementation, the steps for the electronic device to determine the jump order of multiple first nodes included in the first path are as follows: The electronic device decodes multiple speech frames in sequence according to the time order of the multiple speech frames to obtain multiple first nodes; and determines the jump order of the multiple first nodes included in the first path according to the decoding order.
[0140] For example, the time order of the multiple speech frames is speech frame 1 → speech frame 2 → speech frame 3 → speech frame 4 → speech frame 5; the multiple speech frames are decoded in sequence to obtain multiple first nodes as node 1, node 2, node 3, node 4, and node 5; according to the decoding order, the jump order of the multiple first nodes is determined as: node 1 → node 2 → node 3 → node 4 → node 5, and it is determined that speech frame 1 corresponds to node 1, speech frame 2 corresponds to node 2, speech frame 3 corresponds to node 3, speech frame 4 corresponds to node 4, and speech frame 5 corresponds to node 5.
[0141] For example, the phonemes corresponding to node 1, node 2, node 3, node 4, and node 5 in sequence are "x, i, ao, y, i"; the probability value that the phoneme x corresponding to node 1 is consistent with the phoneme of speech frame 1 is determined to obtain the decoding parameter of node 1; the probability value that the phoneme i corresponding to node 2 is consistent with the phoneme of speech frame 2 is determined to obtain the decoding parameter of node 2; the probability value that the phoneme ao corresponding to node 3 is consistent with the phoneme of speech frame 3 is determined to obtain the decoding parameter of node 3; the probability value that the phoneme y corresponding to node 4 is consistent with the phoneme of speech frame 4 is determined to obtain the decoding parameter of node 4; the probability value that the phoneme i corresponding to node 5 is consistent with the phoneme of speech frame 5 is determined to obtain the decoding parameter of node 5.
[0142] 305. The electronic device determines the recognition result of the target speech signal based on the first decoding parameter, the multiple first nodes, and the decoding parameter of each first node, and the recognition result is used to indicate whether to wake up the electronic device.
[0143] In a possible implementation, the electronic device determines the recognition result of the target speech signal according to the speech recognition model. Correspondingly, this step is: The electronic device inputs the first decoding parameter, the multiple first nodes, and the decoding parameter of each first node into the speech recognition model to obtain the recognition result of the target speech signal, and the speech recognition model is used to obtain the recognition result based on the decoding parameter of the path, the multiple nodes included in the path, and the decoding parameter of each node. Optionally, the speech recognition model is a fully connected neural network model.
[0144] In a possible implementation, the electronic device inputs the first decoding parameter, multiple first nodes, and the feature vectors corresponding to the decoding parameters of each first node into a speech recognition model to obtain the recognition result of the target speech signal. Optionally, the first decoding parameter is a path score, and the decoding parameter of the first node is a node score. For example, the first decoding parameter is: 0.0279, and the multiple first nodes and the decoding parameters of each first node are: node 1 and the decoding parameter of node 1 is 0.35, node 2 and the decoding parameter of node 2 is 0.5, node 3 and the decoding parameter of node 3 is 0.25, node 4 and the decoding parameter of node 4 is 0.75, node 5 and the decoding parameter of node 5 is 0.85. The first decoding parameter, the multiple first nodes, and the feature vectors corresponding to the decoding parameters of each first node are: {0.0279, node 1, 0.35, node 2, 0.5, node 3, 0.25, node 4, 0.75, node 5, 0.85}.
[0145] In the embodiment of the present application, since the recognition result is determined by the speech recognition model, and the speech recognition model can determine the recognition result by combining the decoding parameter corresponding to the path, multiple nodes, and the decoding parameters of the nodes, when the speech signal is recognized, the decoding path information such as the decoding parameter corresponding to the path, multiple nodes, and the decoding parameters of the nodes can be considered, so the accuracy of the determined recognition result is improved.
[0146] It should be noted that before obtaining the recognition result through the speech recognition model, the electronic device can first obtain a sample speech signal and train the speech recognition model through the sample speech signal.
[0147] In a possible implementation, the process of the electronic device training the speech recognition model is as follows: The electronic device obtains a sample speech signal, where the sample speech signal includes a first speech signal and a second speech signal. The first speech signal is the speech signal corresponding to the wake-up word, and the second speech signal is the speech signal corresponding to the non-wake-up word; Based on the first speech signal and the second speech signal, the initial recognition model is trained until the accuracy rate of the initial recognition model reaches a preset threshold to obtain the speech recognition model. Optionally, the non-wake-up word is a false wake-up word. The false wake-up word is a wake-up word in the non-wake-up words that can wake up the electronic device through the speech wake-up method in the prior art. In the embodiment of the present application, the value of the preset threshold is not specifically limited and can be set and modified as needed. Optionally, the preset threshold is any value between 80% and 100%, for example: the preset threshold is 85%, 90%, 95%, etc.
[0148] In a possible implementation, the steps for the electronic device to obtain the sample voice signal are as follows: The electronic device receives the voice signal corresponding to the wake-up word and the voice signal corresponding to the non-wake-up word; performs noise addition processing on the voice signal corresponding to the wake-up word to obtain a first voice signal, and performs noise addition processing on the voice signal corresponding to the non-wake-up word to obtain a second voice signal. Optionally, the noise is background noise. Correspondingly, this step is: The electronic device superimposes the voice signal corresponding to the wake-up word and the background noise signal to obtain a first voice signal, and superimposes the voice signal corresponding to the non-wake-up word and the background noise signal to obtain a second voice signal.
[0149] In the embodiments of the present application, since the sample voice signal is obtained by performing noise addition processing on the voice signal, the voice recognition model obtained by training with the sample voice signal has a high anti-noise ability, thereby improving the accuracy of the recognition result determined based on the voice recognition model.
[0150] In a possible implementation, the initial recognition model is trained according to the decoding path information of the first voice signal and the second voice signal in the first decoding graph. Correspondingly, the steps for the electronic device to train the initial recognition model based on the first voice signal and the second voice signal are as follows: The electronic device determines the third path of the multiple voice frames included in the first voice signal in the first decoding graph, and the fourth path of the multiple voice frames included in the second voice signal in the first decoding graph; determines the first path information and the second path information, where the first path information includes the decoding parameters of the third path, the multiple third nodes included in the third path, and the decoding parameters of each third node, and the second path information includes the decoding parameters of the fourth path, the multiple fourth nodes included in the fourth path, and the decoding parameters of each fourth node; trains the initial recognition model based on the first path information and the second path information.
[0151] Optionally, the initial recognition model is a fully connected neural network model, and the input samples are the feature vectors corresponding to the first path information and the second path information. Correspondingly, the steps for the electronic device to train the initial recognition model based on the first path information and the second path information are as follows: The electronic device determines the feature vector corresponding to the first path information and the feature vector corresponding to the second path information, inputs the feature vector corresponding to the first path information and the feature vector corresponding to the second path information into the initial recognition model, and trains the initial recognition model.
[0152] Optionally, the decoding parameter of the path is the path score, and the decoding parameter of the node is the node score. For example, the first path information includes: the decoding parameter 0.025 of the third path, the third node A, the decoding parameter 0.25 of the third node A... the third node P, the decoding parameter 0.5 of the third node P; the feature vector corresponding to the first path information is {0.025, A, 0.25... P, 0.5}. Among them, the number of multiple third nodes is positively correlated with the number of multiple speech frames included in the first speech signal. For example, the second path information includes: the decoding parameter 0.015 of the fourth path, the fourth node a, the decoding parameter 0.35 of the fourth node a... the fourth node p, the decoding parameter 0.45 of the fourth node p; the feature vector corresponding to the second path information is {0.015, a, 0.35... p, 0.45}. Among them, the number of multiple fourth nodes is positively correlated with the number of multiple speech frames included in the second speech signal.
[0153] In a possible implementation, there are multiple sample speech signals, and the dimensions of the feature vectors corresponding to the multiple sample speech signals are the same. Correspondingly, the steps for the electronic device to determine the feature vectors corresponding to the first path information and the second path information are as follows: The electronic device determines the dimensions of the feature vectors corresponding to the first path information and the second path information. If the dimension is less than the preset dimension, the dimension of the feature vector is padded to the preset dimension. If the dimension is greater than the preset dimension, the dimension of the feature vector is truncated to the preset dimension. Optionally, the parameter corresponding to padding the dimension is 0. In the embodiments of the present application, the value of the preset dimension is not specifically limited and can be set and modified as needed. Optionally, the preset dimension is any value between 30 and 100 dimensions. For example: the preset dimension is 50 dimensions, 60 dimensions, 80 dimensions, etc.
[0154] In the embodiments of the present application, since the dimensions of the feature vectors corresponding to the multiple sample speech signals are the same, when training the initial recognition model with feature vectors of the same dimension, the influence of the dimension on the training result is avoided, so the efficiency of training the speech recognition model is improved.
[0155] 306. When the recognition result is used to indicate waking up the electronic device, wake up the electronic device. When the recognition result is used to indicate not waking up the electronic device, do not wake up the electronic device.
[0156] In a possible implementation, the electronic device wakes up the electronic device through a wake-up module. Correspondingly, this step is as follows: When the recognition result is used to indicate waking up the electronic device, the electronic device sends a wake-up instruction to the wake-up module. The wake-up module receives the wake-up instruction and wakes up the electronic device. When the recognition result is used to indicate not waking up the electronic device, the electronic device does not send a wake-up instruction to the wake-up module.
[0157] In a possible implementation, after waking up the electronic device, the electronic device is in a wake-up state. The electronic device collects a new voice signal, recognizes the new voice signal, and obtains a control instruction corresponding to the new voice signal; and controls the electronic device to perform relevant operations or feedback according to the control instruction, so as to realize the control of the voice signal.
[0158] The embodiment of the present application provides a method for recognizing a voice signal. Since the decoding parameters corresponding to the first path and the decoding parameters corresponding to the second path are considered, the similarity relationship between the voice signal and the wake-up signal is considered. By using multiple parameters such as the decoding parameters corresponding to the first path, multiple first nodes on the first path, and the decoding parameters of each first node, the decoding path information corresponding to the voice signal is considered. In this way, both the similarity relationship between the voice signal and the wake-up signal and the decoding path information corresponding to the voice signal are considered, so the accuracy of the recognition result is improved and the false wake-up rate is reduced.
[0159] Figure 4 is a block diagram of a voice signal recognition device provided by the embodiment of the present application. Refer to Figure 4 , the device includes:
[0160] A receiving module 401, configured to receive a target voice signal and determine a plurality of voice frames included in the target voice signal;
[0161] A first determination module 402, configured to determine the first decoding parameters of the first path of the plurality of voice frames in the first decoding graph, and determine the second decoding parameters of the second path of the plurality of voice frames in the second decoding graph. The first decoding graph includes decoding paths corresponding to a plurality of basic voice signals, and the second decoding graph includes decoding paths corresponding to a plurality of wake-up voice signals;
[0162] A second determination module 403, configured to determine a plurality of first nodes included in the first path and the decoding parameters of each first node when the difference between the first decoding parameters and the second decoding parameters is not greater than a preset difference;
[0163] A third determination module 404, configured to determine the recognition result of the target voice signal based on the first decoding parameters, the plurality of first nodes, and the decoding parameters of each first node, where the recognition result is used to indicate whether to wake up the electronic device.
[0164] In a possible implementation, the third determination module 404 is configured to input the first decoding parameters, the plurality of first nodes, and the decoding parameters of each first node into a voice recognition model to obtain the recognition result of the target voice signal. The voice recognition model is used to obtain the recognition result based on the decoding parameters of the path, the plurality of nodes included in the path, and the decoding parameters of each node.
[0165] In another possible implementation, refer to Figure 5, the device further includes a training module 405, and the training module 405 includes:
[0166] An acquisition unit 4051, configured to acquire sample voice signals, where the sample voice signals include a first voice signal and a second voice signal, the first voice signal is a voice signal corresponding to a wake-up word, and the second voice signal is a voice signal corresponding to a non-wake-up word;
[0167] A training unit 4052, configured to train an initial recognition model based on the first voice signal and the second voice signal until the accuracy rate of the initial recognition model reaches a preset threshold, so as to obtain a voice recognition model.
[0168] In another possible implementation manner, the training unit 4052 is configured to determine a third path of multiple voice frames included in the first voice signal in the first decoding graph, and a fourth path of multiple voice frames included in the second voice signal in the first decoding graph; determine first path information and second path information, where the first path information includes decoding parameters of the third path, multiple third nodes included in the third path, and decoding parameters of each third node, and the second path information includes decoding parameters of the fourth path, multiple fourth nodes included in the fourth path, and decoding parameters of each fourth node; train the initial recognition model based on the first path information and the second path information.
[0169] In another possible implementation manner, the acquisition unit 4051 is configured to receive a voice signal corresponding to a wake-up word and a voice signal corresponding to a non-wake-up word; perform noise addition processing on the voice signal corresponding to the wake-up word to obtain a first voice signal, and perform noise addition processing on the voice signal corresponding to the non-wake-up word to obtain a second voice signal.
[0170] In another possible implementation manner, the first determination module 402 is configured to determine decoding parameters of multiple decoding paths of multiple voice frames in the first decoding graph; determine, from the decoding parameters of the multiple decoding paths, the decoding parameter with the largest value as the first decoding parameter of the first path.
[0171] In another possible implementation manner, the first determination module 402 is configured to, for each decoding path in the first decoding graph, determine a basic voice signal corresponding to the decoding path; determine a first language decoding parameter and a first acoustic decoding parameter of multiple voice frames under the decoding path, where the first language decoding parameter is used to represent the matching probability between the multiple voice frames and a word sequence corresponding to the basic voice signal, and the first acoustic decoding parameter is used to represent the matching probability between the multiple voice frames and a first phoneme sequence, and the first phoneme sequence is obtained by decomposing the word sequence; determine the product of the first language decoding parameter and the first acoustic decoding parameter to obtain the decoding parameter of the multiple voice frames under the decoding path.
[0172] In another possible implementation, the first determination module 402 is further configured to determine decoding parameters of multiple decoding paths of multiple speech frames in a second decoding graph; and determine the decoding parameter with the largest value among the decoding parameters of the multiple decoding paths as the second decoding parameter of the second path.
[0173] In another possible implementation, for each decoding path in the second decoding graph, the first determination module 402 is further configured to determine a wake-up speech signal corresponding to the decoding path; determine a second language decoding parameter and a second acoustic decoding parameter of the multiple speech frames under the decoding path, where the second language decoding parameter is used to represent the matching probability between the multiple speech frames and a wake-up word sequence corresponding to the wake-up speech signal, and the second acoustic decoding parameter is used to represent the matching probability between the multiple speech frames and a second phoneme sequence, and the second phoneme sequence is obtained by decomposing the wake-up word sequence; and determine the product of the second language decoding parameter and the second acoustic decoding parameter to obtain the decoding parameter of the multiple speech frames under the decoding path.
[0174] In another possible implementation, each speech frame includes a speech signal with a first preset duration;
[0175] The receiving module 401 is configured to divide the target speech signal according to a preset period to obtain multiple speech frames included in the target speech signal.
[0176] In another possible implementation, the second determination module 403 is configured to determine the jump order of multiple first nodes included in the first path; determine the speech frame corresponding to each first node according to the jump order; and determine the probability value that the phoneme corresponding to each first node is consistent with the phoneme corresponding to the speech frame, and use the probability value as the decoding parameter of each first node.
[0177] The embodiment of the present application provides a speech signal recognition device. Since the decoding parameter corresponding to the first path and the decoding parameter corresponding to the second path are used to consider the similarity relationship between the speech signal and the wake-up signal, and multiple parameters such as the decoding parameter corresponding to the first path, multiple first nodes on the first path, and the decoding parameter of each first node are used to consider the decoding path information corresponding to the speech signal, the similarity relationship between the speech signal and the wake-up signal is considered, and the decoding path information corresponding to the speech signal is considered, so the accuracy of the recognition result is improved and the false wake-up rate is reduced.
[0178] Figure 6The block diagram of an electronic device 600 provided by an exemplary embodiment of the present invention is shown. The electronic device 600 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer or a desktop computer. The electronic device 600 may also be referred to by other names such as a user device, a portable electronic device, a laptop electronic device, a desktop electronic device, etc.
[0179] Generally, the electronic device 600 includes: a processor 601 and a memory 602.
[0180] The processor 601 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor 601 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 601 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 601 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0181] The memory 602 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 602 may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 602 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 601 to implement the voice signal recognition method provided in the method embodiments of the present application.
[0182] In some embodiments, the electronic device 600 may further optionally include: a peripheral device interface 603 and at least one peripheral device. The processor 601, the memory 602, and the peripheral device interface 603 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 603 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 604, a display screen 605, a camera 606, an audio circuit 607, a positioning component 608, and a power supply 609.
[0183] The peripheral device interface 603 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 601 and the memory 602. In some embodiments, the processor 601, the memory 602, and the peripheral device interface 603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 601, the memory 602, and the peripheral device interface 603 may be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0184] The radio frequency circuit 604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 604 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 604 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency circuit 604 may communicate with other electronic devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: a metropolitan area network, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 604 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0185] The display screen 605 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 605 is a touch display screen, the display screen 605 also has the ability to collect touch signals on or above the surface of the display screen 605. The touch signals can be input as control signals to the processor 601 for processing. At this time, the display screen 605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 605, which is disposed on the front panel of the electronic device 600; in other embodiments, there may be at least two display screens 605, which are respectively disposed on different surfaces of the electronic device 600 or are in a foldable design; in still other embodiments, the display screen 605 may be a flexible display screen, which is disposed on a curved surface or a folding surface of the electronic device 600. Even more, the display screen 605 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 605 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0186] The camera module 606 is used to capture images or videos. Optionally, the camera module 606 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the electronic device, and the rear camera is disposed on the back of the electronic device. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to implement functions such as background blurring by fusing the main camera and the depth-of-field camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fused shooting functions. In some embodiments, the camera module 606 may also include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. A two-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0187] The audio circuit 607 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 601 for processing, or input to the radio frequency circuit 604 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 600. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 601 or the radio frequency circuit 604 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 607 may further include a headphone jack.
[0188] The positioning component 608 is used to locate the current geographical location of the electronic device 600 to achieve navigation or LBS (Location Based Service). The positioning component 608 may be a positioning component based on the GPS (Global Positioning System) of the United States, the Beidou system of China, the GLONASS system of Russia, or the Galileo system of the European Union.
[0189] The power supply 609 is used to supply power to each component in the electronic device 600. The power supply 609 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 609 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0190] In some embodiments, the electronic device 600 further includes one or more sensors 610. The one or more sensors 610 include but are not limited to: an acceleration sensor 611, a gyroscope sensor 612, a pressure sensor 613, a fingerprint sensor 614, an optical sensor 615, and a proximity sensor 616.
[0191] The acceleration sensor 611 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the electronic device 600. For example, the acceleration sensor 611 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 601 can control the display screen 605 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 611. The acceleration sensor 611 can also be used for collecting game or user's motion data.
[0192] The gyroscope sensor 612 can detect the body orientation and rotation angle of the electronic device 600. The gyroscope sensor 612 can cooperate with the acceleration sensor 611 to collect the 3D actions of the user on the electronic device 600. Based on the data collected by the gyroscope sensor 612, the processor 601 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0193] The pressure sensor 613 can be disposed on the side frame of the electronic device 600 and / or the lower layer of the display screen 605. When the pressure sensor 613 is disposed on the side frame of the electronic device 600, it can detect the holding signal of the user on the electronic device 600, and the processor 601 can identify the left or right hand or perform a quick operation according to the holding signal collected by the pressure sensor 613. When the pressure sensor 613 is disposed on the lower layer of the display screen 605, the processor 601 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 605. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0194] The fingerprint sensor 614 is used to collect the fingerprint of the user. The processor 601 can identify the user's identity according to the fingerprint collected by the fingerprint sensor 614, or the fingerprint sensor 614 can identify the user's identity according to the collected fingerprint. When the identified user identity is a trusted identity, the processor 601 authorizes the user to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 614 can be disposed on the front, back, or side of the electronic device 600. When there are physical buttons or a manufacturer logo on the electronic device 600, the fingerprint sensor 614 can be integrated with the physical button or the manufacturer logo.
[0195] The optical sensor 615 is used to collect the ambient light intensity. In one embodiment, the processor 601 can control the display brightness of the display screen 605 according to the ambient light intensity collected by the optical sensor 615. Specifically, when the ambient light intensity is high, the display brightness of the display screen 605 is increased; when the ambient light intensity is low, the display brightness of the display screen 605 is decreased. In another embodiment, the processor 601 can also dynamically adjust the shooting parameters of the camera module 606 according to the ambient light intensity collected by the optical sensor 615.
[0196] The proximity sensor 616, also known as a distance sensor, is typically disposed on the front panel of the electronic device 600. The proximity sensor 616 is used to collect the distance between the user and the front of the electronic device 600. In one embodiment, when the proximity sensor 616 detects that the distance between the user and the front of the electronic device 600 is gradually decreasing, the processor 601 controls the display screen 605 to switch from the lit state to the off state; when the proximity sensor 616 detects that the distance between the user and the front of the electronic device 600 is gradually increasing, the processor 601 controls the display screen 605 to switch from the off state to the lit state.
[0197] Those skilled in the art can understand that Figure 6 the structure shown in does not constitute a limitation on the electronic device 600, and may include more or fewer components than shown, or combine certain components, or adopt a different component arrangement.
[0198] The embodiment of the present application also provides a computer-readable storage medium, in which at least one program code is stored, and the at least one program code is loaded and executed by a processor to implement the voice signal recognition method as described in any of the above implementation manners.
[0199] The embodiment of the present application also provides a computer program product, which includes at least one program code, and the at least one program code is loaded and executed by a processor to implement the voice signal recognition method as described in any of the above implementation manners.
[0200] In some embodiments, the computer program involved in the embodiment of the present application can be deployed to be executed on a computer device, or on multiple computer devices located at one place, or alternatively, on multiple computer devices distributed at multiple places and interconnected by a communication network. The multiple computer devices distributed at multiple places and interconnected by a communication network can form a blockchain system.
[0201] The above are only optional embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for recognizing a speech signal, characterized in that: The method comprises: receiving a target speech signal, and determining a plurality of speech frames included in the target speech signal; Determine a first decoding parameter of a first path of the plurality of speech frames in a first decoding graph, and determine a second decoding parameter of a second path of the plurality of speech frames in a second decoding graph, wherein the first decoding graph includes decoding paths corresponding to a plurality of basic speech signals, and the second decoding graph includes decoding paths corresponding to a plurality of wake-up speech signals; When the difference between the first decoding parameter and the second decoding parameter is not greater than a preset difference, determining a decoding parameter of each of the plurality of first nodes included in the first path; Determine a recognition result of the target voice signal based on the first decoding parameter, the plurality of first nodes and the decoding parameter of each first node, wherein the recognition result is used to indicate whether to wake up the electronic device; Wherein, the determining of the first decoding parameters of the first path of the multiple speech frames in the first decoding graph includes: for each decoding path in the first decoding graph, determining the basic speech signal corresponding to the decoding path; determining the first language decoding parameters and the first acoustic decoding parameters of the multiple speech frames under the decoding path, the first language decoding parameters are used to represent the matching probability between the multiple speech frames and the word sequence corresponding to the basic speech signal, and the first acoustic decoding parameters are used to represent the matching probability between the multiple speech frames and the first phoneme sequence, and the first phoneme sequence is obtained based on the decomposition of the word sequence; determining the product of the first language decoding parameter and the first acoustic decoding parameter to obtain the decoding parameters of the multiple speech frames under the decoding path; and determining the decoding parameter with the largest value from the decoding parameters of the multiple decoding paths as the first decoding parameter of the first path.
2. The method according to claim 1, characterized in that The determining the recognition result of the target speech signal based on the first decoding parameter, the plurality of first nodes and the decoding parameter of each first node comprises: The first decoding parameter, the multiple first nodes and the decoding parameter of each first node are input into a speech recognition model to obtain a recognition result of the target speech signal. The speech recognition model is used to obtain the recognition result based on the decoding parameter of the path, the multiple nodes included in the path and the decoding parameter of each node.
3. The method according to claim 2, characterized in that The process of training the speech recognition model includes: Acquire a sample voice signal, where the sample voice signal includes a first voice signal and a second voice signal, where the first voice signal is a voice signal corresponding to a wake-up word, and the second voice signal is a voice signal corresponding to a non-wake-up word; Based on the first speech signal and the second speech signal, an initial recognition model is trained until the accuracy of the initial recognition model reaches a preset threshold, thereby obtaining the speech recognition model.
4. The method according to claim 3, characterized in that The training of the initial recognition model based on the first speech signal and the second speech signal includes: Determine a third path of a plurality of speech frames included in the first speech signal in the first decoding graph, and a fourth path of a plurality of speech frames included in the second speech signal in the first decoding graph; Determine first path information and second path information, wherein the first path information includes a decoding parameter of a third path, a plurality of third nodes included in the third path, and a decoding parameter of each third node, and the second path information includes a decoding parameter of a fourth path, a plurality of fourth nodes included in the fourth path, and a decoding parameter of each fourth node; An initial recognition model is trained based on the first path information and the second path information.
5. The method according to claim 3, characterized in that: The step of obtaining a sample voice signal comprises: Receive a voice signal corresponding to a wake-up word and a voice signal corresponding to a non-wake-up word; The voice signal corresponding to the wake-up word is subjected to noise processing to obtain a first voice signal, and the voice signal corresponding to the non-wake-up word is subjected to noise processing to obtain a second voice signal.
6. The method according to claim 1, characterized in that Determining the second decoding parameters of the second path of the plurality of speech frames in the second decoding graph comprises: Determining decoding parameters of a plurality of decoding paths of the plurality of speech frames in the second decoding graph; From the decoding parameters of the multiple decoding paths, determine a decoding parameter with the largest value as the second decoding parameter of the second path.
7. The method according to claim 6, characterized in that The determining decoding parameters of the plurality of decoding paths of the plurality of speech frames in the second decoding graph comprises: For each decoding path in the second decoding graph, determining a wake-up speech signal corresponding to the decoding path; Determine a second language decoding parameter and a second acoustic decoding parameter of the plurality of speech frames in the decoding path, wherein the second language decoding parameter is used to represent a matching probability between the plurality of speech frames and a wake-up word sequence corresponding to the wake-up speech signal, and the second acoustic decoding parameter is used to represent a matching probability between the plurality of speech frames and a second phoneme sequence, wherein the second phoneme sequence is obtained by decomposing the wake-up word sequence; The product of the second language decoding parameter and the second acoustic decoding parameter is determined to obtain decoding parameters of the multiple speech frames in the decoding path.
8. The method according to claim 1, characterized in that Each speech frame includes a speech signal of a first preset duration; The determining of the plurality of speech frames included in the target speech signal comprises: The target speech signal is divided according to a preset period to obtain a plurality of speech frames included in the target speech signal.
9. The method according to claim 1, characterized in that: The determining of the plurality of first nodes included in the first path and a decoding parameter of each first node comprises: Determining a jump order of a plurality of first nodes included in the first path; Determine, according to the jump order, a speech frame corresponding to each first node; Determine a probability value that the phoneme corresponding to each first node is consistent with the phoneme corresponding to the speech frame, and use the probability value as a decoding parameter of each first node.
10. A speech signal recognition device, characterized in that: The device comprises: A receiving module, configured to receive a target speech signal and determine a plurality of speech frames included in the target speech signal; A first determining module, configured to determine a first decoding parameter of a first path of the plurality of speech frames in a first decoding graph, and to determine a second decoding parameter of a second path of the plurality of speech frames in a second decoding graph, wherein the first decoding graph includes decoding paths corresponding to a plurality of basic speech signals, and the second decoding graph includes decoding paths corresponding to a plurality of wake-up speech signals; A second determining module, configured to determine a decoding parameter of each of the plurality of first nodes included in the first path if a difference between the first decoding parameter and the second decoding parameter is not greater than a preset difference; A third determination module, configured to determine a recognition result of the target voice signal based on the first decoding parameter, the plurality of first nodes, and a decoding parameter of each first node, wherein the recognition result is used to indicate whether to wake up the electronic device; The first determination module is used to determine the first decoding parameters of the first path of the multiple speech frames in the first decoding graph, including: for each decoding path in the first decoding graph, determining the basic speech signal corresponding to the decoding path; determining the first language decoding parameters and the first acoustic decoding parameters of the multiple speech frames under the decoding path, the first language decoding parameters are used to represent the matching probability between the multiple speech frames and the word sequence corresponding to the basic speech signal, and the first acoustic decoding parameters are used to represent the matching probability between the multiple speech frames and the first phoneme sequence, and the first phoneme sequence is obtained based on the decomposition of the word sequence; determining the product of the first language decoding parameter and the first acoustic decoding parameter to obtain the decoding parameters of the multiple speech frames under the decoding path; and determining the decoding parameter with the largest value from the decoding parameters of the multiple decoding paths as the first decoding parameter of the first path.
11. An electronic device, characterized in that: The electronic device comprises: A processor and a memory, wherein at least one program code is stored in the memory, and the at least one program code is loaded and executed by the processor to implement the speech signal recognition method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that: The storage medium stores at least one program code, and the at least one program code is loaded and executed by the processor to implement the speech signal recognition method according to any one of claims 1 to 9.
13. A computer program product, characterized in that The computer program product includes at least one program code, and the at least one program code is loaded and executed by a processor to implement the speech signal recognition method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Voice wake-up method and device, electronic equipment and storage medium
CN110570857A
Awakening method and device of intelligent equipment, electronic equipment and medium
CN111554288A