Multi-wake-up word recognition method, device and equipment based on finite-state machine, and medium

Through the multi-wake word recognition method based on a finite state machine, the problems in the prior art that wake words cannot be flexibly replaced, high data requirements and high recognition delay are solved, and multi-wake word recognition with high accuracy and fast response are achieved.

CN120183390APending Publication Date: 2025-06-20E-SURFING DIGITAL LIFE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510338324.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing voice wake-up products are difficult to achieve flexible replacement of wake-up words, multi-wake word model data requirements and high recognition delay.

Method used

The multi-wake word recognition method based on a finite state machine is adopted, and the complete process from monitoring to recognition is realized by presetting the target wake-up words, constructing the target finite state converter diagram, feature extraction and acoustic model decoding, and quickly switch to the recognition state through the OneShot decoding algorithm.

Benefits of technology

It improves the recognition accuracy and response speed of multiple wake-up words, realizes flexible replacement of wake-up words, and reduces the demand for data annotation and computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183390A_ABST
    Figure CN120183390A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-wake-up-word recognition method and device based on a finite-state machine, equipment and a medium. The method comprises the steps that a target wake-up word is preset; constructing a target finite state converter diagram according to the target wake-up word; in response to the input voice, performing feature extraction on the input voice to obtain a feature frame; inputting the feature frame into an acoustic model to generate a first streaming result; and inputting the first streaming result into the target finite state converter diagram for decoding to obtain an identification state result. The method can improve the recognition accuracy and response speed of multiple wake-up words, and can be widely applied to the technical field of speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition, and in particular to a multi-wake word recognition method, device, equipment and medium based on a finite state machine. Background Art

[0002] Intelligent terminal devices with voice as the entry have been widely used in various fields, including mobile phones, computers, televisions and smart home devices, almost all of which support voice wake-up and voice control. As a pre-module of speech recognition, the voice wake-up module is independently deployed from the speech recognition system and is responsible for connecting the user's voice input and recognition service to achieve seamless interaction. Traditional wake-up words usually only listen to fixed phrases. To save resources, such modules tend to be as lightweight as possible and are usually deployed on hardware accelerators. However, most current voice wake-up products usually preset fixed wake-up words or support switching of a very small number (such as two or three) of wake-up words, and the wake-up words cannot be flexibly replaced. In addition, currently, to achieve accurate recognition of multiple wake-up words, a large number of wake-up word speech samples and noise data in specific scenarios are required for training and optimization. Especially in a multi-user and multi-scenario environment, the phoneme features of different wake-up words may be similar. If the data is insufficient, the model is prone to false wake-up or missed wake-up. And the multi-wake word model has a significantly increased computational load because it needs to listen to multiple specific words, and the latency is higher compared to the single wake-up word. Summary of the Invention

[0003] In view of this, the main purpose of the embodiments of the present invention is to provide a multi-wake word recognition method, device, equipment and medium based on a finite state machine, in order to solve at least one of the existing technical problems. The present invention can improve the recognition accuracy and response speed of multiple wake-up words.

[0004] To achieve the above object, on the one hand, an embodiment of the present invention provides a multi-wake word recognition method based on a finite state machine, the method includes the following steps:

[0005] Preset target wake-up words;

[0006] Construct a target finite state transducer graph according to the target wake-up words;

[0007] In response to the input voice, extract features from the input voice to obtain feature frames;

[0008] Input the feature frames into an acoustic model to generate a first streaming result;

[0009] Input the first streaming result into the target finite state transducer graph for decoding to obtain a recognition state result.

[0010] In some embodiments, a multi-wake word recognition method based on a finite state machine further includes the following steps:

[0011] Generate a trigger status code according to the recognition status result.

[0012] In some embodiments, constructing a target finite state transducer graph according to the target wake-up word includes the following steps:

[0013] Obtain a status node according to the first character of the target wake-up word;

[0014] Obtain the directed jump edges between the status nodes;

[0015] Construct an initial finite state transducer graph according to the status node and the directed jump edges;

[0016] Perform an optimization process on the initial finite state transducer graph to obtain the target finite state transducer graph.

[0017] In some embodiments, performing an optimization process on the initial finite state transducer graph to obtain the target finite state transducer graph includes the following steps:

[0018] Obtain a fault-tolerant wake-up word;

[0019] Obtain a fault-tolerant node according to the second character of the fault-tolerant wake-up word;

[0020] Add the fault-tolerant node and the unknown character edges to the initial finite state transducer graph to obtain an intermediate finite state transducer graph;

[0021] Set self-loops for the status nodes and the fault-tolerant nodes of the intermediate finite state transducer graph, and perform determinization on the intermediate finite state transducer graph to obtain the target finite state transducer graph.

[0022] In some embodiments, inputting the feature frame into an acoustic model to generate a first streaming result includes the following steps:

[0023] Input the feature frame into an acoustic model and output the character probability distribution at each time step;

[0024] Obtain path probabilities according to the character probability distribution through the CTC decoding algorithm;

[0025] Add up all the path probabilities through the CTC decoding algorithm to obtain a target sequence probability;

[0026] Obtain an initial character sequence according to the target sequence probability;

[0027] Delete the blank symbols and the consecutive repeated third characters in the initial character sequence through the CTC decoding algorithm to obtain a target character sequence;

[0028] Concatenate a plurality of the target character sequences to obtain the first streaming result.

[0029] In some embodiments, the step of inputting the first streaming result into the target finite state transducer graph for decoding to obtain an identification status result includes the following steps:

[0030] Read the fourth character of the current time step of the first streaming result;

[0031] Obtain the fifth character of the outgoing edge of the current state of the target finite state transducer graph;

[0032] Determine whether the fifth character matches the fourth character. If the fifth character does not match the fourth character, reset the current state to the initial state. If the fifth character matches the fourth character, update the current state to the next state pointed to by the outgoing edge, read the fourth character of the next time step, and return to the step of obtaining the fifth character of the outgoing edge of the current state of the target finite state transducer graph until all the fourth characters are read to obtain the identification status result.

[0033] In some embodiments, the step of generating a trigger status code according to the identification status result includes the following steps:

[0034] When the identification status result is an intermediate state, the trigger status code is an unawakened state;

[0035] When the identification status result is an end state, the trigger status code is a successfully awakened state.

[0036] To achieve the above object, another aspect of the embodiments of the present invention provides a multi-wake word recognition device based on a finite state machine, and the device includes:

[0037] A first module for presetting target wake words;

[0038] A second module for constructing a target finite state transducer graph according to the target wake words;

[0039] A third module for extracting features from the input voice in response to the input voice to obtain feature frames;

[0040] A fourth module for inputting the feature frames into an acoustic model to generate a first streaming result;

[0041] A fifth module for inputting the first streaming result into the target finite state transducer graph for decoding to obtain an identification status result.

[0042] To achieve the above object, another aspect of the embodiments of the present invention provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned multi wake-word recognition method based on a finite state machine.

[0043] To achieve the above object, another aspect of the embodiments of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned multi wake-word recognition method based on a finite state machine.

[0044] To achieve the above object, another aspect of the embodiments of the present invention provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of a computer device can read the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the computer device executes the above-mentioned multi wake-word recognition method based on a finite state machine.

[0045] The embodiments of the present invention at least include the following beneficial effects: The present invention provides a multi wake-word recognition method, device, equipment and medium based on a finite state machine. The solution presetts target wake words; constructs a target finite state transducer graph according to the target wake words to support the complete process from monitoring to recognition with an orderly state transition; in response to the input voice, extracts features of the input voice to obtain feature frames; inputs the feature frames into an acoustic model to generate a first streaming result; inputs the first streaming result into the target finite state transducer graph for decoding, while ensuring the streaming output of the speech recognition result, quickly switches to the recognition state to obtain the recognition state result, improving the recognition accuracy and response speed of multi wake words. Description of the Drawings

[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0047] Figure 1 is a flowchart of the method provided by the embodiments of the present invention;

[0048] Figure 2 is a schematic diagram of a traditional weighted finite state transducer (WFST) provided by the embodiments of the present invention;

[0049] Figure 3It is a finite state transducer diagram of the wake-up word and command word provided by an embodiment of the present invention;

[0050] Figure 4 It is a schematic diagram of fault tolerance path optimization of the finite state transducer diagram provided by an embodiment of the present invention;

[0051] Figure 5 It is a schematic diagram of the unknown character edge of the finite state transducer diagram provided by an embodiment of the present invention;

[0052] Figure 6 It is a schematic diagram of path self-loop of the finite state transducer diagram provided by an embodiment of the present invention;

[0053] Figure 7 It is a schematic diagram of graph determinization of the finite state transducer diagram provided by an embodiment of the present invention;

[0054] Figure 8 It is a schematic diagram of all nodes of the finite state transducer diagram provided by an embodiment of the present invention and their corresponding fallback states;

[0055] Figure 9 It is a schematic diagram of the network structure of the end-to-end acoustic model provided by an embodiment of the present invention;

[0056] Figure 10 It is a flow chart of the overall structure of the flexible custom wake-up word provided by an embodiment of the present invention;

[0057] Figure 11 It is a flow chart of the overall structure of the wake-up process provided by an embodiment of the present invention;

[0058] Figure 12 It is a schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0059] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present invention. They are only examples of devices and methods consistent with some aspects of the embodiments of the present invention described in detail in the appended claims.

[0060] It should be noted that although the functional modules are divided in the system schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division from that in the system or a different sequence from that in the flowchart. The terms "first / S100" and "second / S200" in the description, claims and the above-mentioned drawings may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "when" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0061] The terms "at least one", "a plurality of", "each", "any one" and the like used in the present invention, where "at least one" includes one, two or more than two, "a plurality of" includes two or more than two, "each" refers to each one of the corresponding plurality, and "any one" refers to any one of the plurality.

[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.

[0063] Before the embodiments of the present invention are described in detail, some nouns and terms involved in the embodiments of the present invention are first described, and the nouns and terms involved in the embodiments of the present invention are applicable to the following explanations.

[0064] A finite state automaton (FSA) is a mathematical model used to describe and process transitions between a finite number of states.

[0065] A finite state transducer (FST) is an extended finite state automaton (FSA) that can perform conversions between two symbol sets. Each state transition of the FST has an input label and an output label, and through these label pairs, the FST can convert an input sequence into an output sequence.

[0066] Connectionist Temporal Classification (CTC) is a training and decoding method for sequence modeling tasks, especially suitable for cases where the lengths of the input and output sequences are different.

[0067] OOV (Out Of Vocabulary) refers to words that did not appear in the training data but appear in the test data or in actual applications. These words may be new words, rare words, misspellings, proper nouns, or words in languages not covered by the model.

[0068] Intelligent terminal devices with voice as the entry have been widely used in various fields, including mobile phones, computers, televisions, and smart home devices, almost all of which support voice wake-up and voice control. As a pre-module of speech recognition, the voice wake-up module is independently deployed from the speech recognition system and is responsible for connecting the user's voice input and recognition services to achieve seamless interaction. Traditional wake-up words usually only listen for fixed phrases. To save resources, such modules tend to be as lightweight as possible and are usually deployed on hardware accelerators.

[0069] With the in-depth application of voice interaction in televisions and smart homes, more and more set-top boxes also support voice wake-up and control functions. Currently, users generally input commands by holding down the voice key on a Bluetooth voice remote control. To further optimize the user experience and simplify the operation process, the multi-wake-up word function can be introduced, enabling users to directly wake up the set-top box without having to trigger it by pressing a button each time. This design is not only convenient but also more in line with the trend of user-friendly interaction in smart home devices.

[0070] However, to achieve a stable multi-wake-up word function on offline devices, the following problems are faced:

[0071] 1. The wake-up words cannot be flexibly replaced

[0072] Most current voice wake-up products usually preset fixed wake-up words or support switching between a very small number (such as two or three) of wake-up words. Users cannot name the device independently, which is lacking in terms of personalization and user experience. Especially in scenarios such as smart homes and in-vehicle devices, users hope to be able to customize wake-up words at will. However, the challenge of achieving flexible replacement of wake-up words lies in the need to dynamically load a large amount of feature data for different wake-up words and ensure the recognition accuracy of each wake-up word, which is extremely difficult.

[0073] 2. High requirements for multi-wake-up word model data

[0074] To achieve accurate recognition of multi-wake-up words, a large number of wake-up word voice samples and noise data in specific scenarios are required for training and optimization. Especially in multi-user and multi-scenario environments, the phoneme features of different wake-up words may be similar. If the data is insufficient, the model is prone to false wake-up or missed wake-up. In addition, the collection and annotation of multi-wake-up word training data also require a high cost investment, which poses a challenge to improving the adaptability of the multi-wake-up model.

[0075] 3. High recognition latency of the multi-wake-up word model

[0076] Multi - wake - word models need to monitor multiple specific words, resulting in a significant increase in computational complexity and higher latency compared to single - wake - word models. This is particularly prominent in offline devices. Due to limited offline processing capabilities, the scale of model parameters and computational overhead need to be minimized as much as possible. However, this also means that a balance must be struck between recognition accuracy and recognition speed to avoid a poor user experience caused by wake - up latency. In addition, the latency problem may be further amplified in multi - task scenarios (such as smart homes), affecting the fluency of interaction.

[0077] In view of this, as Figure 1 shown, embodiments of the present invention provide a multi - wake - word recognition method based on a finite - state machine. The method may include, but is not limited to, steps S100 to S700:

[0078] Step S100, preset a target wake - word.

[0079] Step S200, construct a target finite - state transducer graph according to the target wake - word.

[0080] Step S300, in response to the input speech, extract features from the input speech to obtain feature frames.

[0081] Step S400, input the feature frames into an acoustic model to generate a first streaming result.

[0082] Step S500, input the first streaming result into the target finite - state transducer graph for decoding to obtain a recognition status result.

[0083] In steps S100 to S500 of some embodiments, a new and efficient voice wake - up method is designed. Modeling is achieved through an end - to - end acoustic model, eliminating the need for precisely aligned data, thus automatically generating frame - level label alignment. In addition, a keyword graph is constructed based on a finite - state machine, and an orderly state transition supports the entire process from monitoring to recognition. And when inputting the first streaming result into the target finite - state transducer graph for decoding, while ensuring the streaming output of the speech recognition result, it quickly switches to the recognition state to obtain the recognition status result, improving the recognition accuracy and response speed of multi - wake - words.

[0084] In step S100 of some embodiments, in the voice wake - up setting, any target wake - word can be preset, such as "Hello, Xiaoyi". Users can customize the wake - word by naming the device independently, improving the personalization of voice interaction recognition and the user experience, and enabling flexible replacement of the wake - word.

[0085] In existing traditional weighted finite - state transducers (WFSTs), an acoustic model mapping graph (H), a context relationship graph (C), a dictionary mapping graph (L), and a language model graph (G) are included. AsFigure 2 The WFST shown: Each WFST contains a set of states (nodes) and directed jumps (edges). The input label, output label, and weight are saved on the edges in the format of "input: output / weight". Each graph has at least one start state (such as the "0" node in the figure, represented by a bold single circle) and at least one end state (such as the "8", "12", and "4" nodes, represented by double circles). The weight is statistically calculated by the language model and is represented in the form of negative logarithm in actual use for convenient addition calculation of path costs (such as "you: you / 0.6" indicating that the probability of this path is approximately 0.6).

[0086] In step S200 of some embodiments, based on the structure of a set of states and the directed jumps between states, according to the preset target wake-up word, a target finite state transducer (FST) graph is constructed. The target FST graph does not require additional statistics of the weights based on the language model and only needs to use the probabilities output by the acoustic model to calculate the costs on the paths in the subsequent process, making the structure of the target FST graph lighter and more intuitive. And through the optimization process of the FST graph, the fault tolerance rate and search efficiency of the target FST graph are improved.

[0087] In some embodiments, the FST graph may include a wake-up word or command words. Among them, all characters that appear in the wake-up word or command words must appear in the FST graph and be able to complete self-looping (i.e., self-loop), and the remaining characters are replaced by OOV to ensure smooth decoding. Multiple wake-up words and different types of command words can be set, and similar sounds, front and back nasal sounds, etc. can also be added according to the actual situation to improve the fault tolerance rate. In the subsequent process of decoding using the FST graph, it is divided into two types of triggers: wake-up words and command words. As Figure 3 shown, the wake-up word and different types of command words are distinguished from the configuration file, and the wake-up word and command words need to satisfy the serial trigger of having the wake-up word first and then the corresponding command word, so as to reduce false wake-up. Different command words satisfy parallel trigger, and different state node codes represent different commands.

[0088] In some embodiments, step S200 may include but is not limited to steps S210 to S240:

[0089] Step S210, obtain the state node according to the first character of the target wake-up word;

[0090] Step S220, obtain the directed jump edges between the state nodes;

[0091] Step S230, construct an initial finite state transducer graph according to the state node and the directed jump edges;

[0092] Step S240, perform an optimization process on the initial finite state transducer graph to obtain the target finite state transducer graph.

[0093] In step S210 of some embodiments, according to the characters of the target wake-up word, state nodes of the FST graph are generated in real time. Exemplarily, as Figure 3 shown, if the target wake-up word is "Hello, Xiaojie", corresponding state nodes are generated for each character. The state node code of the state node corresponding to "Hello" is 1, the state node code of the state node corresponding to "good" is 2, the state node code of the state node corresponding to "Hello" is 3, and the state node code of the state node corresponding to "Jie" is 4.

[0094] In steps S220 to S230 of some embodiments, directed jump edges between state nodes are obtained, and conversion rules between states are defined, which determine how to transfer from one state to another state and can carry additional information, such as input symbols and output symbols. Through each state node and the directed jump edges between state nodes, an initial finite state transducer graph can be constructed.

[0095] In step S240 of some embodiments, the initial finite state transducer graph is optimized for fault-tolerant paths and graph determinization to obtain a target finite state transducer graph, which improves the fault tolerance rate and search efficiency of the finite state transducer graph.

[0096] In some embodiments, step S240 may include but is not limited to steps S241 to S244:

[0097] Step S241, obtaining a fault-tolerant wake-up word;

[0098] Step S242, obtaining a fault-tolerant node according to the second character of the fault-tolerant wake-up word;

[0099] Step S243, adding the fault-tolerant node and unknown character edges to the initial finite state transducer graph to obtain an intermediate finite state transducer graph;

[0100] Step S244, setting self-loops for the state nodes and the fault-tolerant nodes of the intermediate finite state transducer graph, and determinizing the intermediate finite state transducer graph to obtain the target finite state transducer graph.

[0101] In steps S241 to S242 of some embodiments, according to the character pronunciation of the target wake-up word, characters with similar sounds or differences in front and back nasals are obtained as the characters of the fault-tolerant wake-up word, and the characters of the fault-tolerant wake-up word are used as the fault-tolerant state nodes of the finite state transducer graph. Optionally, in the command word part of the finite state transducer graph, a fault-tolerant wake-up word can also be set, and the characters of the fault-tolerant command word are set according to the meaning of the target command word, the character pronunciation of the target command word, etc. For example, the commonly used "Make the sound louder" and "Increase the volume" both mean to increase the volume. Exemplarily, asFigure 4 As shown, the target command word is "previous song", and the error-tolerant command words with the same meaning as "previous song" can include "upper song". Then, the target path is "up" → "one" → "song" (such as the blue path shown in Figure 4 ), and the error-tolerant path is "up" → "song" (such as the red path shown in Figure 4 ). By setting the error-tolerant wake-up word and error-tolerant command word, the problem of missing insignificant characters in the acoustic model prediction can be solved, and the user's voice interaction experience can be improved.

[0102] In step S243 of some embodiments, the obtained error-tolerant nodes are added to the initial finite state transducer graph, and unknown character edges OOV can also be added to the wake-up word part and command word part of the initial finite state transducer graph to obtain an intermediate finite state transducer graph, so as to cope with the filler words and the like included in the input speech due to the user's speaking habits. By adding the unknown character edges OOV, the subsequent FST graph can continue to decode when encountering unknown words, rather than directly decoding failure, thus improving the error tolerance rate. Exemplarily, the set target wake-up word and target command word are "Hello Xiaojie, next song". By adding the unknown character edges OOV, an unknown character path of the red path shown in Figure 5 can be obtained, so that the user's wake-up word and command word "Hello, um, Xiaojie, please play the next song" can be correctly recognized. Then, the target path is where ε is an empty character state node.

[0103] In step S244 of some embodiments, self-loops are set for each state node and error-tolerant node of the finite state transducer graph. By setting the self-loops, unknown characters can rotate and consume at the start node (such as the nodes 0 and 4 shown in Figure 6 ), and the single characters that are repeatedly predicted can be consumed at the current state. Exemplarily, as shown in the red self-rotation path in Figure 6 , for "XXX Hello Xiaojie, next next one one one song", the unknown character "XXX" can rotate and consume in the self-loop of the start node 0, and the single characters "next next" and "one one one" that are repeatedly predicted can be consumed in the self-loops of the state nodes 14 and 15 respectively. Thus, when decoding using the finite state transducer graph subsequently, the wake-up word "Hello Xiaojie" and the command word "next song" can be obtained. And the intermediate finite state transducer graph is also determinized to obtain the target finite state transducer graph. The determinization of the graph means that for a given input symbol, there is only one transition state corresponding to it. The determinized graph can complete the conversion in a time proportional to the number of input symbols, improving the search efficiency, which is crucial for real-time processing. Exemplarily, the longest common edge is shared on the same wake-up word / command word, such as Figure 7As shown, the blue edges are the longest common edges shared for the determination of the graph on the same command word / wake-up word. "Hello Xiaojie" and "Hello Xiaoao" share the edge of the three characters "Hello Xiao". "Previous" and "Up" share "Up".

[0104] In some embodiments, the wake-up word and the command word are in series, such as Figure 8 As shown, the state node 4 is the wake-up state of the wake-up word. On the left side of this node is the graph formed by all wake-up words, and on the right side are two types of commands, "Previous" and "Next". There may be multiple error-tolerant paths in the same command, but there is only one end node. In the subsequent decoding process, when the decoding algorithm returns different node states, different types of instructions can be fed back.

[0105] In step S300 of some embodiments, in response to the user's input voice, feature frames are extracted from the audio signal of the input voice, and these feature frames will be used as the input of the acoustic model. Optionally, the input voice can be preprocessed, such as noise reduction, echo cancellation, etc., but not limited to this. During the feature extraction process, the audio signal of the input voice can be divided into frames of a fixed size (for example, each frame is 25 ms, and the frame shift is 10 ms, not limited to this).

[0106] In some embodiments, before the feature frames are input into the acoustic model, the acoustic model is also pre-trained. In the pre-trained acoustic model, training data processing and modeling granularity selection are also performed.

[0107] For training data processing, the embodiments of the present invention adopt end-to-end acoustic model modeling, which can widely utilize existing open-source data resources and improve the robustness and adaptability of the model through diverse data augmentation techniques. Optionally, in terms of processing, all data is uniformly converted into a single-channel format with a sampling rate of 16,000 Hz, and enhancement means such as Gaussian noise, colored noise, reverberation, variable speed and pitch, and distortion are introduced during training to simulate various channel noises and accent changes, so that the model can maintain higher recognition accuracy for the speech inputs of different people in complex environments.

[0108] For modeling granularity selection, the embodiments of the present invention select pinyin combination units suitable for wake-up word and command word recognition for shorter speech segments (such as instructions within four characters), effectively reducing the data requirements and model complexity. Optionally, through un-toned pinyin combination modeling, 100 to 150 of the most commonly used un-toned pinyins are selected from daily Chinese characters, including common pronunciations such as "shi", "zhi", "yi", "li", "yang", etc., which can cover most common instruction words while maintaining recognition accuracy. This pinyin modeling method not only simplifies the model structure but also improves the recognition robustness, especially performing well on resource-constrained devices.

[0109] Before pre-training the acoustic model, an end-to-end acoustic model structure was constructed based on the multi-layer convolutional network and CTC loss function of the Jasper-like model. This network structure only uses one-dimensional convolution, BatchNormalization, ReLU activation, Dropout layer, and residual connection. In the design, it contains multiple blocks, each block consists of multiple sub-blocks, and the input of each block is passed to the last sub-block of the block through the residual connection. As Figure 9 shown, the acoustic model network structure of the embodiment of the present invention is clear and has a hierarchical depth. By implementing acoustic modeling through a fully convolutional network, compared with traditional recurrent networks (such as RNN or LSTM), it has significant parallel processing advantages, which greatly accelerates the model in both the training and inference stages. At the same time, the complexity of the model structure is effectively reduced, the memory requirement is reduced, which is particularly important for the rapid recognition of wake words. Under this design, the deep convolutional network can more accurately capture the subtle differences of different wake words, laying a foundation for achieving high-precision speech wake-up tasks.

[0110] During the pre-training of the acoustic model, CTC is used as the loss function. This is a completely end-to-end acoustic model training, which does not require pre-aligning the data in advance. Only an input sequence and an output sequence are needed for training. In this way, there is no need to align and label the data one by one, and CTC directly outputs the probability of sequence prediction without external post-processing.

[0111] Exemplarily, there are the following settings:

[0112] The input sequence X = (x1, x2,..., x T ) has T time steps;

[0113] The output sequence Y = (y1, y2,..., y u ), with a length of u, and u ≤ T;

[0114] CTC introduces a blank symbol (blank, usually represented by "_") as the padding between target characters to support flexible alignment. Then the calculation steps of the loss are roughly as follows: Given the target sequence Y, CTC will generate multiple possible alignment paths π, and each path is a way of aligning Y. A path may contain the blank symbol "_" (for example, mapping "ni hao xiaojie" to "ni hao xiao jie" or "ni_hao_xiao_jie").

[0115] For the target sequence Y, the probability of each path is calculated through the conditional probability of the input sequence X, and the probability of each path is:

[0116]

[0117] Among them, P(π t |x t ) represents the probability of predicting symbol π t at time step t.

[0118] The CTC loss calculates the total probability P(Y|X) of the target sequence by summing the probabilities of all possible alignment paths:

[0119] P(Y|X) = ∑ π∈align(Y) P(π|X)

[0120] where align(Y) is the set of all possible paths that align with Y.

[0121] The definition of the CTC loss function CTC loss is the logarithmic loss of the total probability of the target sequence:

[0122] CTC loss = -log(P(Y|X))

[0123] In the above manner, the Jasper model can be directly trained from the input sequence and the target sequence without external alignment and annotation.

[0124] In some embodiments, step S400 may include but is not limited to steps S410 to S460:

[0125] Step S410, input the feature frames into the acoustic model, and output the character probability distribution at each time step;

[0126] Step S420, through the CTC decoding algorithm, obtain the path probability according to the character probability distribution;

[0127] Step S430, through the CTC decoding algorithm, sum all the path probabilities to obtain the target sequence probability;

[0128] Step S440, obtain the initial character sequence according to the target sequence probability;

[0129] Step S450, through the CTC decoding algorithm, delete the blank symbols and consecutive repeated third characters in the initial character sequence to obtain the target character sequence;

[0130] Step S460, splice several of the target character sequences to obtain the first streaming result.

[0131] In steps S410 to S440 of some embodiments, the feature frames are input into a pre-trained acoustic model, and the acoustic model can output the character probability distribution P(π t |xt )。For each time step, through the CTC decoding algorithm, the path probability P(π|X) is calculated according to the character probability distribution. By summing up all possible path probabilities, the target sequence probability is calculated, and the initial character sequence corresponding to the maximum target sequence probability is output (such as "__ni ni__hao hao_").

[0132] In step S450 of some embodiments, through the CTC decoding algorithm, the blank symbols and consecutive repeated characters in the initial character sequence are deleted to obtain the target character sequence. Exemplarily, by deleting the blank symbols and consecutive repeated characters in the initial character sequence "__ni ni__hao hao_", the target character sequence "nihao" can be obtained.

[0133] In step S460 of some embodiments, several target character sequences are concatenated to obtain the first streaming result. Exemplarily, after inputting the current frame and decoding it through the CTC decoding algorithm, the target character sequence "ni hao" of the first batch of streaming speech decoding can be obtained. Then, after inputting the next frame and decoding it through the CTC decoding algorithm, the target character sequence "xiao" of the second batch of streaming speech decoding can be obtained. Finally, after inputting the last frame and decoding it through the CTC decoding algorithm, the target character sequence "yi" of the third batch of streaming speech decoding can be obtained. Then, the first streaming result including "ni hao", "xiao", and "yi" can be obtained. The first streaming result is a set of target character sequences of the streaming speech decoding that processes each frame in real-time CTC decoding. In real-time CTC decoding, each time a frame is processed, the CTC decoder will output the result (such as pinyin) of the current batch of streaming speech decoding and temporarily store it, removing the blank character "_" and consecutive repeated non-toned pinyin.

[0134] In step S500 of some embodiments, the first streaming result is input into the target finite state transducer graph for decoding. The OneShot decoding method is adopted, which does not require storing the complete historical information and only records the current node path state thrown each time. This method simplifies the decoding process and improves the operation efficiency. When it is necessary to calculate the cumulative confidence triggered by the wake word, only the sum of the weights output by the acoustic models of the valid trigger paths needs to be summarized, so as to effectively estimate the confidence, while reducing the storage and calculation complexity. This method greatly optimizes the resource consumption and improves the decoding speed, and is particularly suitable for scenarios with high real-time requirements such as wake word detection. Corresponding to Figure 8 , the pseudo-code of the OneShot decoding algorithm is as follows:

[0135] Input: The recognition result word of the acoustic model

[0136] Output: The corresponding output state point

[0137] 1. start: Starting state

[0138] 2. For each outgoing edge of the starting state:

[0139] 3. If the character corresponding to a certain outgoing edge = word:

[0140] 4. The starting state = the next state pointed to by this outgoing edge

[0141] 5. Return the starting state

[0142] 6. The starting state = the fallback state of this starting state

[0143] 7. Return to step 1

[0144] In some embodiments, step S500 may include but is not limited to steps S510 to S530:

[0145] Step S510, read the fourth character of the current time step of the first streaming result;

[0146] Step S520, obtain the fifth character of the outgoing edge of the current state of the target finite state transducer graph;

[0147] Step S530, determine whether the fifth character matches the fourth character. If the fifth character does not match the fourth character, reset the current state to the initial state. If the fifth character matches the fourth character, update the current state to the next state pointed to by the outgoing edge, read the fourth character of the next time step, and return to the step of obtaining the fifth character of the outgoing edge of the current state of the target finite state transducer graph until all the fourth characters are read to obtain the recognition state result.

[0148] In steps S510 to S530 of some embodiments, input the first streaming result into the target finite state transducer graph, read the fourth character of the current time step of the first streaming result. For the fifth character of each outgoing edge of the current state of the target finite state transducer graph, if the fourth character matches the fifth character, point the current state node to the state node corresponding to the outgoing edge of the matching fifth character. At this time, the current state is updated to the state corresponding to the outgoing edge of the matching fifth character, and continue to read the fourth character of the next time step of the first streaming result, return to the step of obtaining the fifth character of the outgoing edge of the current state of the target finite state transducer graph until all the fourth characters are read to obtain the recognition state result. Exemplarily, as Figure 8As shown, for each outgoing edge of the start node, in a loop, if the fourth character "ni" of the current time step of the first streaming result read matches the fifth character "ni" of the outgoing edge of the start node 0 of the FST graph, the start node is pointed to the node 1 corresponding to the outgoing edge. Then, the fourth character "hao" of the next time step of the first streaming result is read. If the fourth character "hao" matches the fifth character "hao" of the outgoing edge of the node 1 of the FST graph, the node 1 is pointed to the node 2 corresponding to the outgoing edge, and so on until all the fourth characters are read. In some embodiments, during the process of determining whether the fifth character matches the fourth character, if the fifth character does not match the fourth character, the current state is reset to the initial state. Exemplarily, as Figure 8 shown, when the fourth character read in the next time step of the first streaming result is "bu", and "bu" does not match the fifth character of any outgoing edge of the node 1, the current state is reset to the start state 0. Optionally, as Figure 8 shown, when the fourth character read in the next time step of the first streaming result is the unknown character "oov", and the fourth character "oov" matches the fifth character "oov" of the outgoing edge of the node 1, the current state node 1 is pointed to the node 5 corresponding to the outgoing edge. Through the fault tolerance of the FST graph, the word outside the vocabulary identified as the unknown character "oov" is consumed through this outgoing edge.

[0149] In some embodiments, a trigger status code can be generated according to the recognition status result. Exemplarily, for the target wake-up word "Hello, Xiaoyi", when the target character sequence "ni hao" in the first batch of streaming speech decoding passes through the FST graph oneshot decoding, the result of the first batch of FST graph oneshot decoding can be obtained, that is, the recognition status result is an intermediate state, and its corresponding trigger status code is the non-wake-up state; when the target character sequence "xiao" in the second batch of streaming speech decoding passes through the FST graph oneshot decoding, the result of the second batch of FST graph oneshot decoding can be obtained, that is, the recognition status result is an intermediate state, and its corresponding trigger status code is the non-wake-up state; when the target character sequence "yi" in the third batch of streaming speech decoding passes through the FST graph oneshot decoding, the result of the second batch of FST graph oneshot decoding can be obtained, that is, the recognition status result is an end state, and its corresponding trigger status code is the successful wake-up state.

[0150] As Figure 10 shown, the overall process of flexibly customizing the wake-up word in the embodiments of the present invention is as follows:

[0151] Step a1: The user inputs and / or speaks the custom wake-up word "Hello, Xiaoyi", and the system captures this audio signal through an audio input device such as a microphone.

[0152] Step a2: The captured audio signal is sent to an acoustic model for processing, which converts the audio signal into a corresponding sequence of pinyin or phonemes, for example, outputting "ni hao xiao yi".

[0153] Step a3: Using the output of the acoustic model, construct an FST graph. This graph defines how to convert the sequence of phonemes output by the acoustic model into the user-defined wake-up word text. During the construction of the FST graph, it also includes:

[0154] Step a31: Add error-tolerant paths to the FST graph to improve the error tolerance rate;

[0155] Step a32: Perform determinization on the FST graph to enable the FST graph to be decoded more efficiently. By determinization, a non-deterministic state machine is converted into a deterministic state machine to ensure that each input corresponds to a unique output;

[0156] Step a33: Minimize the determinized FST graph, removing redundant states and edges in the graph to reduce the complexity of the graph and improve the decoding speed and efficiency.

[0157] As Figure 11 shown, the overall process of the wake-up process in the embodiment of the present invention is as follows:

[0158] Step b1: The user inputs the voice "ni hao xiao yi".

[0159] Step b2: The input voice audio is input into the system in a streaming manner. Each small segment of audio (usually dozens of milliseconds) generates a feature frame through feature extraction and is input into the acoustic model.

[0160] Step b3: According to the maximum target sequence probability, the streaming acoustic model outputs an initial character sequence, such as "__nini__hao hao_"; in real-time CTC decoding, for each frame processed, the CTC decoder will output the pinyin of the current frame and temporarily store it. Removing the blank character "_" and consecutive repeated non-toned pinyin, the target character sequence "ni hao" can be obtained and sent to step b4; then continue to process the next frame and decode, the target character sequence "xiao" can be obtained. Similarly, then the target character sequence "yi" can be obtained, and the first streaming result is "ni hao xiao yi".

[0161] Step b4: Input the first streaming result into the FST graph for oneshot decoding (as Figure 8As shown in the figure, the recognition status result is obtained: for each out-edge of the start node in a loop, if the first input "ni" = the out-edge "ni" of the FST start node 0, then the start node points to the node 1 corresponding to the out-edge; if the second input "hao" = the out-edge "hao" of the FST node 1, then the node 1 points to the node 2 corresponding to the out-edge; if the second input unknown word "oov" = the out-edge "oov" of the FST node 1, then the node 1 points to the node 5 corresponding to the out-edge; if the two inputs are "bu"!= any out-edge of the FST start node 1, then the current state is reset to the start state node 0.

[0162] Step b5: When the recognition status result is the end state, trigger the status code as the successful wake-up state.

[0163] Through the above process, the embodiment of the present invention combines the Oneshot decoding of the finite state machine in the streaming input speech stream, judges in real time whether the user-defined wake-up word is matched, and generates a trigger status code, realizing efficient and accurate wake-up detection.

[0164] The embodiment of the present invention also provides a multi-wake-up word recognition device based on a finite state machine, which can implement the above-mentioned multi-wake-up word recognition method based on a finite state machine. The device includes:

[0165] The first module is used to preset the target wake-up word;

[0166] The second module is used to construct a target finite state transducer graph according to the target wake-up word;

[0167] The third module is used to extract features from the input speech in response to the input speech to obtain feature frames;

[0168] The fourth module is used to input the feature frames into an acoustic model to generate a first streaming result;

[0169] The fifth module is used to input the first streaming result into the target finite state transducer graph for decoding to obtain a recognition status result.

[0170] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0171] The embodiment of the present invention also provides an electronic device, which includes a processor and a memory. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned multi-wake-up word recognition method based on a finite state machine. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0172] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented in the device embodiments of the present invention are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0173] Reference Figure 12 , Figure 12 illustrates the hardware structure of an electronic device according to another embodiment. The electronic device includes:

[0174] A processor 601, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention;

[0175] A memory 602, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 602 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 602 and are called by the processor 601 to execute a multi-wake word recognition method based on a finite state machine according to the embodiments of the present invention;

[0176] An input / output interface 603, which is used to implement information input and output;

[0177] A communication interface 604, which is used to implement communication and interaction between the present device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0178] A bus 605, which transmits information between various components of the device (such as the processor 601, the memory 602, the input / output interface 603, and the communication interface 604);

[0179] Among them, the processor 601, the memory 602, the input / output interface 603, and the communication interface 604 are communicatively connected to each other inside the device through the bus 605.

[0180] An embodiment of the present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above-mentioned multi wake-word recognition method based on a finite state machine.

[0181] It can be understood that the content in the above method embodiments is applicable to this storage medium embodiment. The functions specifically implemented by this storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0182] An embodiment of the present invention also provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute the above-mentioned multi wake-word recognition method based on a finite state machine.

[0183] In summary, the multi wake-word recognition method, device, equipment and medium of the embodiment of the present invention have the following advantages:

[0184] 1. The embodiment of the present invention adopts an end-to-end CTC acoustic model to reduce the data annotation cost. Training is achieved through open-source data resources without the need for precisely aligned data, thus automatically generating frame-level label alignment, avoiding the dependence on precisely aligned data, and significantly reducing the data annotation and customization costs. Moreover, the embodiment of the present invention improves the robustness and adaptability of the model through diverse data augmentation techniques. In data processing, all data is unified into a single-channel format with a sampling rate of 16000Hz. During training, enhancement means such as Gaussian noise, colored noise, reverberation, variable speed and pitch, and distortion are introduced to simulate various channel noises and accent changes, enabling the model to maintain higher recognition accuracy for speech inputs from different people in complex environments, not only reducing the data cost but also further enhancing the adaptability of the model under multi-scene and multi-accent conditions.

[0185] 2. The traditional phoneme-level modeling has a relatively small granularity and is suitable for multi-dialect processing, but there are limitations in speech recognition accuracy; the word-level unit avoids the conversion between phonemes and Chinese characters, but its quantity is huge and the training cost is high. In contrast, the embodiment of the present invention selects untoned pinyin combinations as the modeling unit, which can cover most common command words while maintaining recognition accuracy. This pinyin modeling method not only simplifies the model structure, reduces the computational complexity, but also improves the recognition robustness, which is beneficial to the deployment of resource-constrained devices.

[0186] 3. Compared with traditional keyword training methods that often rely on highly customized, specially labeled high-quality data, which have high data costs and long generation cycles, the embodiments of the present invention implement acoustic modeling through a Jasper-like fully convolutional network. Compared with traditional recurrent networks (such as RNN or LSTM), it has significant parallel processing advantages, not only improving the training and inference speed, but also significantly reducing the memory requirements, and is suitable for wake-up tasks with high real-time requirements. Under this design, the deep convolutional network can more accurately capture the subtle differences of different wake-up words, laying a foundation for achieving high-precision speech wake-up tasks.

[0187] 4. The lightweight design of the FST graph in the embodiments of the present invention further reduces the size of the graph, making it more suitable for deployment on embedded devices or mobile terminals. This composition strategy effectively reduces the memory usage and computational load while ensuring accuracy, meeting the resource limitations of the device.

[0188] 5. The embodiments of the present invention complete the composition of keywords based on a finite state machine, supporting the entire process from monitoring to recognition with an orderly state transition. And an innovative oneshot decoding algorithm is proposed, which can quickly switch to the recognition state while ensuring the streaming output of speech recognition results. The recognition state is divided into an intermediate state and an end state. After reaching the end state, the end side can trigger a successful wake-up, realizing fast state switching and wake-up response. This not only reduces the resource consumption during the decoding process, but also improves the real-time performance of the system, further enhancing the recognition accuracy and response speed.

[0189] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated where the order of various operations is changed and where sub-operations described as part of a larger operation are executed independently.

[0190] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features described may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Thus, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0191] If the described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0192] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device). For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0193] More specific examples (nonexhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0194] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0195] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0196] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

[0197] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.

Claims

1. A multi-wake-up word recognition method based on a finite state machine, characterized in that: The following steps are involved: Pre-set target wake-up words; According to the target wake-up word, construct a target finite state converter graph; In response to input speech, extracting features of the input speech to obtain a feature frame; Inputting the feature frame into an acoustic model to generate a first streaming result; The first streaming result is input into the target finite state converter diagram for decoding to obtain a recognition state result.

2. According to the finite state machine-based multi-wake-up word recognition method of claim 1, it is characterized in that: The following steps are also included: A trigger status code is generated according to the identification status result.

3. According to the finite state machine-based multi-wake-up word recognition method of claim 1, it is characterized in that: The step of constructing a target finite state converter graph according to the target wake-up word comprises the following steps: According to the first character of the target wake-up word, a state node is obtained; Obtaining directed jump edges between the state nodes; Constructing an initial finite state converter graph according to the state nodes and the directed jump edges; The initial finite state converter graph is optimized to obtain the target finite state converter graph.

4. According to the finite state machine-based multi-wake-up word recognition method of claim 3, it is characterized in that: The step of optimizing the initial finite state converter graph to obtain the target finite state converter graph comprises the following steps: Get the error-tolerant wake-up word; Obtaining a fault-tolerant node according to the second character of the fault-tolerant wake-up word; Adding the fault-tolerant node and the unknown character edge to the initial finite state converter graph to obtain an intermediate finite state converter graph; Self-loops are set for the state nodes and the fault-tolerant nodes of the intermediate finite state converter graph, and the intermediate finite state converter graph is determinized to obtain the target finite state converter graph.

5. According to the finite state machine-based multi-wake-up word recognition method of claim 1, it is characterized in that: The step of inputting the feature frame into the acoustic model to generate a first streaming result comprises the following steps: Input the feature frame into the acoustic model and output the character probability distribution at each time step; Obtaining the path probability according to the character probability distribution through the CTC decoding algorithm; Through the CTC decoding algorithm, all the path probabilities are added together to obtain the target sequence probability; According to the target sequence probability, an initial character sequence is obtained; By using a CTC decoding algorithm, blank symbols and consecutively repeated third characters in the initial character sequence are deleted to obtain a target character sequence; Several target character sequences are concatenated to obtain the first streaming result.

6. According to the finite state machine-based multi-wake-up word recognition method of claim 1, it is characterized in that: The step of inputting the first streaming result into the target finite state converter graph for decoding to obtain a recognition state result comprises the following steps: Read the fourth character of the current time step of the first streaming result; Obtain the fifth character of the outgoing edge of the current state of the target finite state converter graph; Determine whether the fifth character matches the fourth character. If the fifth character does not match the fourth character, reset the current state to the initial state. If the fifth character matches the fourth character, update the current state to the next state pointed to by the outgoing edge, read the fourth character of the next time step, and return to the step of obtaining the fifth character of the outgoing edge of the current state of the target finite state converter diagram until all the fourth characters are read to obtain the recognition state result.

7. The method for multiple wake-up word recognition based on a finite state machine according to claim 2, characterized in that: The step of generating a trigger status code according to the identification status result comprises the following steps: When the identification state result is an intermediate state, the trigger state code is an unawakened state; When the identification status result is an end status, the trigger status code is a successful awakening status.

8. A multi-wake-up word recognition device based on a finite state machine, characterized in that: include: The first module is used to pre-set the target wake-up word; A second module is used to construct a target finite state converter diagram according to the target wake-up word; A third module is used for extracting features of the input speech in response to the input speech to obtain a feature frame; A fourth module is used to input the feature frame into an acoustic model to generate a first streaming result; The fifth module is used to input the first streaming result into the target finite state converter diagram for decoding to obtain a recognition state result.

9. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 7.