Speech Recognition Decoding Method, Device, Storage Medium and Computer Equipment
By building a dynamic decoding network in the dynamic decoding process of speech recognition, using blank arc propagation and self-jumping technology, the problems of excessive decoding network and path redundancy are solved, and effective resource compression and recognition efficiency are improved.
Patent Information
- Application Number
- CN202111529639.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-12-14
AI Technical Summary
In the existing speech recognition technology, the decoding network is too large and the decoding path is redundant, resulting in waste of resources and inefficient recognition.
By obtaining the built static decoding network, determining the activation node and activation arc during the dynamic decoding process, controlling the propagation of the blank arc, and performing blank arc self-jumping on the termination node, thereby building a dynamic decoding network to reduce the resource occupation of the static decoding network.
Effectively compress the decoding network, reduces decoding path redundancy, improves recognition efficiency, and reduces memory usage.
Smart Images

Figure CN114155837B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of speech signal processing, and particularly to a speech recognition decoding method, apparatus, storage medium, and computer device. Background Art
[0002] Speech recognition is a technology that converts an audio segment into its corresponding text through certain technical means. Current speech recognition technology is mainly divided into two parts: acoustic calculation and decoding. After an audio segment is segmented into a series of speech frames and the feature vectors are extracted frame by frame, the acoustic model calculates the acoustic posterior based on the feature vector; in the series of acoustic posteriors output by the acoustic model, decoding is performed in combination with a language model (i.e., a decoding network), and during the decoding process, the most matching output sequence is searched, and this sequence is the recognition result.
[0003] Compared with the traditional acoustic model of speech recognition, for each frame, the traditional acoustic model gives a clear acoustic posterior, while the Connectionist Temporal Classification (CTC) acoustic model adds a blank phoneme state. When the state of some frames cannot be clearly attributed to a certain phoneme state, blank is used to represent the output state when these states are uncertain.
[0004] However, the current technology has the disadvantages of an overly large decoding network and redundant decoding paths. Summary of the Invention
[0005] Embodiments of this application provide a speech recognition decoding method, apparatus, storage medium, and computer device, which can effectively compress the decoding network and reduce redundant decoding paths.
[0006] On the one hand, a speech recognition decoding method is provided, and the method includes:
[0007] Obtain a static decoding network that has been constructed;
[0008] During the process of decoding target speech data, determine active nodes and active arcs according to the nodes that have been traversed and the out-arcs of the nodes that have been traversed in the static decoding network, so as to obtain a dynamic decoding network including the active nodes and the active arcs;
[0009] When accessing a first active node corresponding to the current node on the static decoding network, control the first active node to perform one propagation of a blank arc, and the blank arc is used to mark a blank phoneme;
[0010] When the blank arc is activated and propagated along the blank arc starting from the first activation node, a second activation node is newly created, where the first activation node and the second activation node map to the current node on the static decoding network;
[0011] Propagate the second activation node according to the real arc of the current node on the static decoding network that has a mapping relationship with the second activation node, where the real arc is used to mark the output state of the current node;
[0012] Traverse all nodes on the static decoding network. If the termination node on the static decoding network is accessed, a blank arc self-jump is performed on the termination node.
[0013] Optionally, when accessing the first activation node corresponding to the current node on the static decoding network, controlling the first activation node to perform one propagation of the blank arc includes:
[0014] When accessing the first activation node corresponding to the current node on the static decoding network, propagate according to the information carried by the real arc of the current node;
[0015] After accessing all real arcs of the current node, control the first activation node to perform one propagation of the blank arc, and record on the first activation node whether the first activation node propagates along the blank arc;
[0016] If there is a blank arc on the first activation node, control the first activation node to perform empty arc propagation after the blank arc propagation.
[0017] Optionally, propagating the second activation node according to the real arc of the current node on the static decoding network that has a mapping relationship with the second activation node includes:
[0018] When propagating the second activation node, find the current node on the static decoding network that has a mapping relationship with the second activation node;
[0019] Traverse all out-arcs of the current node on the static decoding network. If there is a real arc among the out-arcs of the current node, propagate the information on the second activation node along the real arc to the successor node of the current node.
[0020] Optionally, when traversing all out-arcs of the current node on the static decoding network, it further includes:
[0021] If all out-arcs of the current node are empty arcs, do not propagate the second activation node.
[0022] Optionally, when accessing the termination node on the static decoding network, performing a self-jump of the blank arc on the termination node includes:
[0023] If a node without an outgoing arc is accessed on the static decoding network, determine the node without an outgoing arc as the termination node on the static decoding network;
[0024] After activating the blank arc connected to the termination node, determine the termination node as the successor node of the blank arc connected to the termination node, so as to perform a self-jump of the blank arc on the termination node.
[0025] Optionally, after creating the second activation node, it further includes:
[0026] Record on the second activation node whether the second activation node is generated by the propagation of the blank arc.
[0027] Optionally, the method further includes:
[0028] Until the activation node corresponding to the termination node that has been traversed in the static decoding network is included in the dynamic decoding network, complete the decoding of the target speech data;
[0029] Determine a decoding path according to the activation nodes and activation arcs in the dynamic decoding network;
[0030] Generate a speech recognition result corresponding to the target speech data according to the decoding path.
[0031] On the other hand, a speech recognition decoding device is provided, and the device includes:
[0032] An acquisition unit, configured to acquire a static decoding network that has been constructed;
[0033] A determination unit, configured to determine activation nodes and activation arcs according to the nodes that have been traversed and the outgoing arcs of the nodes that have been traversed in the static decoding network during the process of decoding target speech data, so as to obtain a dynamic decoding network including the activation nodes and the activation arcs;
[0034] A first propagation unit, configured to control the first activation node corresponding to the current node on the static decoding network to perform one propagation of the blank arc, and the blank arc is used to mark a blank phoneme;
[0035] A new unit, which is used to create a second activation node when propagating along the blank arc starting from the first activation node after the blank arc is activated, where the first activation node and the second activation node map to the current node on the static decoding network;
[0036] A second propagation unit, which is used to propagate the second activation node according to the real arc of the current node on the static decoding network that has a mapping relationship with the second activation node, and the real arc is used to mark the output state of the current node;
[0037] A processing unit, which is used to traverse all nodes on the static decoding network. If the termination node on the static decoding network is accessed, a blank arc self-jump is performed on the termination node.
[0038] On the other hand, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the steps in the speech recognition decoding method described in any one of the above embodiments.
[0039] On the other hand, a computer device is provided. The computer device includes a processor and a memory. A computer program is stored in the memory, and the processor is used to execute the steps in the speech recognition decoding method described in any one of the above embodiments by calling the computer program stored in the memory.
[0040] On the other hand, a computer program product is provided, including computer instructions, and the steps in the speech recognition decoding method described in any one of the above embodiments are implemented when the computer instructions are executed by a processor.
[0041] In the embodiments of the present application, a static decoding network that has been constructed is obtained; during the process of decoding target speech data, activation nodes and activation arcs are determined according to the nodes that have been traversed and the out-arcs of the nodes that have been traversed in the static decoding network, so as to obtain a dynamic decoding network including activation nodes and activation arcs; when accessing the first activation node corresponding to the current node on the static decoding network, the first activation node is controlled to perform one propagation of the blank arc, and the blank arc is used to mark the blank phoneme; when the blank arc is activated and propagated along the blank arc starting from the first activation node, a second activation node is newly created, where the first activation node and the second activation node map to the current node on the static decoding network; according to the real arc of the current node on the static decoding network that has a mapping relationship with the second activation node, the second activation node is propagated, and the real arc is used to mark the output state of the current node; all nodes on the static decoding network are traversed, and if the termination node on the static decoding network is accessed, a self-jump of the blank arc is performed on the termination node. In the embodiments of the present application, by not adding optional blank arcs to the static decoding network and only adding blank arcs to the dynamic decoding network during the dynamic decoding process, the resources of the static decoding network are effectively reduced, the decoding network can be effectively compressed, and only one blank propagation is required for each node, greatly reducing the redundancy of the decoding path. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained according to these drawings.
[0043] Figure 1 Schematic diagram of the static decoding network provided by the embodiments of the present application.
[0044] Figure 2 Schematic diagram of the static decoding network of the CTC acoustic model.
[0045] Figure 3 Another schematic diagram of the static decoding network of the CTC acoustic model.
[0046] Figure 4 Schematic diagram of the process of the speech recognition decoding method provided by the embodiments of the present application.
[0047] Figure 5 Schematic diagram of the first application scenario of the speech recognition decoding method provided by the embodiments of the present application.
[0048] Figure 6 Schematic diagram of the second application scenario of the speech recognition decoding method provided by the embodiments of the present application.
[0049] Figure 7 This is a schematic diagram of the third application scenario of the speech recognition and decoding method provided by the embodiments of the present application.
[0050] Figure 8 This is a schematic diagram of the fourth application scenario of the speech recognition and decoding method provided by the embodiments of the present application.
[0051] Figure 9 This is a schematic diagram of the structure of the speech recognition and decoding device provided by the embodiments of the present application.
[0052] Figure 10 This is a schematic diagram of the structure of the computer device provided by the embodiments of the present application. Detailed implementation manners
[0053] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0054] The embodiments of the present application provide a speech recognition and decoding method, device, computer device and storage medium. Specifically, the speech recognition and decoding method of the embodiments of the present application can be executed by a computer device, where the computer device can be a terminal or a server and other devices. The terminal can be a smart phone, a tablet computer, a notebook computer, a smart TV, a smart speaker, a wearable smart device, a personal computer (Personal Computer, PC) and other devices. The terminal can also include a client, and the client can be a video client, a browser client or an instant messaging client, etc. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.
[0055] The embodiments of the present application can be applied to various scenarios such as artificial intelligence, speech recognition, and intelligent transportation.
[0056] First, some nouns or terms that appear in the process of describing the embodiments of the present application are explained as follows:
[0057] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0058] Among them, the key technologies of speech processing technology include automatic speech recognition technology (ASR), text-to-speech technology (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future.
[0059] The acoustic model (AM) is a differential knowledge representation of acoustics, phonetics, environmental variables, speaker gender, accent, etc. It can be represented by an acoustic model based on the Hidden Markov Model (HMM), such as the Gaussian Mixture Model-Hidden Markov Model (GMM-HMM) and the Deep Neural Networks-Hidden Markov Model (DNN-HMM). The Hidden Markov Model is a weighted finite-state automaton in discrete time domain; of course, it can also be an end-to-end acoustic model, such as the Connectionist Temporal Classification (CTC) model, the Long-Short Term Memory (LSTM) model, and the Attention model.
[0060] Among them, compared with the traditional acoustic model for speech recognition, the traditional acoustic model gives a clear acoustic posterior for each frame, while the CTC acoustic model adds a blank state. When the states of some frames cannot be clearly attributed to a specific phoneme state, the blank state is used to represent the output state when the states are uncertain.
[0061] Since the blank state is added to the output states of the CTC acoustic model, the blank state also needs to be added to the decoding network corresponding to the CTC acoustic model for sequence search. During the decoding process, due to the randomness of the positions where the blank state appears, in the decoding network as shown in Figure 1 it is impossible to determine the positions of the arcs containing the blank state. In the current technical solution, it is chosen to add blanks at all possible positions of the blank state in the decoding network, that is, an optional blank is added to each outgoing arc with an output state in the decoding network. The specific addition method is as follows:
[0062] (1) Traverse the already constructed decoding network. During the traversal, check all the outgoing arcs of the nodes. If there is an outgoing arc that is a real arc, where a real arc means there is a state on the arc, then go to step (2); otherwise, traverse the next node.
[0063] For example, as shown in Figure 1 if there are states state_x and state_y on the two outgoing arcs of node 1 respectively, then go to step (2).
[0064] (2) For the above real arc, insert an optional blank arc before it, that is, a blank arc and an empty arc. During decoding, information can be directly propagated to the target node through the empty arc.
[0065] Among them, the empty arc means that there is no information carried on the arc. The blank arc means that there is information about the blank phoneme carried on the arc.
[0066] For example, as shown in Figure 2 if the outgoing arc with the output state of state_x, insert a new node 4. Insert a real arc with the output state of blank and an empty arc on node 1 respectively. Both arcs are connected to node 4, and connect the outgoing arc with the output state of state_x between node 4 and 2. For the outgoing arc with the output state of state_y, insert a new node 5. Insert a real arc with the output state of blank and an empty arc on node 1 respectively. Both arcs are connected to node 5, and connect the outgoing arc with the output state of state_y between node 5 and 3.
[0067] However, the current technology has the disadvantages of a too large decoding network and redundant decoding paths.
[0068] Compared with the decoding network of the non-CTC acoustic model, the decoding network of the CTC acoustic model adds optional blank arcs on its basis. For the decoding network of the non-CTC acoustic model as shown in Figure 1 expand the decoding network of the non-CTC acoustic model to be as shown in Figure 2When decoding the CTC acoustic model decoding network shown, for Figure 1 a real arc on it, a blank arc, an empty arc, and a new node need to be added. Suppose there are 100 real arcs in a decoding network of a non-CTC acoustic model, then 200 additional arcs need to be added to convert it into a decoding network of a CTC acoustic model, resulting in an overly large decoding network.
[0069] When decoding on the decoding network as shown in Figure 2 and propagating node 1, as shown in Figure 3 it is necessary to propagate the information of node 1 through two blank arcs respectively. If there are N real arcs on a node, then N optional blank arcs will be added to this node. During the decoding process, when the information of this node needs to be propagated backward, the information may be propagated along the N blank arcs, that is, propagated N times on the blank, resulting in redundant decoding paths.
[0070] Therefore, the present application proposes a blank decoding method based on the CTC model, which does not require adding optional blank arcs in the static decoding network, but only adds blank arcs to the dynamic decoding network during the dynamic decoding process to reduce the resources of the static decoding network, and can compress the decoding network of the CTC model on a large scale. In addition, this decoding method can minimize the decoding path redundancy while ensuring the equivalence of the decoding space.
[0071] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the priority order of the embodiments.
[0072] Each embodiment of the present application provides a speech recognition decoding method, which can be executed by a terminal or a server, or jointly executed by a terminal and a server; in the embodiments of the present application, the case where the speech recognition decoding method is executed by the server is taken as an example for illustration.
[0073] Please refer to Figures 4 to 8 , Figure 4 which is a schematic flowchart of the speech recognition decoding method provided by the embodiments of the present application, Figures 5 to 8 and
[0074] Step 410, obtain the static decoding network that has been constructed.
[0075] For example, a decoding network without adding optional blanks as shown in Figure 1 can be used as the static decoding network.
[0076] Among them, the basic structure of the static decoding network is a directed graph, which is composed of nodes and directed arcs. A lexeme, as well as the acoustic model information and / or language model information of this lexeme, can be stored on the directed arcs. The acoustic model information generally appears as an acoustic model score, and the language model information generally appears as a language model score. Speech recognition is a process of finding an optimal path on this directed graph according to the input speech data. For example, the directed graph may include nodes and directed arcs connecting each node. Among the nodes, there is a start node and an end node. The start node represents the starting position when decoding the target speech data based on the directed graph in the static decoding network, while the end node represents the ending position when decoding the target speech data based on the directed graph in the static decoding network.
[0077] For example, in the directed graph, nodes can be used to represent phonemes, and the connection relationship between phonemes can be represented by the directed arcs connecting each node.
[0078] For another example, in the directed graph, directed arcs can also be used to represent phonemes, while nodes are a necessary element for expressing the network topology structure. The connection relationship is established between each directed arc through nodes. For example, each node can represent the connection relationship of words, that is, each node can represent the end time point of a word. Each directed arc can represent a possible word (phoneme), as well as information such as the acoustic score, language model score, and time point at which the word occurs.
[0079] Among them, a path in the directed graph includes each node and the directed arcs connecting each node. At this time, the words represented by each node and the directed arcs connecting each node can form a sentence.
[0080] Among them, a phoneme is the smallest speech unit divided according to the natural properties of speech. From an acoustic perspective, a phoneme is the smallest speech unit divided from the perspective of voice quality. From a physiological perspective, one pronunciation action forms one phoneme.
[0081] Step 420, during the process of decoding the target speech data, determine the activated nodes and activated arcs according to the nodes that have been traversed and the out-arcs of the nodes that have been traversed in the static decoding network, so as to obtain a dynamic decoding network including the activated nodes and the activated arcs.
[0082] For example, during the process of decoding the target speech data, feature extraction processing can be performed on the target speech data to be recognized to obtain a feature sequence of the target speech data. Then, search for the path with the highest score in the decoding network for the feature sequence of the target speech data as the decoding path, and use the word string corresponding to the decoding path as the speech recognition result corresponding to the target speech data.
[0083] For example, when performing feature extraction processing on the target speech data to be recognized to obtain the feature sequence of the target speech data, first, according to a certain sampling frequency (such as more than twice the highest frequency of the sound), the speech in the target speech data is converted from the physical state into an analog signal that is discrete in time and continuous in amplitude, and then a digital-form speech signal is formed through analog-to-digital (A / D) conversion.
[0084] For example, before performing feature extraction, preprocessing can also be performed on the digital-form speech signal. Such as pre-emphasis, windowing, framing, endpoint detection, and filtering and other preprocessing. Among them, pre-emphasis can be used to enhance the high-frequency part of the speech signal and make the spectrum of the speech signal smooth; windowing and framing can be used to divide the speech signal into multiple overlapping frames in the form of a rectangular window or a Hamming window according to the time-varying characteristics of the speech signal; endpoint detection can be used to find the starting part and the ending part of the speech signal, and filtering can be used to remove the background noise of the speech signal.
[0085] Then, feature extraction is performed on the preprocessed speech signal to extract the speech features of the target speech data, and a normalized feature sequence of the target speech data is formed according to the time series.
[0086] Among them, the speech features can include time-domain features and frequency-domain features in terms of manifestation forms, can include features based on the human voice generation mechanism in terms of sources, and also include features based on human ear auditory perception. In addition to the aforementioned static speech features, it can also include logarithmic energy or new features formed by splicing dynamic features calculated from static features by first-order and second-order differences.
[0087] For example, the obtained feature sequence of the target speech data is input into a decoder, and each frame of speech data in the target speech data is decoded based on the directed graph in the preset static decoding network in the decoder.
[0088] For example, during the decoding process, the constructed decoding network is called a static decoding network, the nodes that have been traversed on the static decoding network are called active nodes, the out-arcs that have been traversed on the static decoding network are called active arcs, and the topological structure composed of active nodes and active arcs is called a dynamic decoding network.
[0089] Among them, during the decoding process, the constructed static decoding network is traversed.
[0090] Step 430, when accessing the first active node corresponding to the current node on the static decoding network, control the first active node to perform one propagation of the blank arc, and the blank arc is used to mark the blank phoneme.
[0091] Optionally, when accessing the first activation node corresponding to the current node on the static decoding network, controlling the first activation node to perform one propagation of the blank arc includes:
[0092] When accessing the first activation node corresponding to the current node on the static decoding network, perform propagation according to the information carried by the real arcs of the current node;
[0093] After accessing all the real arcs of the current node, control the first activation node to perform one propagation of the blank arc, and record on the first activation node whether the first activation node propagates along the blank arc;
[0094] If there is a blank arc on the first activation node, control the first activation node to perform null arc propagation after the blank arc propagation.
[0095] For example, the currently accessed node on the static decoding network is referred to as the current node. When accessing the current node on the static decoding network, a first activation node having a mapping relationship with the current node can be generated in the dynamic decoding network, and the information carried on the current node is mapped to the first activation node.
[0096] For example, use the decoding network without adding optional blanks as shown in Figure 1 as the static decoding network. The current node is node 1, and the first activation node corresponding to the current node is activation node node1. When accessing the activation node node1 (the first activation node) corresponding to node 1 (the current node) on the static decoding network, the information propagation on the activation node node1 can be as shown in Figure 5 as follows:
[0097] (a) Perform propagation according to the information carried by the real arcs of node 1;
[0098] (b) After accessing all the real arcs state_x and state_y of node 1, control the activation node node1 to perform one propagation of the blank arc. At this time, record on the activation node node1 whether the activation node node1 propagates along the blank arc;
[0099] (c) If there is a blank arc on the activation node node1, control the activation node node1 to perform null arc propagation after the blank arc propagation.
[0100] Among them, if the activation node node1 does not propagate along the blank arc, a flag of False is returned. If there is a blank arc on the activation node node1 and the activation node node1 propagates along the blank arc, a flag of True is returned, and after the blank arc propagation, the empty arc propagation is performed.
[0101] For example, the flag can be understood as defining a variable used to determine whether the entire program is active. This variable is called a "flag" and acts as the traffic signal of the program. For example, the program can continue to run when the flag is True.
[0102] Among them, the current CTC needs to add an optional blank arc in the static decoding network. When accessing the activation node node1, two blank propagations are required. However, for the decoding method proposed in the embodiments of the present application, when accessing the activation node node1, as Figure 5 shown, there is no need to add an optional blank arc in the static decoding network. Only a blank arc needs to be added to the dynamic decoding network during the dynamic decoding process, effectively reducing the resources of the static decoding network. Only one blank propagation is required for each node, significantly reducing the redundant paths, and the reduction amplitude is proportional to the outgoing arcs of the node.
[0103] Step 440, when the blank arc is activated and propagated along the blank arc starting from the first activation node, a second activation node is newly created, where the first activation node and the second activation node map to the current node on the static decoding network.
[0104] Optionally, after the second activation node is newly created, it further includes:
[0105] Recording on the second activation node whether the second activation node is generated by the propagation of the blank arc.
[0106] For example, as Figure 6 shown, when the blank arc is activated and propagated along the blank arc starting from the activation node node1 (the first activation node), an activation node node4 (the second activation node) is newly created. Among them, the second activation node (such as the activation node node4) is a transition node, and the blank arc is the outgoing arc of the first activation node (such as the activation node node1).
[0107] For example, as Figure 6 shown for the activation node node4, the newly created activation node node4 still maps to node 1 of the static decoding network at this time.
[0108] As Figure 7As shown, the activation nodes node1 and node4 are both mapped to node 1 of the static decoding network, and it is recorded on the activation node node4 whether this activation node node4 is generated by the propagation of the blank arc.
[0109] Among them, activation can be understood as being accessed; activating the blank arc can be understood as the blank arc being accessed.
[0110] For example, as Figure 6 or Figure 7 shown, the activation node node2 in the dynamic decoding network is the mapped node corresponding to node 2 on the static decoding network; the activation node node3 in the dynamic decoding network is the mapped node corresponding to node 3 on the static decoding network.
[0111] Among them, since this activation node node4 is the mapped node of node 1, the output state of node 1 is thus on this activation node node4. For example, if the output states of node 1 are state_y and state_y, then the outgoing arcs with the output state of state_y and the outgoing arcs with the output state of state_y correspondingly exist on this activation node node4.
[0112] Step 450, propagate the second activation node according to the real arcs of the current node on the static decoding network that has a mapping relationship with the second activation node, and the real arcs are used to mark the output state of the current node.
[0113] Optionally, the propagating the second activation node according to the real arcs of the current node on the static decoding network that has a mapping relationship with the second activation node includes:
[0114] When propagating the second activation node, find the current node on the static decoding network that has a mapping relationship with the second activation node;
[0115] Traverse all the outgoing arcs of the current node on the static decoding network. If there are real arcs among the outgoing arcs of the current node, then propagate the information on the second activation node along the real arcs to the successor nodes of the current node.
[0116] Optionally, when traversing all the outgoing arcs of the current node on the static decoding network, it further includes:
[0117] If all the outgoing arcs of the current node are blank arcs, then do not propagate the second activation node.
[0118] For example, as Figure 7As shown, when propagating the activation node node4, the mapped node on the static decoding network corresponding to the activation node node4 is found. For example, it can be seen from step 440 that the mapped node of the activation node node4 is node1. Therefore, when propagating along the out-arcs of the activation node node4, all the out-arcs of node1 on the static decoding network are traversed. If an out-arc is a real arc (such as the out-arc with the output state state_y and the out-arc with the output state state_y), the information on the activation node node4 is propagated along this real arc (such as the out-arc with the output state state_y and the out-arc with the output state state_y). For example, the information on the activation node node4 is propagated along the out-arc with the output state state_x to the successor node (activation node node2) corresponding to the out-arc of state_x, and the information on the activation node node4 is propagated along the out-arc with the output state state_y to the successor node (activation node node3) corresponding to the out-arc of state_y. If it is a blank arc, no propagation is performed, that is, the activation node node4 only performs real-arc propagation.
[0119] For example, please refer to Figure 2 and Figure 7 , in Figure 2 , if there are two real arcs for node1, two blank arcs, two empty arcs, and two new nodes need to be added. In the embodiment of the present application, as Figure 7 shown, for the decoding process of node1 with two real arcs, only 1 blank arc and 1 new node need to be added during the dynamic decoding process. Compared with the decoding method shown in Figure 2 , the decoding network of the CTC model can be compressed on a large scale.
[0120] Step 460, traverse all the nodes on the static decoding network. If the termination node on the static decoding network is accessed, a blank arc self-jump is performed on the termination node.
[0121] Optionally, the step of performing a blank arc self-jump on the termination node if the termination node on the static decoding network is accessed includes:
[0122] If a node on the static decoding network without an out-arc is accessed, the node without an out-arc is determined as the termination node on the static decoding network;
[0123] After activating the blank arc connected to the termination node, the termination node is determined as the successor node of the blank arc connected to the termination node, so as to perform a blank arc self-jump on the termination node.
[0124] For example, traverse all the nodes on the static decoding network. When accessing each node, repeat the process from step 410 to step 450. When accessing a termination node (for example, a node without an outgoing arc on the static decoding network is a termination node), perform a self-loop on the blank arc at the termination node. That is, when the blank arc connected to the termination node is activated, the target node of this blank arc is the node itself. The equivalence principle is as Figure 8 shown. So far, the decoding network shown in this schematic diagram is completely equivalent to the decoding network obtained by adding an optional blank to the current CTC in terms of the decoding space; any path that can be obtained on the current CTC decoding network can be obtained on the Figure 8 schematic diagram shown.
[0125] As Figure 8 shown, if the currently accessed node is node 1, then on the dynamic decoding network, this node 1 is the first activation node corresponding to the current node 1, and the Figure 8 node 5 in it is equivalent to the second activation node. Node 2 is the successor node of node 1.
[0126] For example, if the currently accessed node is node 2, then on the dynamic decoding network, this node 2 is the first activation node corresponding to the current node 2, and the Figure 8 node 6 in it is equivalent to the second activation node. Node 3 is the successor node of node 2.
[0127] For example, if the currently accessed node is node 3, then on the dynamic decoding network, this node 3 is the first activation node corresponding to the current node 3, and the Figure 8 node 7 in it is equivalent to the second activation node. Node 4 is the successor node of node 3.
[0128] For example, when accessing node 4, this node 4 has no outgoing arc on the static decoding network. Therefore, this node 4 is a termination node. At this time, the target node of the blank arc added in the dynamic decoding network is this node 4 itself, which is equivalent to determining the termination node 4 as the successor node of the blank arc connected to it, thereby realizing the self-loop of the blank arc.
[0129] Optionally, the method further includes:
[0130] Until the dynamic decoding network contains the activation nodes corresponding to the termination nodes that have been traversed in the static decoding network, complete the decoding of the target voice data;
[0131] Determine the decoding path according to the activation nodes and activation arcs in the dynamic decoding network;
[0132] Generate a speech recognition result corresponding to the target speech data according to the decoding path.
[0133] For example, traverse all nodes on the static decoding network until an activation node corresponding to a termination node that has been traversed in the static decoding network is included in the dynamic decoding network, indicating that the decoding of the target speech data has been completed. Determine the decoding path based on the activation nodes and activation arcs in the dynamic decoding network. After obtaining the decoding path, the phonemes represented by the activation arcs between the activation nodes included in the decoding path can be determined. Then, according to the sequential combination between the activation nodes, a speech recognition result of the target speech data to be recognized is obtained.
[0134] For example, as Figure 8 shown, the directed graph in the dynamic decoding network can represent phonemes through arcs and represent the connection relationship of phonemes with nodes. The decoding path includes a node No. 1 (starting node), a node No. 5, a node No. 2, a node No. 6, a node No. 3, a node No. 7, and a node No. 4 (termination node). Among them, the arc between the node No. 1 and the node No. 5 represents a blank phoneme, and the arc between the node No. 5 and the node No. 2 represents the phoneme "a"; the arc between the node No. 2 and the node No. 6 represents a blank phoneme, and the arc between the node No. 6 and the node No. 3 represents the phoneme "b"; the arc between the node No. 3 and the node No. 7 represents a blank phoneme, and the arc between the node No. 7 and the node No. 4 represents the phoneme "c"; then, the speech recognition result corresponding to the target speech data generated according to this decoding path is "abc".
[0135] Among them, the embodiments of the present application make the decoding network construction process of the CTC model more concise, and the decoding network construction can be completed by following the traditional decoding network construction method. On the premise of no loss of speech recognition effect, the decoding network can be compressed by more than 40%, and the memory occupancy during the recognition process can be reduced by more than 20%.
[0136] Any combination of the above all technical solutions can form an optional embodiment of the present application, which will not be elaborated here one by one.
[0137] In an embodiment of the present application, a static decoding network that has been constructed is obtained; during the process of decoding target speech data, activation nodes and activation arcs are determined according to the nodes that have been traversed in the static decoding network and the outgoing arcs of the nodes that have been traversed, so as to obtain a dynamic decoding network including the activation nodes and the activation arcs; when accessing a first activation node corresponding to the current node on the static decoding network, the first activation node is controlled to perform one propagation of a blank arc, and the blank arc is used to mark a blank phoneme; when the blank arc is activated and propagated along the blank arc starting from the first activation node, a second activation node is newly created, where the first activation node and the second activation node map to the current node on the static decoding network; according to the real arc of the current node on the static decoding network that has a mapping relationship with the second activation node, the second activation node is propagated, and the real arc is used to mark the output state of the current node; all nodes on the static decoding network are traversed, and if the termination node on the static decoding network is accessed, a self-jump of the blank arc is performed on the termination node. In the embodiment of the present application, by not adding optional blank arcs in the static decoding network and only adding blank arcs on the dynamic decoding network during the dynamic decoding process, the resources of the static decoding network are effectively reduced, the decoding network can be effectively compressed, and only one blank propagation is required for each node, greatly reducing the redundancy of the decoding path.
[0138] To facilitate better implementation of the speech recognition decoding method in the embodiment of the present application, the embodiment of the present application also provides a speech recognition decoding device. Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of the speech recognition decoding device provided in the embodiment of the present application. Among them, the speech recognition decoding device 900 may include:
[0139] An acquisition unit 901, configured to acquire a static decoding network that has been constructed;
[0140] A determination unit 902, configured to determine activation nodes and activation arcs according to the nodes that have been traversed in the static decoding network and the outgoing arcs of the nodes that have been traversed during the process of decoding target speech data, so as to obtain a dynamic decoding network including the activation nodes and the activation arcs;
[0141] A first propagation unit 903, configured to control the first activation node to perform one propagation of a blank arc when accessing the first activation node corresponding to the current node on the static decoding network, and the blank arc is used to mark a blank phoneme;
[0142] A new unit 904 is used to create a second activation node when propagating along the blank arc from the first activation node after activating the blank arc, where the first activation node and the second activation node map to the current node on the static decoding network;
[0143] A second propagation unit 905 is used to propagate the second activation node according to the real arc of the current node on the static decoding network that has a mapping relationship with the second activation node, and the real arc is used to mark the output state of the current node;
[0144] A processing unit 906 is used to traverse all nodes on the static decoding network. If the termination node on the static decoding network is accessed, a blank arc self-jump is performed on the termination node.
[0145] Optionally, the first propagation unit 903 can be used to: when accessing the first activation node corresponding to the current node on the static decoding network, propagate according to the information carried by the real arc of the current node; after accessing all real arcs of the current node, control the first activation node to perform a propagation of a blank arc, and record on the first activation node whether the first activation node propagates along the blank arc; if there is a blank arc on the first activation node, control the first activation node to perform a null arc propagation after the blank arc propagation.
[0146] Optionally, the second propagation unit 905 can be used to: when propagating the second activation node, find the current node on the static decoding network that has a mapping relationship with the second activation node; traverse all out-arcs of the current node on the static decoding network. If there is a real arc among the out-arcs of the current node, propagate the information on the second activation node along the real arc to the successor node of the current node.
[0147] Optionally, when the second propagation unit 905 traverses all out-arcs of the current node on the static decoding network, it can also be used to: if all out-arcs of the current node are null arcs, do not propagate the second activation node.
[0148] Optionally, the processing unit 906 can be used to: if a node without an out-arc on the static decoding network is accessed, determine the node without an out-arc as the termination node on the static decoding network; when activating the blank arc connected to the termination node, determine the termination node as the successor node of the blank arc connected to the termination node, so as to achieve a blank arc self-jump on the termination node.
[0149] Optionally, after creating the second activation node, the new unit 904 may also be used to record on the second activation node whether the second activation node is generated by the propagation of the blank arc.
[0150] Optionally, the processing unit 906 may also be used to: complete the decoding of the target speech data until the activation node corresponding to the termination node that has been traversed in the static decoding network is included in the dynamic decoding network; determine the decoding path according to the activation nodes and activation arcs in the dynamic decoding network; and generate a speech recognition result corresponding to the target speech data according to the decoding path.
[0151] It should be noted that the functions of the various modules in the speech recognition decoding device 900 in the embodiments of the present application may be correspondingly referred to the specific implementation manners of any of the above method embodiments, and will not be elaborated here.
[0152] Each unit in the above speech recognition decoding device may be implemented in whole or in part by software, hardware, and their combination. The above units may be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above units.
[0153] The speech recognition decoding device 900 may be integrated, for example, in a terminal or server with a memory and a processor and having computing capabilities, or the speech recognition decoding device 900 is the terminal or server. The terminal may be a smart phone, a tablet computer, a laptop computer, a smart TV, a smart speaker, a wearable smart device, a personal computer (PC), etc. The terminal may also include a client, and the client may be a video client, a browser client, or an instant messaging client, etc. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0154] Figure 10 It is a schematic structural diagram of the computer device provided by the embodiments of the present application, as Figure 10As shown in the figure, the computer device 1000 may include: a communication interface 1001, a memory 1002, a processor 1003, and a communication bus 1004. The communication interface 1001, the memory 1002, and the processor 1003 communicate with each other through the communication bus 1004. The communication interface 1001 is used for the device 1000 to communicate with external devices for data. The memory 1002 can be used to store software programs and modules. The processor 1003 runs the software programs and modules stored in the memory 1002, such as the software programs for the corresponding operations in the foregoing method embodiments.
[0155] Optionally, the processor 1003 may call the software programs and modules stored in the memory 1002 to perform the following operations: obtain the already constructed static decoding network; during the process of decoding the target voice data, determine the activation nodes and activation arcs according to the nodes that have been traversed and the outgoing arcs of the nodes that have been traversed in the static decoding network, so as to obtain a dynamic decoding network including the activation nodes and the activation arcs; when accessing the first activation node corresponding to the current node on the static decoding network, control the first activation node to perform one propagation of the blank arc, and the blank arc is used to mark the blank phoneme; when the blank arc is activated and propagated along the blank arc starting from the first activation node, create a second activation node, where the first activation node and the second activation node map the current node on the static decoding network; according to the real arc of the current node on the static decoding network that has a mapping relationship with the second activation node, propagate the second activation node, and the real arc is used to mark the output state of the current node; traverse all the nodes on the static decoding network, and if the termination node on the static decoding network is accessed, perform a self-jump of the blank arc on the termination node.
[0156] Optionally, the computer device 1000 is the terminal or the server. The terminal may be a smart phone, a tablet computer, a laptop computer, a smart TV, a smart speaker, a wearable smart device, a personal computer, or the like. The server may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0157] Optionally, the present application further provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the foregoing method embodiments are implemented.
[0158] The present application also provides a computer-readable storage medium for storing a computer program. The computer-readable storage medium can be applied to a computer device, and the computer program enables the computer device to execute the corresponding processes in the speech recognition decoding method in the embodiments of the present application. For the sake of brevity, details are not described herein again.
[0159] The present application also provides a computer program product, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute the corresponding processes in the speech recognition decoding method in the embodiments of the present application. For the sake of brevity, details are not described herein again.
[0160] The present application also provides a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute the corresponding processes in the speech recognition decoding method in the embodiments of the present application. For the sake of brevity, details are not described herein again.
[0161] It should be understood that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or by the instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware decoding processor, or executed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0162] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.
[0163] It should be understood that the above memory is by way of example but not limitation. For example, the memory in the embodiments of the present application can also be a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM), etc. That is to say, the memory in the embodiments of the present application is intended to include but not be limited to these and any other suitable types of memory.
[0164] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0165] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0166] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0167] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0168] In addition, the functional units in the embodiments of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0169] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer or a server) to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0170] As described above, it is only the specific implementation manner of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A speech recognition decoding method, characterized in that, The method includes: Obtaining a static decoding network that has been constructed; During the process of decoding target speech data, determining activation nodes and activation arcs based on the nodes that have been traversed and the out-arcs of the nodes that have been traversed in the static decoding network, so as to obtain a dynamic decoding network including the activation nodes and the activation arcs; When accessing a first activation node corresponding to a current node on the static decoding network, controlling the first activation node to perform one propagation of a blank arc, where the blank arc is used to mark a blank phoneme; After activating the blank arc, when propagating along the blank arc starting from the first activation node, creating a second activation node, where the first activation node and the second activation node map to the current node on the static decoding network; Propagating the second activation node according to the real arc of the current node on the static decoding network that has a mapping relationship with the second activation node, where the real arc is used to mark the output state of the current node; Traversing all nodes on the static decoding network, if the termination node on the static decoding network is accessed, then perform a self-jump of the blank arc on the termination node.
2. The voice recognition decoding method according to claim 1, wherein The step of, when accessing a first activation node corresponding to a current node on the static decoding network, controlling the first activation node to perform one propagation of a blank arc, includes: When accessing a first activation node corresponding to a current node on the static decoding network, performing propagation according to the information carried by the real arc of the current node; After accessing all real arcs of the current node, controlling the first activation node to perform one propagation of a blank arc, and recording on the first activation node whether the first activation node propagates along the blank arc; If there is a blank arc on the first activation node, then controlling the first activation node to perform empty arc propagation after the propagation of the blank arc.
3. The voice recognition decoding method according to claim 1, characterized in that, The step of propagating the second activation node according to the real arc of the current node on the static decoding network that has a mapping relationship with the second activation node, includes: When propagating the second activation node, searching for the current node on the static decoding network that has a mapping relationship with the second activation node; Traversing all out-arcs of the current node on the static decoding network, if there is a real arc among the out-arcs of the current node, then propagating the information on the second activation node along the real arc to the successor node of the current node.
4. The voice recognition decoding method according to claim 3, wherein When traversing all out-arcs of the current node on the static decoding network, it further includes: If all out-arcs of the current node are empty arcs, then not propagating the second activation node.
5. The voice recognition decoding method according to claim 3, wherein, The step of, if the termination node on the static decoding network is accessed, then performing a self-jump of the blank arc on the termination node, includes: If a node without an out-arc on the static decoding network is accessed, then determining the node without an out-arc as the termination node on the static decoding network; After activating the blank arc connected to the termination node, determine the termination node as the successor node of the blank arc connected to the termination node, so as to achieve a self-jump of the blank arc on the termination node.
6. The voice recognition decoding method according to claim 1, characterized in that, After the new second activation node is created, it further includes: Record on the second activation node whether the second activation node is generated by the propagation of the blank arc.
7. The speech recognition decoding method according to any one of claims 1-6, characterized in that, The method further includes: Decode the target speech data until the activation node corresponding to the termination node that has been traversed in the static decoding network is included in the dynamic decoding network; Determine the decoding path according to the activation nodes and activation arcs in the dynamic decoding network; Generate the speech recognition result corresponding to the target speech data according to the decoding path.
8. A voice recognition decoding device, characterized in that, The device includes: An acquisition unit, configured to acquire a static decoding network that has been constructed; A determination unit, configured to determine activation nodes and activation arcs according to the nodes that have been traversed in the static decoding network and the outgoing arcs of the nodes that have been traversed during the decoding of the target speech data, so as to obtain a dynamic decoding network including the activation nodes and the activation arcs; A first propagation unit, configured to control the first activation node to perform one propagation of the blank arc when accessing the first activation node corresponding to the current node on the static decoding network, where the blank arc is used to mark a blank phoneme; A new creation unit, configured to create a second activation node when propagating along the blank arc starting from the first activation node after activating the blank arc, where the first activation node and the second activation node map to the current node on the static decoding network; A second propagation unit, configured to propagate the second activation node according to the real arc of the current node on the static decoding network that has a mapping relationship with the second activation node, where the real arc is used to mark the output state of the current node; A processing unit, configured to traverse all nodes on the static decoding network, and if the termination node on the static decoding network is accessed, perform a self-jump of the blank arc on the termination node.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the steps in the speech recognition decoding method according to any one of claims 1-7.
10. A computer device, characterized in that, The computer device includes a processor and a memory, the memory stores a computer program, and the processor is configured to execute the steps in the speech recognition decoding method according to any one of claims 1-7 by calling the computer program stored in the memory.
Citation Information
Patent Citations
Speech recognition method and speech recognition device
CN105513589A
Voice recognition method and system based on deep neural network acoustic model
CN112927682A