Standard term speech recognition system and method based on double AC automaton
By using dual AC automata technology in the speech recognition system to process word elements and standardized terms respectively, the problem of difficulty in detecting speech recognition standardized terms in the prior art is solved, and effective detection of standardized terms and normalized processing of similar words is realized.
Patent Information
- Application Number
- CN202510261278.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-06
AI Technical Summary
The existing AC automata is difficult to use for the standardized term detection of speech recognition, especially when dealing with invalid words such as modal words and similar words, it is difficult to construct an effective pattern string.
The standardized speech recognition system based on dual AC automata is adopted, and the sound is collected through the microphone module, and the speech recognition module performs speech feature extraction and text sequence generation. Combined with the word element extraction AC automata and the standardized word AC automata, the matching and detection of word element and standardized word respectively.
Effective detection of standardized terms for speech recognition is achieved, the influence of modal words is avoided, the robustness of the method is enhanced, and the normalized detection of similar words can be effectively handled.
Smart Images

Figure CN120148482A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition, and particularly to a canonical term speech recognition system and method based on a double AC automaton. Background Art
[0002] The AC automaton is an algorithm for string search. The background of this algorithm is to solve the problem of finding multiple pattern strings (or called "keywords") in a main text string. It can complete the search in linear time and is very efficient.
[0003] The existing AC automaton can efficiently solve the multi-pattern string matching problem of strings, can complete the search in linear time, is particularly suitable for processing the matching problem of a large number of keywords, can also match multiple pattern strings at the same time, without separately matching each pattern, greatly improving the efficiency, and is applicable to various scenarios requiring multi-pattern matching, such as DNA sequence analysis in bioinformatics, keyword filtering in web crawlers, etc. However, it is difficult to be used for the detection of canonical terms in speech recognition. The canonical terms recognized by speech contain invalid words such as filler words, which affect the pattern string matching of the AC automaton. At the same time, the similar words in the canonical terms recognized by speech make it difficult to construct pattern strings. Therefore, it is very necessary to propose a canonical term speech recognition system and method based on a double AC automaton. Summary of the Invention
[0004] The purpose of the present invention is to provide a canonical term speech recognition system and method based on a double AC automaton, while enhancing the robustness of the method, avoiding invalid words such as filler words, and realizing the normalization detection of similar words, so as to solve the existing technical defects and unachievable technical requirements.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A canonical term speech recognition system based on a double AC automaton, comprising:
[0006] A microphone module, configured to collect the sound of the device and send the sound data to the speech recognition module;
[0007] A speech recognition module, configured to receive the sound data sent by the microphone module, extract speech features from the sound data, generate a speech recognition text sequence according to the speech features, and then send the speech recognition text sequence to the token extraction AC automaton;
[0008] A token extraction AC automaton, configured to receive the speech recognition text sequence sent by the speech recognition module and the token list sent by the canonical term decomposition module, match the speech recognition text sequence with the token list, obtain a complete token, generate a corresponding token number sequence, and then send the token number sequence to the canonical term AC automaton;
[0009] The canonical term AC automaton is used to receive the sequence of token numbers sent by the token extraction AC automaton and the list of canonical terms sent by the canonical term decomposition module, match the sequence of token numbers with the list of canonical terms, match a complete canonical term, generate the corresponding canonical term sequence, and then generate the compliance detection module (5) from the canonical term sequence;
[0010] The compliance detection module (5) is used to receive the canonical term sequence sent by the canonical term AC automaton and the compliance list sent by the canonical term decomposition module, and detect the canonical term sequence through the compliance list;
[0011] The canonical term decomposition module generates a list of tokens, a list of canonical terms, and a compliance list according to the canonical terms respectively, and sends the three to the token extraction AC automaton, the canonical term AC automaton, and the compliance detection module (5) respectively.
[0012] A canonical term speech recognition method based on a dual AC automaton includes:
[0013] 1), The microphone module collects the sound of the device, converts the sound into PCM data, and then sends the PCM data to the speech recognition module;
[0014] 2), The speech recognition module generates a speech recognition text sequence according to the sound data
[0015] 2.1), The speech recognition module receives the PCM data sent by the microphone module;
[0016] 2.2), The speech recognition module extracts the speech features of the PCM data;
[0017] 2.3), The speech recognition module generates a speech recognition text sequence through the language model according to the speech features;
[0018] 2.4), The speech recognition module sends the speech recognition text sequence to the token extraction AC automaton;
[0019] 3), The token extraction AC automaton generates a sequence of token numbers according to the speech recognition text sequence and the list of tokens
[0020] 3.1), The token extraction AC automaton receives the speech recognition text sequence sent by the speech recognition module and the list of tokens sent by the canonical term decomposition module;
[0021] 3.2), The token extraction AC automaton generates a corresponding token dictionary tree according to the list of tokens, and the token extraction AC automaton constructs a first failure pointer for the token dictionary tree;
[0022] 3.3), The token extraction AC automaton matches a complete token in the token dictionary tree for the speech recognition text sequence and generates a corresponding token number sequence;
[0023] 3.4), The token extraction AC automaton sends the token number sequence to the standard term AC automaton;
[0024] 4), The standard term AC automaton generates a standard term sequence according to the token number sequence and the standard term list
[0025] 4.1), The standard term AC automaton receives the token number sequence sent by the token extraction AC automaton and the standard term list sent by the standard term decomposition module;
[0026] 4.2), The standard term AC automaton generates a corresponding standard term dictionary tree according to the standard term list and constructs a second failure pointer for the standard term dictionary tree;
[0027] 4.3), The standard term AC automaton matches a complete standard term in the standard term dictionary tree for the token number sequence and generates a corresponding standard term sequence;
[0028] 4.4), The standard term AC automaton generates a compliance detection module (5) for the standard term sequence
[0029] 5), The compliance detection module (5) receives the standard term sequence sent by the standard term AC automaton and the compliance list sent by the standard term decomposition module, and detects the compliance term sequence through the compliance list to check for omissions and incorrect order;
[0030] 6), The standard term decomposition module provides a corresponding list generated according to the standard term for each module
[0031] 6.1), The standard term decomposition module generates a token list according to the standard term and sends the token list to the token extraction AC automaton;
[0032] 6.2), The standard term decomposition module generates a standard term list according to the standard term and sends the standard term list to the standard term AC automaton;
[0033] 6.3), The standard term decomposition module generates a compliance list according to the standard term and sends the compliance list to the compliance detection module (5).
[0034] Preferably, the method for generating the speech recognition text sequence in step 2.3) is:
[0035] The speech recognition module maps the speech features to the phonemes of the acoustic model and uses the known word sequence of the language model to predict the probability of the next word, and finally outputs the most likely speech recognition text sequence.
[0036] Preferably, in the step 3.2), the token list includes tokens and token numbers. A token is composed of one or more characters. Multiple tokens can share the same token number. Each character of the token serves as a node of the token trie. The path from the root node of the token trie to a leaf node on the token trie represents a token. The tree nodes of the token trie contain the next-level tree nodes stored using an AVL tree, and the failure pointer corresponds to the longest matching suffix of the tree node of the token trie.
[0037] In this application, the above content can be understood as the failure pointer of the AC automaton pointing to the longest suffix state of the current state. Note: When the AC automaton performs matching, multiple pattern strings can be matched at the same position.
[0038] Preferably, in the step 3.3), the method for generating the token number sequence is as follows:
[0039] The token extraction AC automaton matches the speech recognition text sequence in the token trie. During the matching process, if a mismatch is found, the already matched partial prefix is discarded, and the failure pointer is used to jump to the tree node of the longest matching state. When the token extraction AC automaton matches a complete token, the token extraction AC automaton generates a token number sequence corresponding to the token.
[0040] Preferably, in the step 4.2), the standard term list includes standard terms and standard term numbers. The standard terms are composed of multiple token numbers. Multiple standard terms can share the same standard term number. And each token number of the standard term serves as a node of the standard term trie. The path from the root node of the standard term trie to a leaf node on the standard term trie represents a standard term, and the failure pointer corresponds to the longest matching suffix of the tree node of the token trie.
[0041] Preferably, in the step 4.3), the method for generating the standard term sequence is as follows:
[0042] The standard term AC automaton matches the token number sequence in the standard term trie. During the matching process, if a mismatch is found, the already matched partial prefix is discarded, and the failure pointer is used to jump to the tree node of the longest matching state. When the standard term AC automaton matches a complete standard term, the standard term AC automaton generates a standard term sequence corresponding to the standard term.
[0043] Preferably, in the step 5), the compliance list includes a standard term number detection sequence. During detection, the compliance detection module (5) matches the standard term sequence with the standard term number detection sequence to detect whether there is any omission and whether the order is incorrect.
[0044] Preferably, in step 6.1), the specification term decomposition module generates a list of word elements from the key words of the specification terms.
[0045] Preferably, the specification term decomposition module encodes and numbers the specification terms to generate a list of specification terms, and the specification term decomposition module generates a compliance list from the specification term numbers.
[0046] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0047] By adopting a double AC automaton, the present invention respectively realizes detection acceleration for word elements and specification terms, enhances the robustness of the method, avoids invalid words such as modal particles, and realizes normalized detection of similar words. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a schematic diagram of the overall logic of the present invention;
[0049] In the figure: microphone module 1, speech recognition module 2, word element extraction AC automaton 3, specification term AC automaton 4, compliance detection module 5, specification term decomposition module 6. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. Figure 1
[0051] Figure 1 Please refer to Figure 1 , embodiments of the present invention:
[0052] Embodiment:
[0053] As Figure 1 shown: A specification term speech recognition system based on a double AC automaton includes:
[0054] A microphone module 1, configured to collect the sound of the device and send the sound data to the speech recognition module 2;
[0055] A speech recognition module 2, configured to receive the sound data sent by the microphone module 1, extract speech features from the sound data, generate a speech recognition text sequence according to the speech features, and then send the speech recognition text sequence to the word element extraction AC automaton 3;
[0056] The token extraction AC automaton 3 is used to receive the speech recognition text sequence sent by the speech recognition module 2 and the token list sent by the specification term decomposition module 6, match the speech recognition text sequence with the token list, obtain a complete token, generate the corresponding token number sequence, and then send the token number sequence to the specification term AC automaton 4;
[0057] The specification term AC automaton 4 is used to receive the token number sequence sent by the token extraction AC automaton 3 and the specification term list sent by the specification term decomposition module 6, match the token number sequence with the specification term list, obtain a complete specification term, generate the corresponding specification term sequence, and then send the specification term sequence to the compliance detection module 5;
[0058] The compliance detection module 5 is used to receive the specification term sequence sent by the specification term AC automaton 4 and the compliance list sent by the specification term decomposition module 6, and detect the specification term sequence through the compliance list;
[0059] The specification term decomposition module 6 generates a token list, a specification term list and a compliance list according to the specification terms respectively, and sends the three to the token extraction AC automaton 3, the specification term AC automaton 4 and the compliance detection module 5 respectively.
[0060] A specification term speech recognition method based on a double AC automaton includes:
[0061] 1), The microphone module 1 collects the sound of the device, converts the sound into PCM data, and then sends the PCM data to the speech recognition module 2;
[0062] 2), The speech recognition module 2 generates a speech recognition text sequence according to the sound data
[0063] 2.1), The speech recognition module 2 receives the PCM data sent by the microphone module 1;
[0064] 2.2), The speech recognition module 2 extracts the speech features from the PCM data;
[0065] 2.3), The speech recognition module 2 maps the speech features to the phonemes of the acoustic model, and uses the known word sequence of the language model to predict the probability of the next word to constrain the output of the acoustic model, and finally outputs the most likely speech recognition text sequence.
[0066] 2.4), The speech recognition module 2 sends the speech recognition text sequence to the token extraction AC automaton 3;
[0067] 3), The token extraction AC automaton 3 generates a token number sequence according to the speech recognition text sequence and the token list
[0068] 3.1), The token extraction AC automaton 3 receives the speech recognition text sequence sent by the speech recognition module 2 and the token list sent by the standard term decomposition module 6;
[0069] 3.2), The token extraction AC automaton 3 generates a corresponding token dictionary tree according to the token list, and the token extraction AC automaton 3 constructs a first failure pointer for the token dictionary tree;
[0070] In the step 3.2), the token list contains tokens and token numbers. A token is composed of one or more characters. Multiple tokens can share a token number. (Multiple tokens share a token number. For example, a fuel card and a fuel additive share a token number, and the fuel card and the fuel additive are treated as the same token for unified processing in the token dictionary tree). Each character of the token is used as a node of the token dictionary tree. The path from the root node of the token dictionary tree to a leaf node on the token dictionary tree represents a token. The tree nodes of the token dictionary tree store the next-level tree nodes using an AVL tree to improve the access speed when there are a large number of identical prefixes among tokens. The failure pointer corresponds to the longest matching suffix of the tree node of the token dictionary tree.
[0071] 3.3), The token extraction AC automaton 3 matches a complete token in the token dictionary tree for the speech recognition text sequence and generates a corresponding token number sequence;
[0072] In the step 3.3), the method for generating the token number sequence is as follows:
[0073] The token extraction AC automaton 3 matches the speech recognition text sequence in the token dictionary tree. If a mismatch is found during the matching process, the already matched partial prefix is discarded, and the failure pointer is used to jump to the tree node in the longest matching state, improving the token matching efficiency. When the token extraction AC automaton 3 matches a complete token, the token extraction AC automaton 3 generates a token number sequence for the token.
[0074] 3.4), The token extraction AC automaton 3 sends the token number sequence to the standard term AC automaton 4;
[0075] In this embodiment, the content of step 3) can process short tokens by using an AC automaton, avoiding the influence of invalid words such as filler words, and having high performance at the same time.
[0076] 4), The standard term AC automaton 4 generates a standard term sequence according to the token number sequence and the standard term list
[0077] 4.1), The standard term AC automaton 4 receives the token number sequence sent by the token extraction AC automaton 3 and the standard term list sent by the standard term decomposition module 6;
[0078] 4.2) The AC automaton 4 for standard terms generates a corresponding trie of standard terms based on the standard term list, and constructs a second failure pointer for the trie of standard terms;
[0079] In the step 4.2), the standard term list contains standard terms and standard term numbers. The standard terms are composed of multiple token numbers. Multiple standard terms can share a standard term number. The standard terms are composed of multiple token numbers. Multiple standard terms can share a standard term number, which improves the robustness of rule term detection, avoids the influence of invalid words such as modal particles, and at the same time avoids the influence of similar words on the construction of standard terms. Multiple standard terms can share the standard term number, so that one step of compliance detection supports different sentence patterns. And each token number of the standard term is used as a node of the trie of standard terms. The path from the root node to a leaf node in the trie of standard terms represents a standard term, and the failure pointer corresponds to the longest matching suffix of the tree node of the token trie.
[0080] 4.3) The AC automaton 4 for standard terms matches the token number sequence to a complete standard term in the trie of standard terms and generates a corresponding standard term sequence;
[0081] In the step 4.3), the method for generating the standard term sequence is as follows:
[0082] The AC automaton 4 for standard terms matches the token number sequence in the trie of standard terms. During the matching process, if a mismatch is found, the already matched partial prefix is discarded, and the failure pointer is used to jump to the tree node in the longest matching state, which improves the matching efficiency of standard terms. When the AC automaton 4 for standard terms matches a complete standard term, the AC automaton 4 for standard terms generates a standard term sequence with the standard term number corresponding to the standard term.
[0083] 4.4) The AC automaton 4 for standard terms generates a compliance detection module 5
[0084] 5) The compliance detection module 5 receives the standard term sequence sent by the AC automaton 4 for standard terms and the compliance list sent by the standard term decomposition module 6, and detects the compliance term sequence through the compliance list to check whether there are omissions and incorrect orders;
[0085] In the step 5), the compliance list contains a standard term number detection sequence. During detection, the compliance detection module 5 matches the standard term sequence with the standard term number detection sequence to detect whether there are omissions and incorrect orders, so as to realize the detection of standard terms.
[0086] 6) The standard term decomposition module 6 provides a corresponding list generated according to the standard terms for each module
[0087] 6.1), the specification term decomposition module 6 generates a list of word elements from the keyword terms of the specification terms and sends the list of word elements to the word element extraction AC automaton 3;
[0088] 6.2), the specification term decomposition module 6 encodes and numbers the specification terms to generate a list of specification terms and sends the list of specification terms to the specification term AC automaton 4;
[0089] 6.3), the specification term decomposition module 6 generates a compliance list from the specification term numbers and sends the compliance list to the compliance detection module 5.
[0090] The above has shown and described the basic principles, main features and advantages of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic features of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claimed invention.
[0091] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that those skilled in the art can understand.
Claims
1. A standard language speech recognition system based on dual AC automata, characterized in that: include: A microphone module (1) is used to collect the sound of the device and send the sound data to the speech recognition module (2); The speech recognition module (2) is used to receive the sound data sent by the microphone module (1), extract the speech features of the sound data, generate a speech recognition text sequence according to the speech features, and then send the speech recognition text sequence to the word unit extraction AC automaton (3); The word unit extraction AC automaton (3) is used to receive the speech recognition text sequence sent by the speech recognition module (2) and the word unit list sent by the standard term decomposition module (6), match the speech recognition text sequence with the word unit list, obtain a complete word unit through matching, generate a corresponding word unit number sequence, and then send the word unit number sequence to the standard term AC automaton (4); The standard terminology AC automaton (4) is used to receive the word unit number sequence sent by the word unit extraction AC automaton (3) and the standard terminology list sent by the standard terminology decomposition module (6), and match the word unit number sequence with the standard terminology list to match a complete standard terminology, and generate a corresponding standard terminology sequence, and then generate the standard terminology sequence into a compliance detection module (5); A compliance detection module (5) is used to receive the standard term sequence sent by the standard term AC automaton (4) and the compliance list sent by the standard term decomposition module (6), and detect the standard term sequence through the compliance list; The standardized terminology decomposition module (6) generates a word-unit list, a standardized terminology list and a compliance list according to the standardized terminology, and sends the three to the word-unit extraction AC automaton (3), the standardized terminology AC automaton (4) and the compliance detection module (5) respectively.
2. A method for speech recognition of standard terms based on dual AC automata, characterized by comprising: 1) The microphone module (1) collects the sound of the device and converts the sound into PCM data, and then sends the PCM data to the speech recognition module (2); 2) Speech recognition module (2) generates speech recognition text sequence based on sound data 2.1) The speech recognition module (2) receives the PCM data sent by the microphone module (1); 2.2) The speech recognition module (2) extracts speech features from the PCM data; 2.3) The speech recognition module (2) generates a speech recognition text sequence according to speech features through a language model; 2.4) The speech recognition module (2) sends the speech recognition text sequence to the word unit extraction AC automaton (3); 3) Word unit extraction AC automaton (3) generates a word unit number sequence based on the speech recognition text sequence and word unit list 3.1) The word unit extraction AC automaton (3) receives the speech recognition text sequence sent by the speech recognition module (2) and the word unit list sent by the standard term decomposition module (6); 3.2) The word-unit extraction AC automaton (3) generates a corresponding word-unit dictionary tree according to the word-unit list, and the word-unit extraction AC automaton (3) constructs an invalidation pointer for the word-unit dictionary tree; 3.3) The word unit extraction AC automaton (3) matches the speech recognition text sequence to a complete word unit in the word unit dictionary tree and generates a corresponding word unit number sequence; 3.4) The word-unit extraction AC automaton (3) sends the word-unit number sequence to the standard terminology AC automaton (4); 4) Standardized terminology AC automaton (4) generates a standardized terminology sequence based on the word number sequence and the standardized terminology list 4.1) The standard terminology AC automaton (4) receives the word unit number sequence sent by the word unit extraction AC automaton (3) and the standard terminology list sent by the standard terminology decomposition module (6); 4.2) The standard terminology AC automaton (4) generates a corresponding standard terminology dictionary tree according to the standard terminology list, and constructs an invalidation pointer for the standard terminology dictionary tree; 4.3) The standard terminology AC automaton (4) matches the word unit number sequence to a complete standard term in the standard terminology dictionary tree and generates a corresponding standard term sequence; 4.4) The standard terminology AC automaton (4) generates a compliance detection module (5) from the standard terminology sequence 5) The compliance detection module (5) receives the sequence of standard terms sent by the standard term AC automaton (4) and the compliance list sent by the standard term decomposition module (6), and detects the sequence of standard terms through the compliance list to see if there are any omissions or sequence interruptions; 6) Standard term decomposition module (6) provides each module with a corresponding list generated based on standard terms. 6.1) The standard term decomposition module (6) generates a word list according to the standard term, and sends the word list to the word extraction AC automaton (3); 6.2) The standard term decomposition module (6) generates a standard term list according to the standard term, and sends the standard term list to the standard term AC automaton (4); 6.3) The standard term decomposition module (6) generates a compliance list based on the standard terminology and sends the compliance list to the compliance detection module (5).
3. A method for speech recognition of standard terms based on dual AC automata according to claim 2, characterized in that: The method for generating the speech recognition text sequence in step 2.3) is: The speech recognition module (2) maps speech features to the phonemes of the acoustic model and uses the known word sequence of the language model to predict the probability of the next word, and finally outputs the most likely speech recognition text sequence.
4. The method for standard language speech recognition based on dual AC automata according to claim 2, characterized in that: In the step 3.2), the word-gram list includes word-grams and word-gram numbers. A word-gram is composed of one or more characters. Multiple word-grams can share a word-gram number. Each character of the word-gram is used as a node of the word-gram dictionary tree. The path from the root node of the word-gram dictionary tree to a leaf node on the word-gram dictionary tree represents a word-gram. The tree node of the word-gram dictionary tree includes the next level tree node stored using a balanced binary tree, and the invalidation pointer corresponds to the longest matching suffix of the tree node of the word-gram dictionary tree.
5. A method for speech recognition of standard terms based on dual AC automata according to claim 4, characterized in that: In step 3.3), the method of generating the word unit number sequence is: The word-unit extraction AC automaton (3) matches the speech recognition text sequence in the word-unit dictionary tree. If a mismatch is found during the matching process, the matched partial prefix is discarded and the invalid pointer is used to jump to the tree node with the longest matching state. When the word-unit extraction AC automaton (3) matches a complete word-unit, the word-unit extraction AC automaton (3) generates a word-unit number sequence by using the word-unit number corresponding to the word-unit.
6. A method for speech recognition of standard terms based on dual AC automata according to claim 2, characterized in that: In the step 4.2), the list of standard terms includes standard terms and standard term numbers. The standard terms are composed of multiple word unit numbers. Multiple standard terms can share one standard term number, and each word unit number of the standard term is used as a node of the standard term dictionary tree. The path from the root node on the standard term dictionary tree to a leaf node on the standard term dictionary tree represents a standard term, and the invalidation pointer corresponds to the longest matching suffix of the tree node of the word unit dictionary tree.
7. A method for speech recognition of standard terms based on dual AC automata according to claim 6, characterized in that: In step 4.3), the method of generating the standard term sequence is as follows: The standard terminology AC automaton (4) matches the word unit number sequence in the standard terminology dictionary tree. If a mismatch is found during the matching process, the matched partial prefix is discarded and the invalid pointer is used to jump to the tree node with the longest matching state. When the standard terminology AC automaton (4) matches a complete standard terminology, the standard terminology AC automaton (4) generates a standard terminology sequence with the standard terminology number corresponding to the standard terminology.
8. The method for standard language speech recognition based on dual AC automata according to claim 2, characterized in that: In the step 5), the compliance list includes a standard terminology number detection sequence. During detection, the compliance detection module (5) matches the standard terminology sequence with the standard terminology number detection sequence to detect whether there are omissions and whether the sequence is continuous.
9. The method for standard language speech recognition based on dual AC automata according to claim 2, characterized in that: In the step 6.1), the standard term decomposition module (6) generates a word-unit list from the key words of the standard term.
10. The method for standard language speech recognition based on dual AC automata according to claim 2, characterized in that: The standard term decomposition module (6) codes and numbers the standard terminology to generate a standard term list, and the standard term decomposition module (6) codes and numbers the standard terminology to generate a compliance list.
Citation Information
Patent Citations
Operation method for Chinese AC automatic machine based on retrieval of keyword dictionary tree
CN105183788A
Multi-mode string matching method and apparatus
CN106959962A
Text normalizing method and device
CN111435595A
Webpage sensitive word detection method, detection system and related device
CN111680128A
Data evaluation method and system based on combination of double dictionary trees and AC automaton
CN119089899A