A method, device and related equipment for intelligent speech recognition based on classification identification
By attaching industry attribute identifiers to voice data, using a general model for initial recognition and then optimizing recognition in industry vertical modules, the problems of high recognition costs and long cycles in vertical fields are solved, achieving efficient and accurate voice recognition.
Patent Information
- Application Number
- CN202111682776.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2041-12-31
AI Technical Summary
In vertical fields or professional application scenarios, existing technologies require a large amount of manual labeling and model training, resulting in high costs, long cycles and low adaptability. In addition, industry corpora are unwilling to be shared, leading to the problem of duplicate construction.
By attaching industry attribute identifiers to the original speech data and using a general speech recognition model for preliminary recognition, we optimize the recognition of texts with low confidence levels in the industry vertical module and select the final result based on the maximum confidence level.
No manual labeling is required, and secondary verification and recognition are used to improve recognition accuracy, reduce costs and cycles, and improve adaptability.
Smart Images

Figure CN114242042B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition and semantic understanding technology of artificial intelligence natural language processing, and in particular to an intelligent speech recognition method, device and related equipment based on classification identification. Background Art
[0002] Speech recognition technology allows machines to convert speech signals into corresponding text or commands through a process of recognition and understanding. The general operating principles and processes of general speech recognition systems include endpoint detection, feature extraction, pattern matching, and model training. First, feature extraction and pattern matching techniques are used to identify a specific acoustic model; then, model training is used to form a specific language model. This modeling process requires a large database of speech and language data (involving linguistic knowledge) to form the corresponding acoustic and language models. The implementation process involves rapidly optimizing within the space of these acoustic and language models to convert speech signals into text.
[0003] Semantic understanding technology is the targeted understanding of information expressed in text form, understanding the exact meaning of information expressed in text. The corresponding system can perform subsequent actions based on the understanding results.
[0004] Whether it is speech recognition or semantic understanding technology, it is necessary to train the corresponding model in advance, or specify the corresponding application field, and refine the recognition and understanding model to ensure the accuracy of recognition and understanding.
[0005] However, in actual cases or applications, such as scenarios for vertical fields or professional applications (such as telecommunications, finance, wealth management, etc.), since the language content of relevant professional field scenarios is generally not trained in advance, the recognition effect is often poor and generally cannot be directly called.
[0006] Vertical fields / specific industries generally adopt the method of independently building proprietary speech recognition engine systems. The problems are:
[0007] 1. Professional system modeling requires manual labeling of a large amount of industry-specific voice data, as well as targeted model training and optimization. This results in high investment costs, long training cycles, low adaptability, and difficulty in rapid deployment.
[0008] 2. Most industry corpora are unwilling to be opened and shared due to industry secrets and other factors, resulting in duplicate construction. Summary of the Invention
[0009] The purpose of the present invention is to provide an intelligent speech recognition method, device and related equipment based on classification identification, aiming to solve the problem of high cost and long cycle caused by manual labeling in the existing technology.
[0010] In a first aspect, an embodiment of the present invention provides an intelligent speech recognition method based on classification identification, comprising:
[0011] S101, performing endpoint detection and feature extraction on raw speech data to obtain speech feature data;
[0012] S102, identifying the industry attribute of the original voice data by using the industry attribute identifier attached when the original voice data is input;
[0013] S103, identifying the speech feature data using a trained general speech recognition model to obtain a recognition result;
[0014] S104, determining whether the confidence level of the recognition result is lower than a preset value, if so, executing step S105, if not, executing step S107;
[0015] S105, according to the industry attributes, inputting the corresponding text with a confidence level lower than a preset value into the corresponding industry vertical module for optimization recognition to obtain an optimized confidence level;
[0016] S106, selecting the text corresponding to the maximum confidence value as the final recognition result based on the confidence value in the recognition result and the optimized confidence value;
[0017] S107: Output the recognition result.
[0018] In a second aspect, an embodiment of the present invention provides an intelligent speech recognition device based on classification identification, characterized by comprising:
[0019] A feature extraction unit, configured to perform endpoint detection and feature extraction on raw speech data to obtain speech feature data;
[0020] An industry attribute identification unit, configured to identify the industry attribute of the original voice data by using the industry attribute identifier attached when the original voice data is input;
[0021] A speech conversion recognition unit, configured to recognize the speech feature data using a trained general speech recognition model to obtain a recognition result;
[0022] An optimization recognition unit is used to input corresponding text with a confidence level lower than a preset value into a corresponding industry vertical module for optimization recognition according to the industry attribute to obtain an optimized confidence level;
[0023] a judgment unit, configured to judge whether the confidence level in the recognition result is lower than a preset value, and if so, to perform the operation through the optimization recognition unit; if not, to output the recognition result;
[0024] A selection unit is used to select the text corresponding to the maximum confidence value as the final recognition result based on the confidence value in the recognition result and the optimized confidence value;
[0025] The output unit is used to output the recognition results.
[0026] In the third aspect, an embodiment of the present invention further provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and runnable on the processor. When the processor executes the computer program, it implements the intelligent speech recognition method based on classification identification described in the first aspect above.
[0027] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the intelligent speech recognition method based on classification identification described in the first aspect above.
[0028] The embodiment of the present invention adds an industry attribute identifier when inputting the original voice data, which facilitates the subsequent call of the corresponding industry vertical module for optimized recognition;
[0029] After being recognized by the general speech recognition model, text with low confidence in the recognition results is considered to be content that the general speech recognition model lacks training. The corresponding industry-specific modules are called again for secondary verification and recognition, and correction is performed to achieve the purpose of improving recognition accuracy;
[0030] No manual labeling is required, and a high accuracy rate is maintained through secondary verification, recognition and correction. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0032] Figure 1 A schematic diagram of a flow chart of an intelligent speech recognition method based on classification identification provided by an embodiment of the present invention;
[0033] Figure 2 This is a result block diagram of the intelligent speech recognition device based on classification identification provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0035] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0036] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0037] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0038] See also Figure 1 , an intelligent speech recognition method based on classification identification, comprising:
[0039] S101, performing endpoint detection and feature extraction on raw speech data to obtain speech feature data;
[0040] S102, identifying the industry attribute of the original voice data by using the industry attribute identifier attached when the original voice data is input;
[0041] S103, identifying the speech feature data using a trained general speech recognition model to obtain a recognition result;
[0042] S104, determining whether the confidence level of the recognition result is lower than a preset value, if so, executing step S105, if not, executing step S107;
[0043] S105, according to the industry attributes, inputting the corresponding text with a confidence level lower than a preset value into the corresponding industry vertical module for optimization recognition to obtain an optimized confidence level;
[0044] S106, selecting the text corresponding to the maximum confidence value as the final recognition result based on the confidence value in the recognition result and the optimized confidence value;
[0045] S107: Output the recognition result.
[0046] In this embodiment, by attaching an industry attribute identifier when inputting the original voice data, it is convenient to call the corresponding industry vertical module for optimized recognition later;
[0047] After being recognized by the general speech recognition model, text with low confidence in the recognition results is considered to be content that the general speech recognition model lacks training. The corresponding industry-specific modules are called again for secondary verification and recognition, and correction is performed to achieve the purpose of improving recognition accuracy;
[0048] Through the above steps, there is no need for manual labeling, and a high accuracy rate is maintained by secondary verification, identification and correction.
[0049] In one embodiment, the voice endpoint detection steps are as follows:
[0050] Divide the original voice data into frames;
[0051] Extract features from each frame of raw speech data;
[0052] Train a classifier on a collection of data frames with known speech and silence signal regions;
[0053] The classifier is used to classify the original speech data after frame division to determine whether it belongs to a speech signal or a silence signal.
[0054] Specifically, the frame processing of the original voice data includes:
[0055] The raw speech data is passed through a high-pass filter with a cutoff frequency of 180-220Hz. This step removes the DC offset and some low-frequency noise from the signal. Although some speech information remains below 200Hz, it does not significantly affect the speech signal.
[0056] Before feature extraction, we first need to divide the audio signal into frames of 20-40ms in length, with a typical frame overlap of 10ms. For example, if the audio signal sampling rate is 16kHz and the window size is 25ms, each frame of data will contain 0.025 □ 16000 = 400 samples. Assuming a frame overlap of 10ms, the first frame starts at sample 0, and the second frame starts at sample 160.
[0057] Specifically, features extracted from each frame of raw speech data include:
[0058] Five features are extracted from each frame of data:
[0059] The first feature extraction is logarithm of frame energy:
[0060]
[0061] The second feature extraction is the zero crossing rate: the number of times each frame of data crosses the zero point;
[0062] The third feature extraction is the normalized autocorrelation coefficient at lag 1:
[0063]
[0064] The fourth feature extraction is the first coefficient of the Pth order linear prediction;
[0065] The fifth feature extraction is the logarithm of the Pth-order linear prediction error;
[0066] In this article, P=12, that is, the order of the linear predictor is 12. x(n) is a frame of original speech data, where n ranges from 1 to L (L is the length of each frame of data).
[0067] Preferably, the SVM library libsvm is selected to simply train an SVM classifier for the classification of speech signals and silence signals.
[0068] In one embodiment, the identifying the industry attribute of the original voice data by using the industry attribute identifier attached when the original voice data is input includes:
[0069] Define industry attributes based on industry classification;
[0070] Construct industry vertical modules based on industry classification.
[0071] In this embodiment, industry attributes include professional fields such as finance and medical care. When users in these fields input raw voice data, they need to attach the industry attributes of the industry, which can be a meaningful label or a label with a specific meaning that can be found through a table.
[0072] In one embodiment, constructing industry vertical modules according to industry classification includes:
[0073] Collect professional terms and industry-specific terms from various industries and classify them;
[0074] Perform feature extraction on the classified text data;
[0075] Build corresponding industry vertical modules through the extracted feature data.
[0076] In this embodiment, the construction of the industry vertical module requires collecting professional terms in each industry field, classifying them, and converting the classified text data into feature data to construct the corresponding industry vertical module.
[0077] In one embodiment, the recognition result includes text and corresponding application service domain confidence.
[0078] In this embodiment, by outputting the confidence level of the corresponding text in the recognition result, the correlation between the text and the field is evaluated, which facilitates appropriate scoring and accuracy prediction of the text recognition accuracy.
[0079] In a preferred embodiment, the recognition result includes text, corresponding full spelling, corresponding abbreviated spelling and corresponding application service field confidence, wherein the confidence includes the confidence of the full spelling or abbreviated spelling combination, and the professional field confidence of various combinations obtained by calling professional dictionaries through full spelling or abbreviated spelling.
[0080] In this embodiment, the accuracy is further improved by outputting the text-related pinyin form (full spelling or abbreviated spelling) and its corresponding confidence level in the recognition result, wherein the pinyin form is obtained by calling the corresponding professional dictionary.
[0081] In one embodiment, the inputting of corresponding text with a confidence level lower than a preset value into the corresponding industry vertical module for optimized recognition based on the industry attribute to obtain an optimized confidence level includes:
[0082] Calling the corresponding industry vertical module according to the industry attributes accompanying the input of the original voice data;
[0083] Perform content search, matching, and disambiguation on the original voice data with the same pronunciation and similar text features through the industry vertical module;
[0084] A search text in at least one of the fields is selected as a supplement to the recognition result, and the corresponding confidence level is calculated again.
[0085] In this embodiment, the industry vertical module matches and disambiguates the original speech data for optimization.
[0086] The same original voice data may be accompanied by multiple industry attributes when input, so there may be multiple industry vertical modules when calling. When selecting supplementary recognition results, you can select multiple, calculate the confidence, compare and select the best ones to improve the recognition accuracy.
[0087] See also Figure 2 , an intelligent speech recognition device 10 based on classification identification, comprising:
[0088] A feature extraction unit 11 is used to perform endpoint detection and feature extraction on the original speech data to obtain speech feature data;
[0089] An industry attribute identification unit 12, configured to identify the industry attribute of the original voice data by using the industry attribute identifier attached when the original voice data is input;
[0090] The speech conversion recognition unit 13 is used to recognize the speech feature data through a trained general speech recognition model to obtain a recognition result;
[0091] The optimization recognition unit 14 is configured to input the corresponding text with a confidence level lower than a preset value into the corresponding industry vertical module for optimization recognition according to the industry attribute to obtain an optimized confidence level;
[0092] The judging unit 15 is configured to judge whether the confidence level of the recognition result is lower than a preset value, and if so, to perform the operation through the optimizing recognition unit; if not, to output the recognition result;
[0093] A selection unit 16 is configured to select the text corresponding to the maximum confidence value as the final recognition result based on the confidence values in the recognition result and the optimized confidence values;
[0094] The output unit 17 is used to output the recognition result.
[0095] In one embodiment, the intelligent speech recognition device based on the classification identifier further includes:
[0096] Definition unit, used to define industry attributes according to industry classification;
[0097] Building unit, used to construct industry vertical modules according to industry classification.
[0098] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the intelligent speech recognition method based on classification identification as described above is implemented.
[0099] A computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to perform the intelligent speech recognition method based on classification identification.
[0100] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. An intelligent speech recognition method based on classification identification, characterized in that: include: S101, performing endpoint detection and feature extraction on raw speech data to obtain speech feature data; S102, identifying the industry attribute of the original voice data by using the industry attribute identifier attached when the original voice data is input; S103, identifying the speech feature data using a trained general speech recognition model to obtain a recognition result; S104, determining whether the confidence level of the recognition result is lower than a preset value, if so, executing step S105, if not, executing step S107; S105: Based on the industry attributes, the corresponding text with a confidence level lower than a preset value is input into the corresponding industry vertical module for optimized recognition to obtain optimized confidence levels. If the same original speech data is input with multiple industry attributes, multiple industry vertical modules are correspondingly invoked for optimized recognition to obtain multiple optimized confidence levels. S106, selecting the text corresponding to the maximum confidence value as the final recognition result based on the confidence value in the recognition result and the optimized confidence value; S107, outputting the recognition result; The identifying of the industry attributes of the original voice data by the industry attribute identifier attached when the original voice data is input includes: defining the industry attributes according to the industry classification; and constructing the industry vertical module according to the industry classification; The construction of industry vertical modules according to industry classification includes: collecting professional terms and industry characteristic terms of various industries and classifying them; extracting features from the classified text data; and constructing corresponding industry vertical modules based on the extracted feature data; According to the industry attributes, the corresponding text with a confidence level lower than a preset value is input into the corresponding industry vertical module for optimized recognition to obtain the optimized confidence level, including: calling the corresponding industry vertical module according to the industry attributes attached when the original voice data is input; searching, matching and disambiguating the content of the original voice data with the same pronunciation and similar text features through the industry vertical module; selecting the search text in at least one field as a supplement to the recognition result, and calculating the corresponding confidence level again.
2. The intelligent speech recognition method based on classification identification according to claim 1, characterized in that: The recognition result includes text and corresponding application service domain confidence.
3. The intelligent speech recognition method based on classification identification according to claim 1, characterized in that: The recognition result includes text, corresponding full spelling, corresponding abbreviated spelling and corresponding application service field confidence, wherein the confidence includes the confidence of the full spelling or abbreviated spelling combination, and the professional field confidence of various combinations obtained by calling professional dictionaries through full spelling or abbreviated spelling.
4. An intelligent speech recognition device based on classification identification, characterized in that: include: A feature extraction unit, configured to perform endpoint detection and feature extraction on raw speech data to obtain speech feature data; An industry attribute identification unit, configured to identify the industry attribute of the original voice data by using the industry attribute identifier attached when the original voice data is input; A speech conversion recognition unit, configured to recognize the speech feature data using a trained general speech recognition model to obtain a recognition result; An optimization recognition unit is used to input corresponding text with a confidence level lower than a preset value into a corresponding industry vertical module for optimization recognition according to the industry attribute to obtain an optimized confidence level; a judgment unit, configured to judge whether the confidence level in the recognition result is lower than a preset value, and if so, to perform the operation through the optimization recognition unit; if not, to output the recognition result; A selection unit is used to select the text corresponding to the maximum confidence value as the final recognition result based on the confidence value in the recognition result and the optimized confidence value; An output unit, used for outputting recognition results; The identifying of the industry attributes of the original voice data by the industry attribute identifier attached when the original voice data is input includes: defining the industry attributes according to the industry classification; and constructing the industry vertical module according to the industry classification; The construction of industry vertical modules according to industry classification includes: collecting professional terms and industry characteristic terms of various industries and classifying them; extracting features from the classified text data; and constructing corresponding industry vertical modules based on the extracted feature data; According to the industry attributes, the corresponding text with a confidence level lower than a preset value is input into the corresponding industry vertical module for optimized recognition to obtain the optimized confidence level, including: calling the corresponding industry vertical module according to the industry attributes attached when the original voice data is input; searching, matching and disambiguating the content of the original voice data with the same pronunciation and similar text features through the industry vertical module; selecting the search text in at least one field as a supplement to the recognition result, and calculating the corresponding confidence level again.
5. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the intelligent speech recognition method based on classification identification according to any one of claims 1 to 3 is implemented.
6. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to perform the intelligent speech recognition method based on classification identification according to any one of claims 1 to 3.
Citation Information
Patent Citations
Speech recognition method, device and equipment and storage medium
CN111402861A
Voice processing method and device
CN112259081A