Speech recognition method and device, electronic equipment and storage medium
By combining context and semantic information to identify easy-to-mix word segmentation in the coal mine industry, and correcting pronunciation similarity, the problems of low speech recognition rate and high error recognition rate in the industry are solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510141661.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-13
AI Technical Summary
In the coal mine industry, existing speech recognition technology has problems such as low recognition rate and high misidentification rate due to noise interference and the diversity of professional terms.
In the context of the pronunciation to be recognized, the word segmentation exists in the initial speech recognition result based on semantic information, and the word segmentation error is corrected using the pronunciation similarity between the word segmentation and the context corresponding setting word segmentation, and the accurate speech recognition result is obtained.
It improves the accuracy and robustness of speech recognition, reduces the misidentification rate in noise interference and professional term recognition, and is suitable for the coal mining industry and other complex context scenarios.
Smart Images

Figure CN119993128A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition technology, and in particular to a speech recognition method, device, electronic equipment and storage medium. Background Art
[0002] Speech recognition is a technology that converts human speech signals into text information through computer technology and algorithms. It is widely used in various fields to improve the efficiency and convenience of human-computer interaction. However, in the specific scenario of the coal mining industry, the application of speech recognition technology faces many challenges.
[0003] At present, speech recognition is mostly based on general speech recognition models. However, due to the noise interference of the coal mine working environment and the diversity of professional terminology, the existing speech recognition solutions still have problems such as low recognition rate and high error rate in practical applications. Summary of the invention
[0004] The present invention provides a speech recognition method, device, electronic equipment and storage medium to solve the defects in the prior art.
[0005] The present invention provides a speech recognition method, comprising the following steps: Perform speech recognition on the speech to be recognized to obtain an initial speech recognition result; In the context of the speech to be recognized, based on semantic information of the speech to be recognized, identifying easily confused segmentations in the initial speech recognition result; Based on the pronunciation similarity between the easily confused segmented words and the segmented words corresponding to the context, segmentation errors are corrected on the initial speech recognition result to obtain a speech recognition result of the speech to be recognized.
[0006] According to a speech recognition method provided by the present invention, in the context of the speech to be recognized, based on the semantic information of the speech to be recognized, identifying the easily confused segmentation words in the initial speech recognition result includes: In the context of the speech to be recognized, based on semantic information of the speech to be recognized, identifying a first easily confused participle present in the initial speech recognition result; Based on the pronunciation deviation of each participle in the initial speech recognition result, a second easily confused participle in the initial speech recognition result is identified, wherein the pronunciation deviation of each participle is used to characterize the deviation between the speaker pronunciation corresponding to each participle and the standard pronunciation.
[0007] According to a speech recognition method provided by the present invention, the pronunciation deviation of each word segment is determined based on the following steps: Perform phoneme recognition on the standard pronunciation of each participle to obtain a standard phoneme recognition result of each participle; Based on the standard phoneme recognition results of each segmented word and the phonemes included in each segmented word, the pronunciation deviation of each segmented word is determined.
[0008] According to a speech recognition method provided by the present invention, performing speech recognition on a speech to be recognized to obtain an initial speech recognition result includes: Determine a plurality of hot words in the scene to which the to-be-recognized speech belongs; In the process of performing speech recognition on the speech to be recognized, if the pronunciation similarity between the recognized current participle and any hot word is greater than a threshold, the corresponding hot word is used as the participle in the initial speech recognition result.
[0009] According to a speech recognition method provided by the present invention, speech recognition is performed on a speech to be recognized to obtain an initial speech recognition result; in the context of the speech to be recognized, based on the semantic information of the speech to be recognized, easily confused segmentations present in the initial speech recognition result are recognized; based on the pronunciation similarity between the easily confused segmentations and the segmentations set corresponding to the context, segmentation errors are corrected on the initial speech recognition result to obtain the speech recognition result of the speech to be recognized, including: Based on the speech recognition model, speech recognition is performed on the speech to be recognized to obtain an initial speech recognition result; in the context of the speech to be recognized, based on the semantic information of the speech to be recognized, easily confused participles in the initial speech recognition result are recognized; based on the pronunciation similarity between the easily confused participles and the participles set corresponding to the context, the initial speech recognition result is corrected for the word segmentation to obtain the speech recognition result of the speech to be recognized.
[0010] According to a speech recognition method provided by the present invention, the structure of the speech recognition model is determined based on the following steps: Collecting historical speech in the scene to which the speech to be recognized belongs; Extracting acoustic features of the historical speech; Based on the characteristic rules of the acoustic features, the structure of the speech recognition model is determined.
[0011] According to a speech recognition method provided by the present invention, the speech recognition model is obtained by fine-tuning a pre-trained model based on sample speech in a scene to which the speech to be recognized belongs.
[0012] The present invention also provides a speech recognition device, comprising the following units: The first recognition unit is used to perform speech recognition on the speech to be recognized and obtain an initial speech recognition result; A second recognition unit is used to recognize easily confused participles in the initial speech recognition result based on the semantic information of the speech to be recognized in the context of the speech to be recognized; The word segmentation error correction unit is used to perform word segmentation error correction on the initial speech recognition result based on the pronunciation similarity between the easily confused word segments and the context corresponding setting word segments to obtain the speech recognition result of the speech to be recognized.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-mentioned speech recognition methods is implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the speech recognition method described in any one of the above is implemented.
[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the speech recognition method described above is implemented.
[0016] The speech recognition method, device, electronic device and storage medium provided by the present invention use the context of the speech to be recognized to characterize the background information and contextual relationship of the communication scene, and the semantic information of the speech to be recognized to characterize the intrinsic meaning and specific direction of the speech to be recognized, so that in the context of the speech to be recognized, combined with its semantic information, it is possible to more comprehensively understand the exact meaning of the vocabulary in a specific scene, so that the easily confused words in the initial speech recognition results can be accurately recognized. In addition, the present invention sets the pronunciation similarity between the easily confused word segmentation and the context correspondingly, which can avoid the problem of misrecognition caused by the similarity between the easily confused word and the acoustic feature, and then accurately correct the initial speech recognition result, and obtain the speech recognition result of the speech to be recognized. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is a flow chart of the speech recognition method provided by the present invention.
[0019] Figure 2 It is a flow chart of another speech recognition method provided by the present invention.
[0020] Figure 3 It is a structural schematic diagram of the speech recognition device provided by the present invention.
[0021] Figure 4It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0023] In the specific and complex working environment of the coal mining industry, the application of speech recognition technology has encountered unprecedented challenges. Coal mining operations are often accompanied by high-intensity mechanical noise and complex environmental noise, which greatly interferes with the clarity of speech signals and brings great difficulties to speech recognition.
[0024] At the same time, the coal mining industry has a large number of professional and unique terms, which often exceed the recognition range of general speech recognition models, resulting in frequent problems such as low recognition rate and high misrecognition rate. Therefore, how to overcome noise interference and improve the recognition ability of professional terms in the special scenario of the coal mining industry has become a key issue that needs to be solved in current speech recognition technology.
[0025] In this regard, the present invention provides a speech recognition method, which can be applied not only to the special scenario of the coal mining industry, but also to other scenarios with a large number of professional terms or complex contexts, such as medical scenarios, educational scenarios, smart home scenarios, etc. In order to facilitate understanding of the technical solution of the present invention, the following embodiments are all described by taking the application in the coal mining industry scenario as an example.
[0026] in, Figure 1 It is a flow chart of the speech recognition method provided by the present invention, such as Figure 1 As shown, the method includes step 110 , step 120 and step 130 .
[0027] Step 110: Perform speech recognition on the speech to be recognized to obtain an initial speech recognition result.
[0028] Here, the speech to be recognized is speech data that needs to be recognized. The speech to be recognized can be speech data input by a user in real time through a sound collection device, or can be pre-stored or received speech data, which is not specifically limited in the embodiment of the present invention.
[0029] In addition, the initial speech recognition result refers to the text sequence obtained after the speech to be recognized is converted by speech recognition technology. The initial speech recognition result may contain some errors, such as incomplete recognition, wrong words, or the recognized text is irrelevant to the original speech content, so further processing is required to improve the accuracy.
[0030] Taking the coal mine scenario as an example, the speech to be recognized is "The underground equipment has failed. Please send someone to repair it as soon as possible. The equipment number is 1203." After the speech to be recognized is processed by speech recognition technology, the initial speech recognition result may be "The underground equipment has failed. Please send someone to repair it as soon as possible. The equipment number is 1203." In this example, although the initial speech recognition result roughly conveys the main content of the speech, there may be some problems. For example, in the part of "The equipment number is 1203", due to the limitations of speech recognition, the number "1203" may be mistakenly recognized as "1203" with a similar pronunciation (here "1" is the common name for the number "1" in spoken Chinese).
[0031] For example, before performing speech recognition on the speech to be recognized, the speech to be recognized may be preprocessed to filter out noise therein and convert complex and irregular speech signals into digital information to improve speech recognition efficiency and success rate. For example, preprocessing operations such as pre-emphasis and framing may be performed on the speech to be recognized.
[0032] Step 120: In the context of the speech to be recognized, based on the semantic information of the speech to be recognized, identify easily confused word segmentations in the initial speech recognition result.
[0033] Given the diversity and complexity of speech signals, as well as differences in speaker characteristics (such as habits, speaking speed, and intonation) and the influence of environmental noise and background interference, the initial speech recognition results inevitably have certain errors. Specifically, due to the high similarity of acoustic features and the inherent limitations of speech recognition algorithms, the initial recognition results often contain easily confused words, which may be misjudged as other words with similar pronunciations but different meanings during the recognition process, thus causing misunderstandings or deviations in the initial recognition results at the semantic level. Among them, easily confused words refer to those words that are close or similar in pronunciation but different in semantics. For example, "early shift" and "boss" have similar pronunciations but different semantics.
[0034] Considering that the context of the speech to be recognized is used to represent the background information and contextual relationship of the communication scene, the semantic information of the speech to be recognized is used to represent its intrinsic meaning and specific direction. Then, in the context of the speech to be recognized, combined with its semantic information, it is possible to more comprehensively understand the exact meaning of the vocabulary in a specific scene, so that it is possible to accurately recognize and correct the easily confused words in the initial speech recognition results. If easily confused words are only recognized in the context of the speech to be recognized, while ignoring their specific semantic content, the lack of sufficient semantic support may lead to a deviation in the understanding of the meaning of the words, thereby causing misrecognition or confusion problems. Conversely, if easily confused words are identified only based on the semantic information of the speech to be recognized, while ignoring its contextual background, it may be difficult to distinguish between words with similar meanings but different usages in different contexts due to a lack of understanding of the vocabulary usage scenarios, which will also cause recognition errors.
[0035] That is to say, the embodiment of the present invention deeply integrates the semantic information of the speech to be recognized in the context of the speech to be recognized to identify and correct easily confused words in the initial speech recognition results. This fully takes into account the limitations that may be encountered when the recognition relies solely on the context or solely on semantic information, thereby being able to more accurately and comprehensively identify easily confused words.
[0036] Optionally, natural language processing (NLP) technology can be used in combination with deep learning algorithms (such as Transformer models, BERT and other pre-trained language models) to deeply analyze and extract semantic information in the context of the speech to be recognized through semantic parsing, context understanding and other means, effectively distinguish the subtle differences between easily confused words in specific contexts, and thus accurately identify and correct easily confused word segmentations in the initial speech recognition results.
[0037] Step 130: Based on the pronunciation similarity between the easily confused segmented words and the segmented words corresponding to the context, the initial speech recognition result is corrected for the segmented words to obtain the speech recognition result of the speech to be recognized.
[0038] Specifically, context-corresponding word segmentation refers to a set of correct word segmentations that are preset in a specific context based on language habits, professional knowledge, or terminology standards in a specific field. This word segmentation set is intended to accurately reflect the correct combination and customary usage of words in that context. Since context-corresponding word segmentation fully considers the constraints of contextual factors on word combination and usage, it can effectively narrow the scope of easily confused word segmentations in a specific context, that is, reduce those word combinations that may cause ambiguity or misidentification in different contexts, thereby improving the accuracy and efficiency of word segmentation error correction.
[0039] In addition, considering that there is often similarity in acoustic features between easily confused segmentation words and context corresponding segmentation words, they are easily misrecognized during speech recognition. Since pronunciation similarity is used to characterize the closeness of two segmentation words in pronunciation, based on the pronunciation similarity between easily confused segmentation words and context corresponding segmentation words, the problem of misrecognition caused by the similarity of easily confused words and acoustic features can be avoided, and the initial speech recognition result can be accurately corrected to obtain the corrected initial speech recognition result, that is, the speech recognition result of the speech to be recognized.
[0040] Exemplarily, if the pronunciation similarity between the easily confused participle and any set participle in the context is greater than a threshold, the corresponding set participle is added to the candidate participle set, and the set participle with the greatest pronunciation similarity is selected from the candidate participle set to replace the easily confused participle in the initial speech recognition result, so as to obtain the speech recognition result of the speech to be recognized.
[0041] For example, in the context of "shift report" in the coal mine scenario, the defined shift types include "morning shift, mid-shift, night shift", which are set as participles in this specific context. However, due to differences in user accents, the speech recognition system often mistakenly recognizes "morning shift" as "boss" in this context. Given that it is illogical for users to express "boss" in the context of "shift report", combined with the pronunciation similarity between "boss" and the set participle shown in the acoustic analysis, it can be determined that the pronunciation between "boss" and "morning shift" is highly similar. Therefore, it can be reasonably inferred that the "boss" that is misrecognized in this context should actually refer to "morning shift". Based on this, when the system subsequently processes the speech recognition results in this context, if it encounters "boss" again, it will automatically correct it to "morning shift". However, it is worth noting that this automatic correction mechanism is limited to the context of "shift report". In other non-related contexts, the word "boss" will retain its original meaning and will not be automatically converted to "morning shift".
[0042] In the speech recognition method provided by the embodiment of the present invention, the context of the speech to be recognized is used to characterize the background information and contextual relationship of the communication scene, and the semantic information of the speech to be recognized is used to characterize the intrinsic meaning and specific direction of the speech to be recognized, and then in the context of the speech to be recognized, combined with its semantic information, it is possible to more comprehensively understand the exact meaning of the vocabulary in a specific scene, so that it is possible to accurately recognize and correct the easily confused words in the initial speech recognition results. In addition, the embodiment of the present invention sets the pronunciation similarity between the easily confused word segmentations based on the correspondence between the easily confused word segmentations and the context, which can avoid the problem of misrecognition caused by the similarity between the easily confused word segmentations and the acoustic features, and then accurately perform segmentation error correction on the initial speech recognition results to obtain the speech recognition results of the speech to be recognized.
[0043] Based on the above embodiment, in the context of the speech to be recognized, based on the semantic information of the speech to be recognized, identifying the easily confused segmented words in the initial speech recognition result includes: In the context of the speech to be recognized, based on semantic information of the speech to be recognized, identifying a first easily confused participle present in the initial speech recognition result; Based on the pronunciation deviation of each segmented word in the initial speech recognition result, a second easily confused segmented word in the initial speech recognition result is identified, and the pronunciation deviation of each segmented word is used to characterize the deviation between the speaker pronunciation corresponding to each segmented word and the standard pronunciation.
[0044] Considering that in the context of the speech to be recognized, the confusing word recognition of the initial speech recognition result is based on the semantic information of the speech to be recognized, which mainly relies on the semantic relationship and contextual information of the vocabulary, and does not take into account the pronunciation characteristics of the speaker. Therefore, the first easily confused word segmentation in the initial speech recognition result obtained based on the context of the speech to be recognized and the semantic information recognition may be incomplete.
[0045] On this basis, considering that the pronunciation deviation of each segmented word in the initial speech recognition result is used to characterize the deviation between the speaker pronunciation corresponding to each segmented word and the standard pronunciation, the pronunciation deviation of each segmented word can be used to obtain pronunciation similarity or pronunciation feature information, and then based on the pronunciation deviation of each segmented word in the initial speech recognition result, the second easily confused segmented word in the initial speech recognition result can be identified. Since the second easily confused segmented word is determined based on pronunciation similarity or pronunciation feature, it can make up for the deficiency of the pronunciation feature of the first easily confused segmented word obtained based on semantic information recognition, so as to complete and enhance the recognition of easily confused words.
[0046] In addition, if the second easily confused segmented words are identified based solely on the pronunciation deviations of each segmented word in the initial speech recognition results, since this method ignores the importance of semantic coherence and contextual context, the identified second easily confused segmented word set may not be comprehensive enough, and may miss words that should be judged as easily confused based on semantics and context in a specific context, but whose pronunciation deviations are actually not sufficient to trigger an alarm alone.
[0047] Based on this, the embodiment of the present invention integrates the first easily confused word segmentation (obtained based on semantic information and context recognition) and the second easily confused word segmentation (obtained based on pronunciation deviation and pronunciation feature recognition), and thus can take into account the differences in pronunciation characteristics while combining the information of semantic coherence and contextual context, thereby achieving accurate and comprehensive recognition of easily confused words in the speech to be recognized, thereby improving the accuracy and robustness of speech recognition.
[0048] Based on any of the above embodiments, the pronunciation deviation of each word segment is determined based on the following steps: Perform phoneme recognition on the standard pronunciation of each participle to obtain a standard phoneme recognition result of each participle; Based on the standard phoneme recognition results of each segmented word and the phonemes included in each segmented word, the pronunciation deviation of each segmented word is determined.
[0049] Specifically, phoneme recognition refers to decomposing the standard pronunciation of each word into the smallest sound unit that constitutes these pronunciations, and the smallest sound unit is the phoneme. For example, the word "cat" can be decomposed into three phonemes during phoneme recognition: [k] [æ][t]. Among them, the standard pronunciation of each word can be understood as the pronunciation of the word expressed in accordance with the standardized pronunciation rules and methods.
[0050] After performing phoneme recognition on the standard pronunciation of each segmented word, a phoneme sequence contained in each segmented word under the standard pronunciation is obtained. The phoneme sequence can be understood as the standard phoneme recognition result of each segmented word.
[0051] It is worth noting that the speech to be recognized may come from the speaker's expression in a dialect, and different languages have unique phoneme systems and pronunciation rules. This means that for the same participle, if it is expressed in different languages, the phonemes it contains may be different.
[0052] The phonemes contained in each word follow the phoneme system corresponding to the language currently used by the speaker. If the speaker does not use a standard language (such as Mandarin), the phonemes contained in each word may differ from the phonemes in the standard phoneme recognition results of the corresponding word, which means that the pronunciation deviation is shown.
[0053] In the case of pronunciation deviation, it indicates that the speaker may be communicating in dialect. At this time, when performing speech recognition on the segmented words expressed in the dialect, the probability of misrecognition will be relatively high. Therefore, the embodiment of the present invention can identify the second easily confused segmented words in the initial speech recognition result by evaluating the pronunciation deviation of each segmented word in the initial speech recognition result, thereby effectively improving the recognition accuracy of the dialect speech.
[0054] Based on any of the above embodiments, performing speech recognition on the speech to be recognized to obtain an initial speech recognition result includes: Determine multiple hot words in the scene to which the speech to be recognized belongs; In the process of performing speech recognition on the speech to be recognized, if the pronunciation similarity between the current segmented word obtained by recognition and any hot word is greater than a threshold, the corresponding hot word is used as the segmented word in the initial speech recognition result.
[0055] Considering that the same pronunciation may correspond to different semantic segmentations, and these segmentations correspond to their own professional terms or common words in different scenarios, these professional terms may be used frequently in certain scenarios, but may appear less frequently in other scenarios, or even be regarded as incorrect words. In order to further improve the accuracy of speech recognition, the embodiment of the present invention is committed to determining multiple hot words in the scenario to which the speech to be recognized belongs. These hot words refer to professional terms or common expressions that may appear frequently and have clear meanings in specific scenarios.
[0056] Taking the coal mine scene as an example, the corresponding hot words may include "tunnel" and "coal mining machine", which appear frequently in the coal mine operation environment. When performing speech recognition on the speech to be recognized, even if the probability value of the current segmentation is low, if its pronunciation similarity with any hot word exceeds the preset threshold, it indicates that in the current scene, the segmentation is very likely to be the corresponding hot word. Therefore, the hot word can be output as the segmentation in the initial speech recognition result.
[0057] For example, a phoneme sequence pronounced as "hang dao" may be mistakenly recognized as the general word "channel" in daily contexts. However, in the specific application scenario of coal mines, by adding the professional term "laneway" to the hot word library, when the speech recognition process is performed on the speech to be recognized, if the pronunciation is the same as the hot word in the hot word library, it can be accurately recognized as "laneway" based on the priority of the hot word library, thus effectively avoiding the phenomenon of misrecognition.
[0058] Based on any of the above embodiments, speech recognition is performed on the speech to be recognized to obtain an initial speech recognition result; in the context of the speech to be recognized, based on the semantic information of the speech to be recognized, easily confused segmented words in the initial speech recognition result are recognized; based on the pronunciation similarity between the easily confused segmented words and the corresponding setting segmented words in the context, segmentation errors are corrected on the initial speech recognition result to obtain a speech recognition result of the speech to be recognized, including: Based on the speech recognition model, speech recognition is performed on the speech to be recognized to obtain an initial speech recognition result; in the context of the speech to be recognized, based on the semantic information of the speech to be recognized, easily confused segmented words in the initial speech recognition result are recognized; based on the pronunciation similarity between the easily confused segmented words and the context correspondingly set segmented words, the initial speech recognition result is corrected for segmentation errors to obtain a speech recognition result of the speech to be recognized.
[0059] Here, when performing speech recognition on the speech to be recognized, a speech recognition model can be used for recognition, that is, the speech to be recognized is input into the speech recognition model, and the speech recognition model performs speech recognition on the speech to be recognized to obtain an initial speech recognition result; in the context of the speech to be recognized, based on the semantic information of the speech to be recognized, the easily confused segmented words in the initial speech recognition result are recognized; based on the pronunciation similarity between the easily confused segmented words and the corresponding segmented words in the context, the initial speech recognition result is corrected for segmentation errors to obtain a speech recognition result of the speech to be recognized. Among them, the speech recognition model can be trained based on sample speech and sample speech recognition labels corresponding to the sample speech.
[0060] Based on any of the above embodiments, the structure of the speech recognition model is determined based on the following steps: Collect historical speech in the scene to which the speech to be recognized belongs; Extract acoustic features of historical speech; Based on the characteristic rules of acoustic features, the structure of the speech recognition model is determined.
[0061] Given that analog landlines need to be deployed in mine environments, they are connected through wired links that are several kilometers long. During this process, the signal is susceptible to loss and various interferences, resulting in the audio quality of landline channels being much lower than that of ordinary 4G mobile communication channels. In addition, after the landline signal is processed by the program-controlled switch, the sampling frequency is reduced to 8kHz, which is incompatible with the 16kHz recognition engine used by the mainstream speech recognition model, resulting in a significant reduction in the speech recognition effect, or even recognition failure.
[0062] In this regard, the embodiment of the present invention collects historical speech in the scene to which the speech to be recognized belongs (such as the coal mine scene). Since the characteristic laws of the acoustic features of the historical speech can reflect the transmission characteristics and quality requirements of the speech signal in the scene, the structure of the speech recognition model can be customized or optimized based on the characteristic laws of the acoustic features of these historical speech. The recognition engine of the speech recognition model under this structure matches the speech transmission speed (such as 8k sampling frequency) in the scene to which the speech to be recognized belongs, thereby ensuring that the speech recognition model can accurately and efficiently recognize the speech signal in the scene.
[0063] Among them, the acoustic features of the historical speech here may include frequency characteristics, noise level, distortion degree, dynamic range, etc. The characteristic laws of acoustic features can be understood as the statistical laws or changing trends of acoustic features in specific scenarios, such as the stability of frequency characteristics, the volatility of noise levels, the increasing trend of distortion degree, etc.
[0064] For example, if the characteristic law of acoustic features is that the frequency characteristics of speech signals in coal mine scenarios are relatively stable, and the noise level fluctuates within a certain range, and the degree of distortion gradually increases with the increase of transmission distance, then the structure of the speech recognition model can be designed to adopt a recognition engine that adapts to the 8k sampling frequency, enhance noise suppression and filtering functions, introduce a distortion compensation mechanism, etc.
[0065] Based on any of the above embodiments, the speech recognition model is obtained by fine-tuning the pre-trained model based on sample speech in the scene to which the speech to be recognized belongs.
[0066] Specifically, the pre-trained model can be understood as a pre-trained speech recognition model, which may be trained on a large general corpus and has basic speech recognition capabilities, but it may not be fully suitable for specific scenarios, such as coal mine scenarios.
[0067] In order to improve the recognition accuracy of the speech recognition model in the scenario to which the speech to be recognized belongs, it is necessary to collect a large number of sample speech in the corresponding scenario. These sample speech can contain professional terms, common expressions and specific contextual relationships of the corresponding scenario so that the model can learn these particularities and complexities.
[0068] Based on the sample speech in the corresponding scenario, the pre-trained model is fine-tuned to obtain a speech recognition model. Fine-tuning means further training the model on the sample speech in a specific scenario so that it can better adapt to and recognize the speech data in the corresponding scenario. Fine-tuning usually includes adjusting the parameters of the model, optimizing the model structure, or introducing domain-related features.
[0069] The fine-tuned speech recognition model not only retains the basic recognition capabilities of the pre-trained model, but also optimizes the scene to which the speech to be recognized belongs. Therefore, the speech recognition model can better understand the meaning and context of professional terms in the corresponding scene, thereby improving the accuracy of speech recognition in the corresponding scene.
[0070] Based on any of the above embodiments, Figure 2 It is a flow chart of another speech recognition method provided by the present invention, such as Figure 2 As shown, the method includes: First, historical speech in the coal mine scenario is collected, and the structure of the speech recognition model is determined based on the characteristic laws of the acoustic features of the historical speech. The recognition engine of the speech recognition model under this structure matches the speech transmission speed (such as 8k sampling frequency) in the scenario to which the speech to be recognized belongs, thereby ensuring that the speech recognition model can accurately and efficiently recognize the speech signal in this scenario.
[0071] Next, based on the structure of the speech recognition model, a pre-trained model with the same structure is obtained. The pre-trained model may be trained on a large number of general corpora and has basic speech recognition capabilities, but it may not be fully applicable to specific scenarios, such as coal mine scenarios.
[0072] Sample speech in coal mine scenarios (including sample speech of professional terms in coal mine scenarios) is obtained, and the pre-trained model is fine-tuned based on the sample speech to obtain a speech recognition model.
[0073] After obtaining the speech recognition model, the speech to be recognized in the coal mine scenario is input into the speech recognition model, and the speech recognition model performs speech recognition on the speech to be recognized to obtain the initial speech recognition result; in the context of the speech to be recognized, based on the semantic information of the speech to be recognized, the easily confused segmented words in the initial speech recognition result are recognized; based on the pronunciation similarity between the easily confused segmented words and the context corresponding to the set segmented words, the initial speech recognition result is corrected for the segmented words, and the speech recognition result of the speech to be recognized is obtained. In the process of performing speech recognition on the speech to be recognized, if the pronunciation similarity between the current segmented word obtained by recognition and any hot word is greater than the threshold, the corresponding hot word is used as the segmented word in the initial speech recognition result. In the hot word, it is a hot word in the coal mine scenario.
[0074] The speech recognition device provided by the present invention is described below. The speech recognition device described below and the speech recognition method described above can be referred to each other.
[0075] Based on any of the above embodiments, Figure 3 : is a schematic diagram of the structure of the speech recognition device provided by the present invention, such as Figure 3 As shown, the device comprises: The first recognition unit 310 is used to perform speech recognition on the speech to be recognized and obtain an initial speech recognition result; The second recognition unit 320 is used to recognize easily confused segment words in the initial speech recognition result based on the semantic information of the speech to be recognized in the context of the speech to be recognized; The word segmentation error correction unit 330 is used to perform word segmentation error correction on the initial speech recognition result based on the pronunciation similarity between the easily confused word segments and the context corresponding setting words, so as to obtain the speech recognition result of the speech to be recognized.
[0076] Based on any of the above embodiments, in the context of the speech to be recognized, based on the semantic information of the speech to be recognized, identifying easily confused segmented words in the initial speech recognition result includes: In the context of the speech to be recognized, based on semantic information of the speech to be recognized, identifying a first easily confused participle present in the initial speech recognition result; Based on the pronunciation deviation of each segmented word in the initial speech recognition result, a second easily confused segmented word in the initial speech recognition result is identified, and the pronunciation deviation of each segmented word is used to characterize the deviation between the speaker pronunciation corresponding to each segmented word and the standard pronunciation.
[0077] Based on any of the above embodiments, the pronunciation deviation of each word segment is determined based on the following steps: Perform phoneme recognition on the standard pronunciation of each participle to obtain a standard phoneme recognition result of each participle; Based on the standard phoneme recognition results of each segmented word and the phonemes included in each segmented word, the pronunciation deviation of each segmented word is determined.
[0078] Based on any of the above embodiments, performing speech recognition on the speech to be recognized to obtain an initial speech recognition result includes: Determine multiple hot words in the scene to which the speech to be recognized belongs; In the process of performing speech recognition on the speech to be recognized, if the pronunciation similarity between the current segmented word obtained by recognition and any hot word is greater than a threshold, the corresponding hot word is used as the segmented word in the initial speech recognition result.
[0079] Based on any of the above embodiments, speech recognition is performed on the speech to be recognized to obtain an initial speech recognition result; in the context of the speech to be recognized, based on the semantic information of the speech to be recognized, easily confused segmented words in the initial speech recognition result are recognized; based on the pronunciation similarity between the easily confused segmented words and the corresponding setting segmented words in the context, segmentation errors are corrected on the initial speech recognition result to obtain a speech recognition result of the speech to be recognized, including: Based on the speech recognition model, speech recognition is performed on the speech to be recognized to obtain an initial speech recognition result; in the context of the speech to be recognized, based on the semantic information of the speech to be recognized, easily confused segmented words in the initial speech recognition result are recognized; based on the pronunciation similarity between the easily confused segmented words and the context correspondingly set segmented words, the initial speech recognition result is corrected for segmentation errors to obtain a speech recognition result of the speech to be recognized.
[0080] Based on any of the above embodiments, the structure of the speech recognition model is determined based on the following steps: Collect historical speech in the scene to which the speech to be recognized belongs; Extract acoustic features of historical speech; Based on the characteristic rules of acoustic features, the structure of the speech recognition model is determined.
[0081] Based on any of the above embodiments, the speech recognition model is obtained by fine-tuning the pre-trained model based on sample speech in the scene to which the speech to be recognized belongs.
[0082] Figure 4 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 4As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430 and a communication bus 440, wherein the processor 410, the communication interface 420 and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the speech recognition method, which includes: performing speech recognition on the speech to be recognized to obtain an initial speech recognition result; in the context of the speech to be recognized, based on the semantic information of the speech to be recognized, identifying the easily confused segmentation words in the initial speech recognition result; based on the pronunciation similarity between the easily confused segmentation words and the segmentation words corresponding to the context, performing segmentation error correction on the initial speech recognition result to obtain the speech recognition result of the speech to be recognized.
[0083] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0084] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the speech recognition method provided by the above-mentioned methods, and the method includes: performing speech recognition on the speech to be recognized to obtain an initial speech recognition result; in the context of the speech to be recognized, based on the semantic information of the speech to be recognized, identifying easily confused participles existing in the initial speech recognition result; based on the pronunciation similarity between the easily confused participles and the participles set corresponding to the context, performing segmentation error correction on the initial speech recognition result to obtain the speech recognition result of the speech to be recognized.
[0085] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the speech recognition method provided by the above-mentioned methods, the method comprising: performing speech recognition on a speech to be recognized to obtain an initial speech recognition result; in the context of the speech to be recognized, based on the semantic information of the speech to be recognized, identifying easily confused participles present in the initial speech recognition result; based on the pronunciation similarity between the easily confused participles and the participles set corresponding to the context, performing segmentation error correction on the initial speech recognition result to obtain a speech recognition result of the speech to be recognized.
[0086] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0087] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A speech recognition method, characterized in that: include: Perform speech recognition on the speech to be recognized to obtain an initial speech recognition result; In the context of the speech to be recognized, based on semantic information of the speech to be recognized, identifying easily confused segmentations in the initial speech recognition result; Based on the pronunciation similarity between the easily confused segmented words and the segmented words corresponding to the context, segmentation errors are corrected on the initial speech recognition result to obtain a speech recognition result of the speech to be recognized.
2. The speech recognition method according to claim 1, characterized in that: The step of identifying easily confused segmented words in the initial speech recognition result based on semantic information of the speech to be recognized in the context of the speech to be recognized includes: In the context of the speech to be recognized, based on semantic information of the speech to be recognized, identifying a first easily confused participle present in the initial speech recognition result; Based on the pronunciation deviation of each participle in the initial speech recognition result, a second easily confused participle in the initial speech recognition result is identified, wherein the pronunciation deviation of each participle is used to characterize the deviation between the speaker pronunciation corresponding to each participle and the standard pronunciation.
3. The speech recognition method according to claim 2, characterized in that: The pronunciation deviation of each word segment is determined based on the following steps: Perform phoneme recognition on the standard pronunciation of each participle to obtain a standard phoneme recognition result of each participle; Based on the standard phoneme recognition results of each segmented word and the phonemes included in each segmented word, the pronunciation deviation of each segmented word is determined.
4. The speech recognition method according to any one of claims 1 to 3, characterized in that: The performing speech recognition on the speech to be recognized to obtain an initial speech recognition result includes: Determine a plurality of hot words in the scene to which the to-be-recognized speech belongs; In the process of performing speech recognition on the speech to be recognized, if the pronunciation similarity between the recognized current participle and any hot word is greater than a threshold, the corresponding hot word is used as the participle in the initial speech recognition result.
5. The speech recognition method according to any one of claims 1 to 3, characterized in that: The method comprises: performing speech recognition on the speech to be recognized to obtain an initial speech recognition result; in the context of the speech to be recognized, based on the semantic information of the speech to be recognized, identifying easily confused segmentations in the initial speech recognition result; performing segmentation error correction on the initial speech recognition result based on the pronunciation similarity between the easily confused segmentations and the segmentations set corresponding to the context to obtain the speech recognition result of the speech to be recognized, including: Based on the speech recognition model, speech recognition is performed on the speech to be recognized to obtain an initial speech recognition result; in the context of the speech to be recognized, based on the semantic information of the speech to be recognized, easily confused participles present in the initial speech recognition result are recognized; based on the pronunciation similarity between the easily confused participles and the participles set corresponding to the context, the initial speech recognition result is corrected for word segmentation to obtain the speech recognition result of the speech to be recognized.
6. The speech recognition method according to claim 5, characterized in that: The structure of the speech recognition model is determined based on the following steps: Collecting historical speech in the scene to which the speech to be recognized belongs; Extracting acoustic features of the historical speech; Based on the characteristic rules of the acoustic features, the structure of the speech recognition model is determined.
7. The speech recognition method according to claim 5, characterized in that: The speech recognition model is obtained by fine-tuning a pre-trained model based on sample speech in the scene to which the speech to be recognized belongs.
8. A speech recognition device, characterized in that: include: The first recognition unit is used to perform speech recognition on the speech to be recognized and obtain an initial speech recognition result; A second recognition unit is used to recognize easily confused participles in the initial speech recognition result based on the semantic information of the speech to be recognized in the context of the speech to be recognized; The word segmentation error correction unit is used to perform word segmentation error correction on the initial speech recognition result based on the pronunciation similarity between the easily confused word segments and the context corresponding setting word segments to obtain the speech recognition result of the speech to be recognized.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the speech recognition method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the speech recognition method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Medicine name identification method and device based on voice information and electronic equipment
CN120783752A
Drug name recognition method and device based on voice information and electronic equipment
CN120783752B