A deep learning-based Chinese speech recognition system and method
By constructing a Chinese speech recognition system based on deep learning, we have achieved accurate recognition and grammatical verification of Chinese speech, solved the problem of incorrect recognition results in existing systems, and improved the recognition effect.
Patent Information
- Application Number
- CN202210848331.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-19
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2042-07-19
AI Technical Summary
Existing speech recognition systems are unable to effectively verify and correct logical errors and typos in Chinese speech recognition results, resulting in poor recognition performance.
A deep learning-based Chinese speech recognition system is constructed, including a speech acquisition module, a speech recognition module, and a correction module. Through the processes of speech acquisition, speech recognition, and grammar correction, the recognition accuracy is ensured.
It improves the accuracy of Chinese speech recognition, ensures the logical correctness and grammatical correctness of the recognition results, and enhances the recognition effect.
Smart Images

Figure CN115240655B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech recognition technology, and in particular to a Chinese speech recognition system and method based on deep learning. Background Technology
[0002] Currently, with the rapid improvement of computer processing power, speech recognition technology has developed rapidly. The application of speech recognition technology is increasingly changing people's production and lifestyles, and it is widely used in fields such as voice input systems, voice control systems and intelligent dialogue query systems.
[0003] However, most speech recognition systems at present can only perform simple recognition of the speech to be recognized and cannot verify the recognized speech text according to the semantics of Chinese, which leads to logical or grammatical errors in the recognition results. At the same time, they cannot effectively correct the typos in the recognition results, which greatly reduces the speech recognition effect.
[0004] Therefore, this invention provides a Chinese speech recognition system and method based on deep learning. Summary of the Invention
[0005] This invention provides a Chinese speech recognition system and method based on deep learning, which uses a Chinese speech recognition model to sequentially recognize the acquired Chinese speech segments to be recognized, and corrects the recognized speech text according to Chinese grammar, thereby ensuring the accuracy of Chinese speech recognition and improving the effect of Chinese speech recognition.
[0006] This invention provides a Chinese speech recognition system based on deep learning, comprising:
[0007] The speech acquisition module is used to receive Chinese speech segments to be recognized in real time and sort the Chinese speech segments to be recognized based on time series.
[0008] The speech recognition module is used to build a Chinese speech recognition model, and based on the sorting results, it sequentially inputs the obtained Chinese speech segments to be recognized into the Chinese speech recognition model for speech recognition to obtain speech text;
[0009] The correction module is used to correct the grammar of the obtained speech text based on the preset Chinese grammar, so as to obtain the final speech recognition text.
[0010] Preferably, a deep learning-based Chinese speech recognition system includes a speech acquisition module comprising:
[0011] A voice acquisition unit is used to monitor the user's current acoustic characteristics in real time and determine the user's current voice state based on the acoustic characteristics, wherein the voice state includes speaking and not speaking;
[0012] The voice recording unit is used to acquire the Chinese voice spoken by the user when the voice state is in the speaking state, and to store the acquired Chinese voice to obtain the Chinese voice segment to be recognized.
[0013] Preferably, a deep learning-based Chinese speech recognition system includes a speech recording unit comprising:
[0014] The speech processing subunit is used to acquire the obtained Chinese speech segment to be recognized, and to perform spectral analysis on the Chinese speech segment to be recognized to obtain the audio spectrum corresponding to the Chinese speech segment to be recognized.
[0015] The speech filtering subunit is used to determine the first peak frequency point of the Chinese speech segment to be identified at each time based on the audio spectrum, and at the same time, to obtain the noise audio spectrum corresponding to the noise signal, and to determine the second peak frequency point of the noise signal based on the noise audio spectrum.
[0016] The speech filtering subunit is used to compare the first peak frequency point with the second peak frequency point, filter out target peak frequency points where the first peak frequency point is greater than the second peak frequency point, and determine the Chinese speech segment to be recognized corresponding to the target peak frequency point as a valid Chinese speech segment to be recognized.
[0017] Preferably, a deep learning-based Chinese speech recognition system includes a speech acquisition module comprising:
[0018] The time determination unit is used to acquire the obtained Chinese speech segment to be identified and process the Chinese speech segment to be identified to obtain the speech signal corresponding to each frame;
[0019] The time determination unit is further configured to determine the time domain information of the Chinese speech segment to be identified based on the speech signal corresponding to each frame, and match the time domain information with the speech signal corresponding to each frame;
[0020] The sorting unit is used to determine the time series corresponding to the Chinese speech segment to be identified based on the matching result, and to sort the Chinese speech segment to be identified in ascending order of the time series, wherein the Chinese speech segment to be identified is at least one segment.
[0021] Preferably, in a deep learning-based Chinese speech recognition system, the sorting unit includes:
[0022] The result acquisition subunit is used to acquire the sorting result of the Chinese speech segments to be recognized, and to determine the target number of the Chinese speech segments to be recognized based on the sorting result;
[0023] The tag acquisition subunit is used to extract the acoustic features of the Chinese speech segment to be identified and determine the speech type of the Chinese speech segment to be identified based on the acoustic features.
[0024] The tagging subunit is used to obtain a target number of tagging tags from a preset tag database based on the speech type, and to tag the Chinese speech segment to be recognized based on the target number of tagging tags.
[0025] Preferably, a deep learning-based Chinese speech recognition system includes a speech recognition module comprising:
[0026] The data acquisition unit is used to acquire speech training text and read the speech training text from a preset speech library using different accents to obtain audio data of the speech training text with different accents.
[0027] A data processing unit is configured to preprocess the audio data, convert the audio data into a corresponding spectrogram based on the preprocessing result, and determine the effective region in the audio data based on the spectrogram.
[0028] The model building unit is used to determine the feature parameters of the audio data based on the effective region, and at the same time, obtain the correspondence between Chinese Pinyin and Chinese characters, train the feature parameters based on the correspondence, and build a Chinese speech recognition model based on the training results.
[0029] The speech recognition unit is used to sequentially input the acquired Chinese speech segments to be recognized into the Chinese speech recognition model, and analyze the received Chinese speech segments to be recognized based on the preset syntax analysis tree in the Chinese speech recognition model to determine the start and end points of each sentence in the Chinese speech segments to be recognized.
[0030] The speech recognition unit is used to perform a first segmentation on each of the Chinese speech segments to be recognized based on the start point and the end point, and to obtain a set of sentences for each of the Chinese speech segments to be recognized based on the first segmentation result, and to extract the syllable attributes contained in each Chinese speech sentence in the set of sentences.
[0031] The speech recognition unit is used to perform a second segmentation on each Chinese speech based on the syllable attributes, and to obtain the Chinese words contained in each Chinese speech based on the second segmentation result.
[0032] The speech recognition unit is further configured to extract the pronunciation features of the Chinese words, and process the pronunciation features based on the correspondence between the Chinese pinyin and Chinese characters to obtain the word text corresponding to the Chinese words;
[0033] The text splicing unit is used to splice the text corresponding to the Chinese words contained in each Chinese speech sentence to obtain the speech text corresponding to the Chinese speech segment to be identified.
[0034] Preferably, a deep learning-based Chinese speech recognition system includes a speech recognition unit comprising:
[0035] The speech recognition subunit is used to acquire a set of sentences for each Chinese speech segment to be recognized based on the first segmentation result, and simultaneously construct an acoustic model and perform acoustic recognition on each Chinese speech in the set of sentences based on the acoustic model.
[0036] The identity determination subunit is used to determine the sound features corresponding to the Chinese speech of adjacent sentences based on the acoustic recognition results, and to compare the sound features corresponding to the Chinese speech of the adjacent sentences.
[0037] The result determination subunit is used to determine that if the comparison results indicate that the sound features corresponding to the Chinese speech of adjacent sentences are consistent, the users corresponding to the Chinese speech of adjacent sentences are the same, and the speech text corresponding to the Chinese speech of adjacent sentences is uniformly labeled; otherwise, the users corresponding to the Chinese speech of adjacent sentences are different, and the speech text corresponding to the Chinese speech of adjacent sentences is distinguished and labeled.
[0038] Preferably, in a deep learning-based Chinese speech recognition system, the correction module includes:
[0039] The text acquisition unit is used to acquire the Chinese speech segment to be recognized, and at the same time, construct a pronunciation change recognition model, and input the Chinese speech segment to be recognized into the pronunciation change recognition model for processing to obtain the intonation information of the Chinese speech segment to be recognized;
[0040] The intent determination unit is used to acquire the speech text obtained after recognizing the Chinese speech segment to be recognized, and to combine the intonation information with the speech text to determine the target intent of the Chinese speech segment to be recognized.
[0041] The semantic determination unit is used to perform semantic analysis on the speech text based on the target intent, obtain the semantic analysis result, and at the same time, obtain the preset Chinese grammar verification rules and perform grammar verification on the speech text based on the semantic analysis result.
[0042] The grammar correction unit is used to determine the target position of the abnormal speech text in the speech text when the grammar check result determines that there is an error in the speech text, and to determine the logical relationship of the context of the abnormal speech text based on the target position.
[0043] The grammar correction unit is used to split the abnormal speech text at the target location into N text keywords, and to reorganize the N text keywords based on the logical relationship and preset Chinese grammar rules to obtain the corrected speech text.
[0044] A text verification unit is used to perform text verification on the corrected speech text based on the target intent, and to determine the differing characters in the speech text based on the verification result, and to determine the target pinyin of the differing characters;
[0045] The text replacement unit is used to map the target pinyin to each preset noun in the preset noun library, and determine the target replacement text based on the mapping result;
[0046] The text replacement unit is further configured to replace the differing text based on the target replacement text, and obtain the final speech recognition text based on the replacement result.
[0047] Preferably, in a deep learning-based Chinese speech recognition system, the correction module includes:
[0048] A speech recognition text acquisition unit is used to acquire the final speech recognition text and determine the text size of the final speech recognition text;
[0049] The capacity allocation unit is used to allocate target storage space in a preset storage area based on the text size, and to store the final speech recognition text in the target storage space.
[0050] This invention provides a Chinese speech recognition method based on deep learning, comprising:
[0051] Step 1: Receive the Chinese speech segments to be recognized in real time and sort them according to the time series.
[0052] Step 2: Construct a Chinese speech recognition model, and based on the sorting results, input the obtained Chinese speech segments to be recognized into the Chinese speech recognition model for speech recognition to obtain the speech text;
[0053] Step 3: Based on the preset Chinese grammar, perform grammatical correction on the obtained speech text to obtain the final speech recognition text.
[0054] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings.
[0055] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0056] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0057] Figure 1 This is a structural diagram of a deep learning-based Chinese speech recognition system according to an embodiment of the present invention;
[0058] Figure 2 This is a structural diagram of a speech acquisition module in a deep learning-based Chinese speech recognition system according to an embodiment of the present invention.
[0059] Figure 3 This is a flowchart of a Chinese speech recognition method based on deep learning in an embodiment of the present invention. Detailed Implementation
[0060] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0061] Example 1:
[0062] This embodiment provides a Chinese speech recognition system based on deep learning, such as... Figure 1 As shown, it includes:
[0063] The speech acquisition module is used to receive Chinese speech segments to be recognized in real time and sort the Chinese speech segments to be recognized based on time series.
[0064] The speech recognition module is used to build a Chinese speech recognition model, and based on the sorting results, it sequentially inputs the obtained Chinese speech segments to be recognized into the Chinese speech recognition model for speech recognition to obtain speech text;
[0065] The correction module is used to correct the grammar of the obtained speech text based on the preset Chinese grammar, so as to obtain the final speech recognition text.
[0066] In this embodiment, the Chinese speech segment to be recognized refers to the set of sentences received for speech recognition, and each sentence is a speech segment.
[0067] In this embodiment, the time series is used to characterize the chronological order in which different Chinese speech segments to be identified occur, that is, the order in which different Chinese speech segments to be identified are spoken.
[0068] In this embodiment, sorting the Chinese speech segments to be identified based on time series means sorting the acquired speech segments to be identified according to their chronological order.
[0069] In this embodiment, speech text refers to the text information obtained after recognizing the received Chinese speech segment to be recognized, that is, the Chinese character information corresponding to the Chinese speech segment to be recognized.
[0070] In this embodiment, the preset Chinese grammar is pre-defined, including limiting the position and logical relationship of the subject and verb.
[0071] The beneficial effects of the above technical solution are: by constructing a Chinese speech recognition model to sequentially recognize the acquired Chinese speech segments to be recognized, and correcting the recognized speech text according to Chinese grammar, the accuracy of Chinese speech recognition is ensured and the effect of Chinese speech recognition is improved.
[0072] Example 2:
[0073] Based on Example 1, this example provides a Chinese speech recognition system based on deep learning, such as... Figure 2 As shown, the voice acquisition module includes:
[0074] A voice acquisition unit is used to monitor the user's current acoustic characteristics in real time and determine the user's current voice state based on the acoustic characteristics, wherein the voice state includes speaking and not speaking;
[0075] The voice recording unit is used to acquire the Chinese voice spoken by the user when the voice state is in the speaking state, and to store the acquired Chinese voice to obtain the Chinese voice segment to be recognized.
[0076] In this embodiment, the acoustic features are used to determine whether the user is currently speaking.
[0077] The beneficial effects of the above technical solution are: by accurately judging the user's current speaking behavior, it enables timely acquisition and storage of the generated Chinese speech when the user speaks, thus facilitating the accurate and effective recognition of the user's Chinese speech.
[0078] Example 3:
[0079] Based on Example 2, this example provides a Chinese speech recognition system based on deep learning, wherein the speech recording unit includes:
[0080] The speech processing subunit is used to acquire the obtained Chinese speech segment to be recognized, and to perform spectral analysis on the Chinese speech segment to be recognized to obtain the audio spectrum corresponding to the Chinese speech segment to be recognized.
[0081] The speech filtering subunit is used to determine the first peak frequency point of the Chinese speech segment to be identified at each time based on the audio spectrum, and at the same time, to obtain the noise audio spectrum corresponding to the noise signal, and to determine the second peak frequency point of the noise signal based on the noise audio spectrum.
[0082] The speech filtering subunit is used to compare the first peak frequency point with the second peak frequency point, filter out target peak frequency points where the first peak frequency point is greater than the second peak frequency point, and determine the Chinese speech segment to be recognized corresponding to the target peak frequency point as a valid Chinese speech segment to be recognized.
[0083] In this embodiment, audio spectrogram refers to converting the Chinese speech segment to be identified into the corresponding audio form, in order to distinguish the valid speech signal from the noise signal in the Chinese speech segment to be identified.
[0084] In this embodiment, the first peak frequency point refers to the audio value of the Chinese speech segment to be identified in each frame in the time domain.
[0085] In this embodiment, the noise audio spectrum refers to the audio format corresponding to various noises.
[0086] In this embodiment, the second peak frequency point refers to the magnitude of the audio value corresponding to the noise.
[0087] In this embodiment, the target peak frequency point refers to the Chinese voice signal where the first peak frequency point is greater than the second peak frequency point.
[0088] In this embodiment, the effective Chinese speech segment to be identified refers to the speech signal obtained without other interference factors after removing the noise signal in the Chinese speech segment to be identified.
[0089] The beneficial effects of the above technical solution are: by converting the acquired Chinese speech segment to be recognized into the corresponding audio spectrum, and at the same time acquiring the audio spectrum corresponding to the noise signal, the noise signal in the Chinese speech segment to be recognized can be removed through the audio spectrum, thereby ensuring the effectiveness of the Chinese speech segment to be recognized and improving the recognition effect of the Chinese speech segment to be recognized.
[0090] Example 4:
[0091] Based on Example 1, this example provides a Chinese speech recognition system based on deep learning, wherein the speech acquisition module includes:
[0092] The time determination unit is used to acquire the obtained Chinese speech segment to be identified and process the Chinese speech segment to be identified to obtain the speech signal corresponding to each frame;
[0093] The time determination unit is further configured to determine the time domain information of the Chinese speech segment to be identified based on the speech signal corresponding to each frame, and match the time domain information with the speech signal corresponding to each frame;
[0094] The sorting unit is used to determine the time series corresponding to the Chinese speech segment to be identified based on the matching result, and to sort the Chinese speech segment to be identified in ascending order of the time series, wherein the Chinese speech segment to be identified is at least one segment.
[0095] In this embodiment, time-domain information refers to the time range involved in the received Chinese speech segment to be recognized.
[0096] In this embodiment, the time series is used to characterize the specific time corresponding to each Chinese speech segment to be identified.
[0097] The beneficial effects of the above technical solution are: by determining the time domain information involved in the Chinese speech segment to be identified, the specific time sequence of each Chinese speech segment to be identified can be confirmed, thereby facilitating the sorting of the obtained Chinese speech segments to be identified by the specific time sequence, improving the recognition efficiency of the Chinese speech segments to be identified, and ensuring the recognition effect of the Chinese speech segments to be identified.
[0098] Example 5:
[0099] Based on Example 4, this example provides a Chinese speech recognition system based on deep learning, wherein the sorting unit includes:
[0100] The result acquisition subunit is used to acquire the sorting result of the Chinese speech segments to be recognized, and to determine the target number of the Chinese speech segments to be recognized based on the sorting result;
[0101] The tag acquisition subunit is used to extract the acoustic features of the Chinese speech segment to be identified and determine the speech type of the Chinese speech segment to be identified based on the acoustic features.
[0102] The tagging subunit is used to obtain a target number of tagging tags from a preset tag database based on the speech type, and to tag the Chinese speech segment to be recognized based on the target number of tagging tags.
[0103] In this embodiment, the target number is used to characterize the specific number of Chinese speech segments to be identified.
[0104] In this embodiment, acoustic features refer to the sound characteristics of the Chinese speech segment to be identified, including timbre and intonation.
[0105] In this embodiment, the preset tag database is pre-set and used to store the tagging labels corresponding to different speech types.
[0106] In this embodiment, the tag refers to a marker symbol that can be used to distinguish different Chinese speech segments to be recognized. The tag can quickly distinguish different Chinese speech segments to be recognized, and also makes it easier to determine the speech type of the Chinese speech segment to be recognized.
[0107] The beneficial effects of the above technical solution are: by determining the acoustic features of the Chinese speech segment to be identified, and by accurately and effectively judging the speech type of the speech segment to be identified based on the acoustic features, it is easy to select appropriate labels to mark different Chinese speech segments to be identified according to the speech type, thus ensuring the orderliness of the recognition of the Chinese speech segment to be identified, and also improving the recognition efficiency and accuracy.
[0108] Example 6:
[0109] Based on Example 1, this example provides a Chinese speech recognition system based on deep learning, wherein the speech recognition module includes:
[0110] The data acquisition unit is used to acquire speech training text and read the speech training text from a preset speech library using different accents to obtain audio data of the speech training text with different accents.
[0111] A data processing unit is configured to preprocess the audio data, convert the audio data into a corresponding spectrogram based on the preprocessing result, and determine the effective region in the audio data based on the spectrogram.
[0112] The model building unit is used to determine the feature parameters of the audio data based on the effective region, and at the same time, obtain the correspondence between Chinese Pinyin and Chinese characters, train the feature parameters based on the correspondence, and build a Chinese speech recognition model based on the training results.
[0113] The speech recognition unit is used to sequentially input the acquired Chinese speech segments to be recognized into the Chinese speech recognition model, and analyze the received Chinese speech segments to be recognized based on the preset syntax analysis tree in the Chinese speech recognition model to determine the start and end points of each sentence in the Chinese speech segments to be recognized.
[0114] The speech recognition unit is used to perform a first segmentation on each of the Chinese speech segments to be recognized based on the start point and the end point, and to obtain a set of sentences for each of the Chinese speech segments to be recognized based on the first segmentation result, and to extract the syllable attributes contained in each Chinese speech sentence in the set of sentences.
[0115] The speech recognition unit is used to perform a second segmentation on each Chinese speech based on the syllable attributes, and to obtain the Chinese words contained in each Chinese speech based on the second segmentation result.
[0116] The speech recognition unit is further configured to extract the pronunciation features of the Chinese words, and process the pronunciation features based on the correspondence between the Chinese pinyin and Chinese characters to obtain the word text corresponding to the Chinese words;
[0117] The text splicing unit is used to splice the text corresponding to the Chinese words contained in each Chinese speech sentence to obtain the speech text corresponding to the Chinese speech segment to be identified.
[0118] In this embodiment, the preset speech library is pre-set to store different accents, thereby facilitating accurate and effective training of the Chinese speech recognition model.
[0119] In this embodiment, the voice training text is pre-set, and the text information corresponding to the voice is known.
[0120] In this embodiment, preprocessing may involve noise reduction or other processing of the audio data.
[0121] In this embodiment, the spectrogram refers to a spectrum analysis view, with time on the horizontal axis, frequency on the vertical axis, and the coordinate point value being the energy of the speech data.
[0122] In this embodiment, the effective region refers to the audio data that has been filtered to extract speech segments with key representational information.
[0123] In this embodiment, the feature parameters refer to the values of the audio data and the fluctuation range of the tone.
[0124] In this embodiment, the preset syntax analysis tree is pre-set and used to identify the acquired Chinese speech segments to be identified according to Chinese grammar, thereby improving the accuracy and efficiency of recognition.
[0125] In this embodiment, the start point and the end point are used to characterize the beginning and end of each sentence.
[0126] In this embodiment, the first splitting refers to splitting the Chinese speech segment to be identified into multiple sentences, with each sentence as a unit.
[0127] In this embodiment, the sentence set refers to the set obtained after splitting the Chinese speech segment to be recognized into multiple sentences.
[0128] In this embodiment, the syllable attribute refers to whether the words contained in each sentence are monosyllabic or disyllabic.
[0129] In this embodiment, the second splitting refers to splitting each Chinese speech into multiple Chinese words, with each word as a unit.
[0130] In this embodiment, pronunciation features refer to the pronunciation characteristics of each Chinese word.
[0131] In this embodiment, the vocabulary text refers to the Chinese characters corresponding to the vocabulary sounds in each Chinese speech sentence.
[0132] In this embodiment, the received Chinese speech segment to be recognized is analyzed based on the preset syntax analysis tree in the Chinese speech recognition model, including:
[0133] The obtained Chinese speech segment to be identified is obtained, and the Chinese speech segment to be identified is converted into a corresponding feature vector, and the feature sequence corresponding to the Chinese speech segment to be identified is determined based on the feature vector;
[0134] The word sequence recognized by the Chinese speech recognition model for the Chinese speech segment to be recognized is calculated based on the feature sequence, and the recognition accuracy of the Chinese speech segment to be recognized is calculated based on the word sequence. The specific steps include:
[0135] The word sequence identified for the Chinese speech segment to be identified is calculated according to the following formula:
[0136] M=argmax[log2 P(α|m)+η*log2 P(m)];
[0137] Where M represents the word sequence identified for the Chinese speech segment to be identified; P(α|m) represents the acoustic model, which represents the probability that the output acoustic feature is the feature sequence α when the preset word sequence is m, and the value range is (0, 1); P(m) represents the language model, which represents the probability value of the preset word sequence m appearing in the feature sequence, and the value range is (0, 1); η represents an adjustable parameter, and the value range is (0, 1); argmax[·] represents the function that calculates the set of functions, specifically representing the maximum vocabulary set obtained when the conditions for the acoustic model and the language model to identify the Chinese speech segment to be identified are met;
[0138] The total number of words K identified for the Chinese speech segment to be identified is determined based on the word sequence M;
[0139] The recognition accuracy of the Chinese speech segment to be recognized is calculated using the following formula:
[0140]
[0141] in, ω represents the recognition accuracy of the Chinese speech segment to be recognized, and its value ranges from (0, 1); ω represents the error factor, and its value ranges from (0.02, 0.05); K represents the total number of words recognized in the Chinese speech segment to be recognized; δ represents the number of words incorrectly recognized in the Chinese speech segment to be recognized; σ represents the number of words that were missed in the Chinese speech segment to be recognized.
[0142] Compare the calculated recognition accuracy with the preset accuracy.
[0143] If the recognition accuracy is greater than or equal to the preset accuracy, the recognition of the Chinese speech segment to be recognized is deemed qualified.
[0144] Otherwise, the recognition of the Chinese speech segment to be recognized is deemed unqualified, and the speech recognition of the Chinese speech segment to be recognized is re-performed until the recognition accuracy is greater than or equal to the preset accuracy.
[0145] The aforementioned feature vector refers to the sentence vector obtained after vectorizing the Chinese speech segment to be recognized.
[0146] The aforementioned feature sequence refers to the vocabulary sequence corresponding to each feature vector.
[0147] The aforementioned preset accuracy rate is pre-set and is used to measure whether the recognition accuracy of the Chinese speech segment to be recognized meets the preset requirements.
[0148] The beneficial effects of the above technical solution are as follows: by acquiring speech training text and reading it through different accents, the system can effectively acquire different accents. Secondly, by processing and training the audio data corresponding to the speech training text, the system can accurately and reliably construct a Chinese speech recognition model. Finally, the system can segment and recognize the acquired Chinese speech segments to be recognized through the Chinese speech recognition model, thereby ensuring the accuracy and effectiveness of the recognition of the Chinese speech segments to be recognized and ensuring that the obtained speech text is accurate and effective.
[0149] Example 7:
[0150] Based on Example 6, this example provides a Chinese speech recognition system based on deep learning, wherein the speech recognition unit includes:
[0151] The speech recognition subunit is used to acquire a set of sentences for each Chinese speech segment to be recognized based on the first segmentation result, and simultaneously construct an acoustic model and perform acoustic recognition on each Chinese speech in the set of sentences based on the acoustic model.
[0152] The identity determination subunit is used to determine the sound features corresponding to the Chinese speech of adjacent sentences based on the acoustic recognition results, and to compare the sound features corresponding to the Chinese speech of the adjacent sentences.
[0153] The result determination subunit is used to determine that if the comparison results indicate that the sound features corresponding to the Chinese speech of adjacent sentences are consistent, the users corresponding to the Chinese speech of adjacent sentences are the same, and the speech text corresponding to the Chinese speech of adjacent sentences is uniformly labeled; otherwise, the users corresponding to the Chinese speech of adjacent sentences are different, and the speech text corresponding to the Chinese speech of adjacent sentences is distinguished and labeled.
[0154] In this embodiment, the acoustic model is used to analyze the sound characteristics of the Chinese speech segment to be identified, including timbre and tone.
[0155] In this embodiment, the sound features may be the thickness or other characteristics of the Chinese speech corresponding to adjacent sentences.
[0156] In this embodiment, uniform annotation refers to labeling the Chinese speech of adjacent sentences as belonging to the same user, thereby facilitating the differentiation of the obtained speech text.
[0157] In this embodiment, the distinguishing annotation refers to labeling the Chinese speech of adjacent sentences as different users, thereby facilitating the differentiation of the obtained speech text.
[0158] The beneficial effects of the above technical solution are: by constructing an acoustic model and using the acoustic model to identify the acoustic features of adjacent sentences in the Chinese speech segment to be identified, it is easier to accurately determine the users corresponding to different sentences in the Chinese speech segment to be identified, thereby facilitating the orderly management of the identified speech text and improving the recognition effect of the Chinese speech segment to be identified.
[0159] Example 8:
[0160] Based on Example 1, this example provides a Chinese speech recognition system based on deep learning, wherein the correction module includes:
[0161] The text acquisition unit is used to acquire the Chinese speech segment to be recognized, and at the same time, construct a pronunciation change recognition model, and input the Chinese speech segment to be recognized into the pronunciation change recognition model for processing to obtain the intonation information of the Chinese speech segment to be recognized;
[0162] The intent determination unit is used to acquire the speech text obtained after recognizing the Chinese speech segment to be recognized, and to combine the intonation information with the speech text to determine the target intent of the Chinese speech segment to be recognized.
[0163] The semantic determination unit is used to perform semantic analysis on the speech text based on the target intent, obtain the semantic analysis result, and at the same time, obtain the preset Chinese grammar verification rules and perform grammar verification on the speech text based on the semantic analysis result.
[0164] The grammar correction unit is used to determine the target position of the abnormal speech text in the speech text when the grammar check result determines that there is an error in the speech text, and to determine the logical relationship of the context of the abnormal speech text based on the target position.
[0165] The grammar correction unit is used to split the abnormal speech text at the target location into N text keywords, and to reorganize the N text keywords based on the logical relationship and preset Chinese grammar rules to obtain the corrected speech text.
[0166] A text verification unit is used to perform text verification on the corrected speech text based on the target intent, and to determine the differing characters in the speech text based on the verification result, and to determine the target pinyin of the differing characters;
[0167] The text replacement unit is used to map the target pinyin to each preset noun in the preset noun library, and determine the target replacement text based on the mapping result;
[0168] The text replacement unit is further configured to replace the differing text based on the target replacement text, and obtain the final speech recognition text based on the replacement result.
[0169] In this embodiment, the pronunciation change recognition model is used to identify the pronunciation changes in the Chinese speech segment to be recognized.
[0170] In this embodiment, intonation information refers to the tone changes of the Chinese speech segment to be identified, which helps to determine the user's speech intent.
[0171] In this embodiment, the target intent refers to the expressive purpose corresponding to the Chinese speech segment to be identified.
[0172] In this embodiment, semantic analysis refers to analyzing the acquired speech text to determine the meaning expressed by the speech text.
[0173] In this embodiment, the preset Chinese grammar verification rules are pre-set and used to verify the grammar of the Chinese speech segment to be recognized.
[0174] In this embodiment, abnormal speech text refers to the text information corresponding to grammatical errors in the acquired speech text.
[0175] In this embodiment, the target location refers to the position of the abnormal speech text in the obtained speech text.
[0176] In this embodiment, text keywords refer to the Chinese words contained in the sentence containing the abnormal voice text.
[0177] In this embodiment, the preset Chinese grammar rules are pre-defined.
[0178] In this embodiment, text verification refers to verifying the text in the obtained speech text in order to determine whether there are any erroneous characters.
[0179] In this embodiment, the difference text refers to the erroneous text present in the obtained speech text.
[0180] In this embodiment, the target pinyin refers to the pronunciation of the different characters.
[0181] In this embodiment, the preset terminology database is pre-set and used to store different words.
[0182] In this embodiment, the target replacement text refers to a Chinese character that is homophonous with the different text but has a different font.
[0183] The beneficial effects of the above technical solution are: by performing grammatical verification on the obtained speech text, the grammatical errors in the speech text can be corrected in a timely manner when they exist; secondly, after grammatical correction, the form of Chinese characters in the speech text is verified, which ensures the accuracy of the final speech recognition text and guarantees the recognition effect of the Chinese speech segments to be recognized.
[0184] Example 9:
[0185] Based on Example 1, this example provides a Chinese speech recognition system based on deep learning, wherein the correction module includes:
[0186] A speech recognition text acquisition unit is used to acquire the final speech recognition text and determine the text size of the final speech recognition text;
[0187] The capacity allocation unit is used to allocate target storage space in a preset storage area based on the text size, and to store the final speech recognition text in the target storage space.
[0188] In this embodiment, the preset storage area is pre-defined and includes storage spaces of different sizes.
[0189] In this embodiment, the target storage space refers to the storage area used to store the final acquired speech recognition text.
[0190] The beneficial effects of the above technical solution are: by determining the file size of the final obtained speech recognition text, it is easier to allocate corresponding storage space for the speech recognition text and store the speech recognition text, thereby improving the preservation effect of the recognition results of the Chinese speech segments to be recognized, and thus ensuring the recognition effect of the Chinese speech segments to be recognized.
[0191] Example 10:
[0192] This embodiment provides a Chinese speech recognition method based on deep learning, such as... Figure 3 As shown, it includes:
[0193] Step 1: Receive the Chinese speech segments to be recognized in real time and sort them according to the time series.
[0194] Step 2: Construct a Chinese speech recognition model, and based on the sorting results, input the obtained Chinese speech segments to be recognized into the Chinese speech recognition model for speech recognition to obtain the speech text;
[0195] Step 3: Based on the preset Chinese grammar, perform grammatical correction on the obtained speech text to obtain the final speech recognition text.
[0196] The beneficial effects of the above technical solution are: by constructing a Chinese speech recognition model to sequentially recognize the acquired Chinese speech segments to be recognized, and correcting the recognized speech text according to Chinese grammar, the accuracy of Chinese speech recognition is ensured and the effect of Chinese speech recognition is improved.
[0197] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A Chinese speech recognition system based on deep learning, characterized in that, include: The speech acquisition module is used to receive Chinese speech segments to be recognized in real time and sort the Chinese speech segments to be recognized based on time series. The speech recognition module is used to build a Chinese speech recognition model, and based on the sorting results, it sequentially inputs the obtained Chinese speech segments to be recognized into the Chinese speech recognition model for speech recognition to obtain speech text; The correction module is used to perform grammatical correction on the obtained speech text based on preset Chinese grammar to obtain the final speech recognition text; The speech recognition module includes: The data acquisition unit is used to acquire speech training text and read the speech training text from a preset speech library using different accents to obtain audio data of the speech training text with different accents. A data processing unit is configured to preprocess the audio data, convert the audio data into a corresponding spectrogram based on the preprocessing result, and determine the effective region in the audio data based on the spectrogram. The model building unit is used to determine the feature parameters of the audio data based on the effective region, and at the same time, obtain the correspondence between Chinese Pinyin and Chinese characters, train the feature parameters based on the correspondence, and build a Chinese speech recognition model based on the training results. The speech recognition unit is used to sequentially input the acquired Chinese speech segments to be recognized into the Chinese speech recognition model, and analyze the received Chinese speech segments to be recognized based on the preset syntax analysis tree in the Chinese speech recognition model to determine the start and end points of each sentence in the Chinese speech segments to be recognized. The speech recognition unit is used to perform a first segmentation on each of the Chinese speech segments to be recognized based on the start point and the end point, and to obtain a set of sentences for each of the Chinese speech segments to be recognized based on the first segmentation result, and to extract the syllable attributes contained in each Chinese speech sentence in the set of sentences. The speech recognition unit is used to perform a second segmentation on each Chinese speech based on the syllable attributes, and to obtain the Chinese words contained in each Chinese speech based on the second segmentation result. The speech recognition unit is further configured to extract the pronunciation features of the Chinese words, and process the pronunciation features based on the correspondence between the Chinese pinyin and Chinese characters to obtain the word text corresponding to the Chinese words; The text splicing unit is used to splice the word text corresponding to the Chinese words contained in each Chinese speech sentence to obtain the speech text corresponding to the Chinese speech segment to be identified. The received Chinese speech segment to be recognized is analyzed based on the preset syntax tree in the Chinese speech recognition model, including: The obtained Chinese speech segment to be identified is obtained, and the Chinese speech segment to be identified is converted into a corresponding feature vector, and the feature sequence corresponding to the Chinese speech segment to be identified is determined based on the feature vector; The word sequence recognized by the Chinese speech recognition model for the Chinese speech segment to be recognized is calculated based on the feature sequence, and the recognition accuracy of the Chinese speech segment to be recognized is calculated based on the word sequence. The specific steps include: The word sequence identified for the Chinese speech segment to be identified is calculated according to the following formula: ; in, This represents the word sequence identified from the Chinese speech segment to be identified; The acoustic model represents the predefined word sequence as... In this case, the output acoustic features are a feature sequence. The probability of , and the value range is (0, 1); Representational language model, representing a predefined word sequence The probability value of a feature sequence, and the range of the value is (0, 1). This indicates an adjustable parameter with a value range of (0, 1). This represents the function that takes the set of functions, specifically the maximum set of words obtained when the acoustic model and language model meet the conditions for recognizing the Chinese speech segment to be recognized; Based on the word sequence Determine the total number of words identified in the Chinese speech segment to be identified. ; The recognition accuracy of the Chinese speech segment to be recognized is calculated using the following formula: ; in, This represents the recognition accuracy of the Chinese speech segment to be recognized, and its value ranges from (0, 1). This represents the error factor, and its value ranges from (0.02 to 0.05). This represents the total number of words identified in the Chinese speech segment to be identified. ; This indicates the number of words incorrectly identified in the Chinese speech segment to be identified; This indicates the number of words that were missed in the Chinese speech segment to be identified; Compare the calculated recognition accuracy with the preset accuracy. If the recognition accuracy is greater than or equal to the preset accuracy, the recognition of the Chinese speech segment to be recognized is deemed qualified. Otherwise, the recognition of the Chinese speech segment to be recognized is deemed unqualified, and the speech recognition of the Chinese speech segment to be recognized is re-performed until the recognition accuracy is greater than or equal to the preset accuracy.
2. The Chinese speech recognition system based on deep learning according to claim 1, characterized in that, The voice acquisition module includes: A voice acquisition unit is used to monitor the user's current acoustic characteristics in real time and determine the user's current voice state based on the acoustic characteristics, wherein the voice state includes speaking and not speaking; The voice recording unit is used to acquire the Chinese voice spoken by the user when the voice state is in the speaking state, and to store the acquired Chinese voice to obtain the Chinese voice segment to be recognized.
3. The Chinese speech recognition system based on deep learning according to claim 2, characterized in that, The voice recording unit includes: The speech processing subunit is used to acquire the obtained Chinese speech segment to be recognized, and to perform spectral analysis on the Chinese speech segment to be recognized to obtain the audio spectrum corresponding to the Chinese speech segment to be recognized. The speech filtering subunit is used to determine the first peak frequency point of the Chinese speech segment to be identified at each time based on the audio spectrum, and at the same time, to obtain the noise audio spectrum corresponding to the noise signal, and to determine the second peak frequency point of the noise signal based on the noise audio spectrum. The speech filtering subunit is used to compare the first peak frequency point with the second peak frequency point, filter out target peak frequency points where the first peak frequency point is greater than the second peak frequency point, and determine the Chinese speech segment to be recognized corresponding to the target peak frequency point as a valid Chinese speech segment to be recognized.
4. The Chinese speech recognition system based on deep learning according to claim 1, characterized in that, The voice acquisition module includes: The time determination unit is used to acquire the obtained Chinese speech segment to be identified and process the Chinese speech segment to be identified to obtain the speech signal corresponding to each frame; The time determination unit is further configured to determine the time domain information of the Chinese speech segment to be identified based on the speech signal corresponding to each frame, and match the time domain information with the speech signal corresponding to each frame; The sorting unit is used to determine the time series corresponding to the Chinese speech segment to be identified based on the matching result, and to sort the Chinese speech segment to be identified in ascending order of the time series, wherein the Chinese speech segment to be identified is at least one segment.
5. A Chinese speech recognition system based on deep learning according to claim 4, characterized in that, The sorting unit includes: The result acquisition subunit is used to acquire the sorting result of the Chinese speech segments to be recognized, and to determine the target number of the Chinese speech segments to be recognized based on the sorting result; The tag acquisition subunit is used to extract the acoustic features of the Chinese speech segment to be identified and determine the speech type of the Chinese speech segment to be identified based on the acoustic features. The tagging subunit is used to obtain a target number of tagging tags from a preset tag database based on the speech type, and to tag the Chinese speech segment to be recognized based on the target number of tagging tags.
6. The Chinese speech recognition system based on deep learning according to claim 1, characterized in that, The speech recognition unit includes: The speech recognition subunit is used to acquire a set of sentences for each Chinese speech segment to be recognized based on the first segmentation result, and simultaneously construct an acoustic model and perform acoustic recognition on each Chinese speech in the set of sentences based on the acoustic model. The identity determination subunit is used to determine the sound features corresponding to the Chinese speech of adjacent sentences based on the acoustic recognition results, and to compare the sound features corresponding to the Chinese speech of the adjacent sentences. The result determination subunit is used to determine that if the comparison results indicate that the sound features corresponding to the Chinese speech of adjacent sentences are consistent, the users corresponding to the Chinese speech of adjacent sentences are the same, and the speech text corresponding to the Chinese speech of adjacent sentences is uniformly labeled; otherwise, the users corresponding to the Chinese speech of adjacent sentences are different, and the speech text corresponding to the Chinese speech of adjacent sentences is distinguished and labeled.
7. A Chinese speech recognition system based on deep learning according to claim 1, characterized in that, The correction module includes: The text acquisition unit is used to acquire the Chinese speech segment to be recognized, and at the same time, construct a pronunciation change recognition model, and input the Chinese speech segment to be recognized into the pronunciation change recognition model for processing to obtain the intonation information of the Chinese speech segment to be recognized; The intent determination unit is used to acquire the speech text obtained after recognizing the Chinese speech segment to be recognized, and to combine the intonation information with the speech text to determine the target intent of the Chinese speech segment to be recognized. The semantic determination unit is used to perform semantic analysis on the speech text based on the target intent, obtain the semantic analysis result, and at the same time, obtain the preset Chinese grammar verification rules and perform grammar verification on the speech text based on the semantic analysis result. The grammar correction unit is used to determine the target position of the abnormal speech text in the speech text when the grammar check result determines that there is an error in the speech text, and to determine the logical relationship of the context of the abnormal speech text based on the target position. The grammar correction unit is used to split the abnormal speech text at the target location into N text keywords, and to reorganize the N text keywords based on the logical relationship and preset Chinese grammar rules to obtain the corrected speech text. A text verification unit is used to perform text verification on the corrected speech text based on the target intent, and to determine the differing characters in the speech text based on the verification result, and to determine the target pinyin of the differing characters; The text replacement unit is used to map the target pinyin to each preset noun in the preset noun library, and determine the target replacement text based on the mapping result; The text replacement unit is further configured to replace the differing text based on the target replacement text, and obtain the final speech recognition text based on the replacement result.
8. A Chinese speech recognition system based on deep learning according to claim 1, characterized in that, The correction module includes: A speech recognition text acquisition unit is used to acquire the final speech recognition text and determine the text size of the final speech recognition text; The capacity allocation unit is used to allocate target storage space in a preset storage area based on the text size, and to store the final speech recognition text in the target storage space.
9. A Chinese speech recognition method based on deep learning, characterized in that, include: Step 1: Receive the Chinese speech segments to be recognized in real time and sort them according to the time series. Step 2: Construct a Chinese speech recognition model, and based on the sorting results, input the obtained Chinese speech segments to be recognized into the Chinese speech recognition model for speech recognition to obtain the speech text; Step 3: Based on the preset Chinese grammar, perform grammatical correction on the obtained speech text to obtain the final speech recognition text; Step 2 includes: Acquire speech training text, and read the speech training text with different accents from a preset speech library to obtain audio data of the speech training text with different accents; The audio data is preprocessed, and the audio data is converted into a corresponding spectrogram based on the preprocessing result. The effective regions in the audio data are determined based on the spectrogram. Based on the effective region, the feature parameters of the audio data are determined. At the same time, the correspondence between Chinese Pinyin and Chinese characters is obtained, and the feature parameters are trained based on the correspondence. A Chinese speech recognition model is constructed based on the training results. The acquired Chinese speech segments to be recognized are sequentially input into the Chinese speech recognition model, and the received Chinese speech segments to be recognized are analyzed based on the preset syntax analysis tree in the Chinese speech recognition model to determine the start and end points of each sentence in the Chinese speech segments to be recognized. Each Chinese speech segment to be identified is firstly split based on the starting point and ending point, and a set of sentences for each Chinese speech segment to be identified is obtained based on the first split result. The syllable attributes contained in each Chinese speech sentence in the set of sentences are then extracted. Based on the syllable attributes, each Chinese speech sentence is split into two parts, and the Chinese words contained in each Chinese speech sentence are obtained based on the results of the second split. The speech recognition unit is also used to extract the pronunciation features of the Chinese words, and process the pronunciation features based on the correspondence between the Chinese pinyin and Chinese characters to obtain the word text corresponding to the Chinese words; The vocabulary text corresponding to the Chinese words contained in each Chinese speech sentence is concatenated to obtain the speech text corresponding to the Chinese speech segment to be identified. The received Chinese speech segment to be recognized is analyzed based on the preset syntax tree in the Chinese speech recognition model, including: The obtained Chinese speech segment to be identified is obtained, and the Chinese speech segment to be identified is converted into a corresponding feature vector, and the feature sequence corresponding to the Chinese speech segment to be identified is determined based on the feature vector; The word sequence recognized by the Chinese speech recognition model for the Chinese speech segment to be recognized is calculated based on the feature sequence, and the recognition accuracy of the Chinese speech segment to be recognized is calculated based on the word sequence. The specific steps include: The word sequence identified for the Chinese speech segment to be identified is calculated according to the following formula: ; in, This represents the word sequence identified from the Chinese speech segment to be identified; The acoustic model represents the sequence of words in the preset word sequence. In this case, the output acoustic features are a feature sequence. The probability of , and the value range is (0, 1); Representational language model, representing a predefined word sequence The probability value of a feature sequence, and the range of the value is (0, 1). This indicates an adjustable parameter with a value range of (0, 1). This represents the function that takes the set of functions, specifically the maximum set of words obtained when the acoustic model and language model meet the conditions for recognizing the Chinese speech segment to be recognized; Based on the word sequence Determine the total number of words identified in the Chinese speech segment to be identified. ; The recognition accuracy of the Chinese speech segment to be recognized is calculated using the following formula: ; in, This represents the recognition accuracy of the Chinese speech segment to be recognized, and its value ranges from (0, 1). This represents the error factor, and its value ranges from (0.02 to 0.05). This represents the total number of words identified in the Chinese speech segment to be identified. ; This indicates the number of words incorrectly identified in the Chinese speech segment to be identified; This indicates the number of words that were missed in the Chinese speech segment to be identified; Compare the calculated recognition accuracy with the preset accuracy. If the recognition accuracy is greater than or equal to the preset accuracy, the recognition of the Chinese speech segment to be recognized is deemed qualified. Otherwise, the recognition of the Chinese speech segment to be recognized is deemed unqualified, and the speech recognition of the Chinese speech segment to be recognized is re-performed until the recognition accuracy is greater than or equal to the preset accuracy.
Citation Information
Patent Citations
A speech recognition method
CN109215637A
Voice activity detection method and device
CN110689877A
Speech signal screening method, device, audio equipment and system
CN111883183A
Natural language processing system based on artificial intelligence
CN112164403A