Voice recognition method, device, electronic device and storage medium
By determining the voice confidence of a speech segment in a speech recognition device and using high-confidence text information to correct low-confidence text information, the problem of low accuracy in non-standard speech recognition is solved, and higher speech recognition accuracy and universal applicability are achieved.
Patent Information
- Application Number
- CN202211199117.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-29
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2042-09-29
AI Technical Summary
Existing speech recognition devices have a low recognition accuracy when recognizing non-standardized speech, and there is a lack of effective methods to improve it.
By obtaining the voice segments in the original voice information, the voice confidence of the sound object corresponding to each voice segment is determined, and the text information with high voice confidence is used to correct the text information with low voice confidence, thereby improving the accuracy of voice recognition.
It improves the accuracy of speech recognition without affecting the recognition accuracy of standard speech. It is universal and does not require the preparation of a large amount of accented data training model in advance.
Smart Images

Figure CN115641849B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech processing technology, and in particular to a speech recognition method, device, electronic device, and computer-readable storage medium. Background Art
[0002] Currently, speech recognition devices are widely used in various locations and situations due to their ease of use. Most speech recognition devices use models trained on a standardized language (such as Mandarin) to recognize other speech, including standardized speech and non-standardized speech (such as regional dialects). Current speech recognition devices typically implement speech recognition based on speech recognition engines. However, because people from different regions have different accents, using a speech recognition engine based on standardized speech to recognize speakers with non-standardized speech often results in low recognition accuracy. Furthermore, there is currently no effective solution for improving the recognition accuracy of standardized speech recognition engines. Summary of the Invention
[0003] In view of this, the main purpose of this application is to provide a speech recognition method, device, electronic device and computer-readable storage medium, which can improve the speech recognition accuracy of the speech recognition engine.
[0004] To achieve the above objectives, the technical solution of this application is implemented as follows:
[0005] In a first aspect, an embodiment of the present application provides a speech recognition method, comprising:
[0006] Acquire at least two speech segments belonging to different sounding objects contained in the original speech information;
[0007] Determining a speech confidence of a sounding object corresponding to each of the at least two speech segments;
[0008] Identifying at least two text messages corresponding to the at least two speech segments, and using the text information of the target speech confidence corresponding to the sound object to correct the text information corresponding to the sound object of other speech confidences, wherein the target speech confidence is higher than the other speech confidences;
[0009] The at least two pieces of text information after correction are output.
[0010] In the above solution, the method of using the text information of the target speech confidence corresponding to the sound object to correct the text information corresponding to other speech confidence corresponding to the sound object includes:
[0011] Determine at least two words in the at least two text messages that belong to different sound objects and meet the target similarity, and use the words in the text information corresponding to the sound object corresponding to the target speech confidence to correct the words in the text information corresponding to the sound object corresponding to other speech confidences.
[0012] In the above solution, determining at least two words in the at least two text messages that belong to different utterance objects and meet the target similarity includes:
[0013] Performing word segmentation processing on each of the at least two text messages to obtain key words contained in each text message;
[0014] Determining a first similarity between the key words contained in different text information;
[0015] At least two keywords whose first similarity meets the target similarity are obtained from the key words contained in each of the at least two text messages; the at least two keywords are at least two words in the at least two text messages that belong to different voice objects and meet the target similarity.
[0016] In the above solution, the original voice information includes at least one sub-voice information collected sequentially according to the first time interval, and the at least two voice segments belonging to different sounding objects included in the original voice information are obtained, including:
[0017] Preprocessing any first sub-speech information among the at least one collected sub-speech information to obtain second sub-speech information represented by speech features of the first sub-speech information;
[0018] performing segmentation and clustering processing on the second sub-speech information to obtain at least two sub-speech segments belonging to different sounding objects contained in the first sub-speech information;
[0019] The sub-speech segments belonging to the same sounding object in each first sub-speech information are clustered to obtain at least two speech segments belonging to different sounding objects contained in the original speech information.
[0020] In the above solution, the segmentation and clustering of the second sub-speech information to obtain at least two sub-speech segments belonging to different sounding objects contained in the first sub-speech information includes:
[0021] Segmenting and clustering the second sub-speech information to obtain at least two intermediate speech segments belonging to different sounding objects;
[0022] In the case that at least one of the intermediate speech segments contains at least two different sound objects, each of the at least two intermediate speech segments is re-segmented and clustered until each of the at least two sub-speech segments belonging to different sound objects obtained contains only one sound object.
[0023] In the above solution, determining the speech confidence of the sounding object corresponding to each of the at least two speech segments includes:
[0024] Performing sentence division on any first speech segment of the at least two speech segments to obtain a plurality of sentences contained in the first speech segment;
[0025] Recognize any first sentence among the multiple sentences to obtain a sentence confidence score corresponding to the first sentence;
[0026] The speech confidence corresponding to the first speech segment is determined based on each of the sentence confidences.
[0027] In the above solution, the step of identifying any first sentence among the multiple sentences and obtaining a sentence confidence score corresponding to the first sentence includes:
[0028] Determine the type of the first statement; the first statement is any one of the multiple statements;
[0029] Inputting the first sentence into a recognition model corresponding to the type to obtain a probability that the first sentence is correctly recognized;
[0030] A confidence level of a statement corresponding to the first statement is obtained based on the probability.
[0031] In a second aspect, an embodiment of the present application further provides a speech recognition device, comprising:
[0032] an acquisition module, a determination module, a correction module and an output module, wherein;
[0033] The acquisition module is used to acquire at least two speech segments belonging to different sounding objects contained in the original speech information;
[0034] The determining module is used to determine the voice confidence of the sound object corresponding to each voice segment of the at least two voice segments;
[0035] The correction module is configured to identify at least two text messages corresponding to the at least two speech segments, and use the text information of the sound object corresponding to the target speech confidence to correct the text information corresponding to the sound object corresponding to other speech confidences, wherein the target speech confidence is higher than the other speech confidences;
[0036] The output module is used to output the at least two corrected text messages.
[0037] In a third aspect, an embodiment of the present application further provides an electronic device, including:
[0038] The memory is used to store executable instructions; the processor is used to implement the above-mentioned speech recognition method when executing the executable instructions stored in the memory.
[0039] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned speech recognition method are implemented.
[0040] The embodiments of the present application provide a speech recognition method, device, electronic device, and computer-readable storage medium. The method comprises: obtaining at least two speech segments belonging to different sound-producing objects contained in the original speech information, determining the speech confidence of the sound-producing object corresponding to each of the at least two speech segments, identifying at least two text information corresponding to the at least two speech segments, and using the text information of the sound-producing object with high speech confidence to correct the text information of the sound-producing object with low speech confidence. The speech recognition method provided by the present application, by identifying the speech segments corresponding to each sound-producing object in the original speech information, and determining the speech confidence corresponding to each sound-producing object, uses the text information corresponding to the speech segments with high speech confidence to correct the text information corresponding to other speech segments. In this way, in the speech recognition process, the accuracy of speech recognition can be improved by using certain speech segments in the original speech information obtained in real time to correct other speech segments, and the training time of the standard speech recognition engine obtained through a large amount of known data is saved, thereby improving the efficiency of speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 This is a schematic diagram of an application scenario of the speech recognition method provided in an embodiment of the present application;
[0042] Figure 2A-2F This is an optional flowchart of the speech recognition method provided in the embodiment of the present application;
[0043] Figure 3 This is an example diagram of the speech recognition results provided by the embodiment of the present application;
[0044] Figure 4 This is a structural diagram of a speech recognition device provided in an embodiment of the present application;
[0045] Figure 5 It is a schematic diagram of the composition structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0046] In order to more clearly illustrate the purpose, technical solutions and advantages of the embodiments of the present application, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. It should be understood that the following description of the embodiments is intended to explain and illustrate the overall concept of the embodiments of the present application and should not be construed as limiting the embodiments of the present application. In the specification and drawings, the same or similar reference numerals refer to the same or similar parts or components. For the sake of clarity, the drawings are not necessarily drawn to scale, and some well-known parts and structures may be omitted in the drawings.
[0047] Currently, one approach to improving the accuracy of speech recognition engines is to collect a large amount of accented speech data to train a standard acoustic model. This enhances the model's generalization and, therefore, improves speech recognition accuracy. Another approach is to collect a portion of accented speech data and, based on the standard acoustic model trained on standard speech, use speaker adaptation technology to modify the model. However, these approaches have the following issues:
[0048] 1. Not universally applicable. Both of the above methods require the preparation of accented data in advance, including voice data and corresponding text information, to improve the performance of the standard acoustic model. However, when a speech recognition device leaves the factory, it is not certain where it will be used in the future, nor is it aware of the standard accent in that region. Therefore, improving the standard acoustic model is highly targeted to a specific region and is not universally applicable.
[0049] 2. Affecting the accuracy of standard speech recognition. Training the standard acoustic model with data of different accents or making corrections to the standard acoustic model can improve the recognition accuracy of non-standard speech to a certain extent, but it will also reduce the recognition accuracy of standard speech.
[0050] Based on the problems existing in the related art, an embodiment of the present application provides a speech recognition method, which obtains at least two speech segments belonging to different sound-producing objects contained in the original speech information, determines the speech confidence of the sound-producing object corresponding to each of the at least two speech segments, identifies at least two text information corresponding to the at least two speech segments, and uses the text information of the sound-producing object with high speech confidence to correct the text information of the sound-producing object with low speech confidence. The method provided in the present application improves the speech recognition accuracy of the speech recognition engine by identifying the speech confidence of the sound-producing object corresponding to each speech segment included in the original speech information, and using the text information corresponding to the high speech confidence to correct the text information corresponding to the low speech confidence. The present application uses certain speech segments in the original speech information obtained in real time to correct other speech segments, and does not require the preparation of a large amount of accented speech data in advance to train the standard acoustic model, so it will not affect the recognition accuracy of the speech recognition engine for standard speech.
[0051] The following describes exemplary applications of the electronic device provided in the embodiments of the present application. The electronic device provided in the embodiments of the present application can be implemented as various types of terminals such as laptops, tablet computers, desktop computers, mobile devices, etc., and can also be implemented as a server.
[0052] Figure 1 Schematic diagram of an application scenario of the speech recognition method provided in the embodiment of the present application. Figure 1 As shown, the speech recognition system 100 provided in the embodiment of the present application includes a recognition terminal 1001, a network 1002 and an output terminal 1003, wherein the recognition terminal 1001 can collect speech information generated by the sound-emitting object based on a speech collection device such as a microphone, and a computer program for recognizing speech information is also running. When the recognition terminal 1001 has both recognition and display functions, the recognition terminal 1001 can use the method of the embodiment of the present application to obtain and output the speech recognition results for display to the user; when the recognition terminal 1001 only has the recognition function, the recognition terminal can send the recognition results to the output terminal 1003 with a display function in a wireless or wired manner through the network 1002 for display. In other words, the recognition terminal 1001 and the output terminal 1003 can be implemented in hardware in the same device or in different devices.
[0053] Refer to Figure 2, which is an optional flow chart of the speech recognition method provided in an embodiment of the present application, which will be explained in conjunction with the steps shown in Figure 2.
[0054] Step 101: Acquire at least two speech segments belonging to different sound-producing objects contained in original speech information.
[0055] In some embodiments, the original voice information may refer to a conversation between at least two speakers, collected using a voice collection device, such as a microphone. The speakers may have different accents, such as a Mandarin accent or a Hubei accent. The original voice information may include at least one sub-voice information collected sequentially at a first time interval.
[0056] for Figure 2A Step 101 shown can be performed by Figure 2B Steps 1011 to 1013 are implemented and are described below.
[0057] Step 1011: Preprocess any first sub-voice information among the at least one collected sub-voice information to obtain second sub-voice information represented by voice features of the first sub-voice information.
[0058] During actual processing, the speech recognition system can process each collected speech message individually, or it can process the speech messages collected over a period of time. Therefore, the original speech message may include at least one sub-speech message, which is the speech message that the speech recognition system can collect at one time. The first time interval can be manually set based on the processing capability of the speech recognition system. For example, the first time interval can be 5 seconds. Since the processing method for each sub-speech message in the original speech message is the same, this application only uses the processing of the first sub-speech message as an example to illustrate.
[0059] In some embodiments, before performing speech recognition on the first sub-speech information, the collected first sub-speech information needs to be preprocessed to convert the first sub-speech information into second sub-speech information represented by speech features that facilitate subsequent recognition. The preprocessing may include filtering, framing, and windowing the first sub-speech information. The speech features may be represented based on Mel-Frequency Cepstral Coefficients (MFCCs) or Perceptual Linear Predictive (PLP).
[0060] Step 1012: Segment and cluster the second sub-speech information to obtain at least two sub-speech segments belonging to different sounding objects contained in the first sub-speech information.
[0061] In some embodiments, see Figure 2C , Figure 2B Step 1012 shown can be performed by Figure 2C Steps 10121 to 10122 are implemented and are described below.
[0062] Step 10121: Segment and cluster the second sub-speech information to obtain at least two intermediate speech segments belonging to different sound-producing objects.
[0063] In a specific implementation, when segmenting the second sub-speech information, the second sub-speech information can be segmented at least once with gradually decreasing durations according to actual needs, and the segmented speech segments can be clustered to obtain at least two intermediate speech segments belonging to different sounding objects. For example, the first segmentation can be performed with a first unit duration, the second segmentation can be performed with a second unit duration, and the third segmentation can be performed with a third unit duration, etc., wherein the first unit duration is greater than the second unit duration, which is greater than the third unit duration. The results obtained from each segmentation are clustered to obtain multiple intermediate speech segments corresponding to each cluster.
[0064] Step 10122: When at least one of the intermediate speech segments contains at least two different sound objects, each of the at least two intermediate speech segments is re-segmented and clustered until each of the at least two sub-speech segments belonging to different sound objects contains only one sound object.
[0065] Here, if the intermediate speech segment includes multiple sound objects, it means that the intermediate speech segment needs to be further segmented into smaller unit durations. In a specific implementation, if the number of unit duration segmentations reaches a set number (for example, 10 times), and the resulting intermediate speech segment still contains at least two different sound objects, the intermediate speech segment can be segmented and clustered using frame duration as the unit duration, so that each sub-segment contains only one sound object.
[0066] In order to better illustrate this process, the following examples are used to illustrate the embodiments of the present application. For example, when the second sub-speech information is segmented and clustered, it is first automatically segmented with a first unit time length, such as 1s, and then the sound objects to which each speech segment obtained by segmentation belongs are identified, and each speech segment is clustered according to the sound object to obtain multiple first intermediate speech segments; if it is determined that there is at least one first intermediate speech segment containing different sound objects in the multiple first intermediate speech segments, then each first intermediate speech segment in the multiple first intermediate speech segments is re-segmented and clustered with a second unit time length, such as 0.5s, to obtain multiple second intermediate speech segments; if it is determined that there is at least one second intermediate speech segment containing different sound objects in the multiple second intermediate speech segments, then the multiple second intermediate speech segments are further segmented and clustered with a third unit time length (for example, 0.1s), and the above process is iterated with gradually decreasing segmentation unit time lengths until each intermediate speech segment obtained by segmentation contains only one sound object, at which time each intermediate speech segment is a sub-speech segment.
[0067] Step 1013: Cluster the sub-speech segments belonging to the same sounding object in each first sub-speech information to obtain at least two speech segments belonging to different sounding objects contained in the original speech information.
[0068] Here, each first sub-voice information is processed according to the method from step 10121 to step 10122, that is, after processing the entire original voice information, the sub-voice segments corresponding to the same sound object are clustered together, and the original voice information can be obtained, including at least two voice segments belonging to different sound objects.
[0069] Continue to see Figure 2A, step 102: determining the speech confidence of the sounding object corresponding to each speech segment in the at least two speech segments.
[0070] It should be noted that the current standard acoustic models are mostly standard acoustic models trained using Mandarin. Therefore, the recognition results obtained by using the current standard acoustic model to recognize speech with a Mandarin accent have a higher confidence level than the recognition results obtained by recognizing speech with other accents. In other words, the closer the speech is to the pronunciation standard of the standard acoustic model, the higher the speech confidence level of the recognition results obtained by using the current standard acoustic model.
[0071] In actual application, since a speech segment may contain multiple sentences, the recognition of the speech segment is usually performed using a standard acoustic model in units of sentences. Therefore, in some embodiments, Figure 2B , Figure 2A Step 102 shown can be performed by Figure 2B Steps 1021 to 1023 are implemented and are described below.
[0072] Step 1021: Divide any first speech segment of the at least two speech segments into sentences to obtain a plurality of sentences contained in the first speech segment.
[0073] Here, when dividing the first speech segment into sentences, the speaker's pause can be used as a sign of sentence division, or the speaker's two pauses can be used as signs of sentence division. This embodiment does not limit the sign of sentence division.
[0074] Step 1022: Identify any first sentence among the multiple sentences and obtain a sentence confidence level corresponding to the first sentence.
[0075] In some embodiments, see Figure 2D , Figure 2B Step 1022 shown can be performed by Figure 2D Steps 10221 to 10223 are implemented and are described below.
[0076] Step 10221: Determine the type of the first statement; the first statement is any one of the multiple statements.
[0077] Here, the type of the first sentence can be English, Chinese, or other languages.
[0078] Step 10222: Input the first sentence into a recognition model corresponding to the type to obtain a probability that the first sentence is correctly recognized.
[0079] It should be noted that any sentence input into the standard acoustic model will have a corresponding score in the modeling unit of the standard acoustic model. In the specific implementation, the score is represented by the probability that the first sentence is correctly recognized. The higher the score, the higher the confidence of the sentence. For example, see Figure 3 , Figure 3 This is an example diagram of the speech recognition results provided by the embodiment of the present application. Figure 3 The accent of the speaker spk0 is Mandarin, and the accent of spk1 is Hubei accent. Since the pronunciation and intonation of the word "owner" in Hubei accent are quite different from those in Mandarin, when using the standard acoustic model trained with Mandarin for speech recognition, the score of the sentence of the speaker spk0 is higher. Accordingly, the speech confidence of the speaker spk0 is greater than that of the speaker spk1.
[0080] In a specific implementation, the modeling unit in the standard acoustic model will recognize the first sentence according to the type of the first sentence using a recognition model corresponding to the type, so as to obtain a score corresponding to the first sentence. For example, when the type of the first sentence is Chinese, the first sentence is input into the Chinese recognition model for recognition. The evaluation method of the Chinese recognition model is based on the initials or finals of pinyin. For example, when the first sentence is "hello", the corresponding initials or finals are "n, i, h, ao". During recognition, the recognition model will give the probability of the initials or finals output in each frame. For the above example, if the probability of the first frame outputting "n" is 0.98, the probability of "l" is 0.95, the probability of the second frame outputting "i" is 0.98, and the probability of the third frame outputting "i" is 0.98. The probability of the frame outputting "h" is 0.98, and the probability of the fourth frame outputting "ao" is 0.98. The probability that the Chinese recognition model outputs the above four recognition results as other initials or finals is 0. According to this method, after obtaining the output results of the four recognitions, the average probability corresponding to the output results is calculated, and the probability of "nihao" is 0.98, and the probability of "lihao" is 0.9725. Then, the probability of "hello" being correctly recognized is 0.98. At the same time, the score corresponding to the first sentence "hello" is also 0.98.
[0081] Step 10223: Obtain the statement confidence corresponding to the first statement based on the probability.
[0082] In some embodiments, the probability that the first sentence is correctly recognized may be directly used as the confidence level of the sentence corresponding to the first sentence.
[0083] Continue to see Figure 2B , step 1023, determining the speech confidence corresponding to the first speech segment based on each of the sentence confidences.
[0084] In some embodiments, after the score given to each sentence by the modeling unit in the standard acoustic model, i.e., the sentence confidence, the average speech confidence corresponding to multiple sentences can be further obtained based on the sentence confidence corresponding to each sentence in the multiple sentences, and the average speech confidence can be used as the speech confidence of the first speech segment.
[0085] Continue to see Figure 2A Step 103: Identify at least two text messages corresponding to the at least two voice segments, and use the text information of the sound object corresponding to the target voice confidence to correct the text information corresponding to the sound object corresponding to other voice confidences, wherein the target voice confidence is higher than the other voice confidences.
[0086] In some embodiments, at least two speech segments are decoded using a standard acoustic model to obtain text information corresponding to the speech segments. Because each speech segment corresponds to a different speaker, who may come from different regions and have different accents, different speech segments have different speech confidence levels. In specific implementations, the higher the speech confidence level of a speech segment, the closer the speaker's accent is to the standard accent. Therefore, the text information corresponding to the speech segment with a high speech confidence level can be used to correct the text information corresponding to the speech segment with a low speech confidence level.
[0087] In some embodiments, see Figure 2E , Figure 2E This is an optional flow chart of the speech recognition method provided in the embodiment of the present application. Figure 2A In step 103 shown, the step of using the text information of the target speech confidence corresponding to the sound object to correct the text information corresponding to other speech confidence corresponding to the sound object can be achieved by Figure 2E Step 1031 is implemented.
[0088] Step 1031: Determine at least two words in the at least two text messages that belong to different sound objects and meet the target similarity, and use the words in the text information corresponding to the sound object corresponding to the target speech confidence to correct the words in the text information corresponding to the sound object corresponding to other speech confidences.
[0089] Here, based on any word contained in a text message, the similarity between the word and the words contained in other text messages can be calculated. If the similarity between the two words meets the target similarity, the word corresponding to the text message with high confidence is used to correct the word corresponding to the text message with low confidence. For example, there are two text messages, text message A and text message B. Text message A is: Prepare the documents needed for tomorrow's meeting, which contains the word "document"; text message B is: Already prepared, there are three articles, which contain the word "article". If it is recognized that the speech confidence corresponding to text message A is greater than the speech confidence corresponding to text message B, and the similarity between the words "file" and "article" meets the target similarity, then the word "article" contained in text message B is replaced with "file" to complete the correction.
[0090] See further Figure 2F , Figure 2F This is an optional flow chart of the speech recognition method provided in the embodiment of the present application. Figure 2E In step 1031 shown in FIG, the step of determining at least two words in the at least two text messages that belong to different sound objects and meet the target similarity can be performed by Figure 2F Steps 10311 to 10313 are implemented and are described below.
[0091] Step 10311: perform word segmentation processing on each of the at least two text messages to obtain key words contained in each text message.
[0092] In some embodiments, the keyword is a collective term for the words obtained through word segmentation. During word segmentation, the text information can be segmented as a morpheme. A morpheme is the smallest sound-meaning combination in a language that has pronunciation and meaning. Words without actual meaning, such as modal particles such as "ne" and "ah," cannot be considered a morpheme.
[0093] Step 10312: Determine a first similarity between the key words contained in different text information.
[0094] Step 10313: Obtain at least two keywords whose first similarity meets the target similarity from the key words contained in each of the at least two text messages; the at least two keywords are at least two words in the at least two text messages that belong to different voice objects and meet the target similarity.
[0095] In some embodiments, a first similarity between a first keyword included in any text message and a second keyword included in other text messages is calculated. The first keyword is any keyword among the keywords included in the selected text message, and the second keyword is any keyword among the keywords included in other text messages other than the selected text message. After the first keyword included in the selected text message is processed, the next unprocessed text message is selected as the selected text message, and the above operation is repeated to process the keywords included in the other text messages until the key words included in all text messages are processed.
[0096] Continue to see Figure 2A , step 104: output the at least two corrected text messages.
[0097] In some embodiments, when the object of correction is a single word or phrase, the sentence in which the word or phrase is located can be directly output. In other embodiments, the at least two corrected text messages can be output on one or more devices with display functions. The device can be an electronic device that integrates correction and display functions, or it can be an electronic device that only provides display functions.
[0098] The speech recognition method provided in an embodiment of the present application obtains at least two speech segments belonging to different sound-producing objects contained in the original speech information, determines the speech confidence of the sound-producing object corresponding to each of the at least two speech segments, identifies at least two text information corresponding to the at least two speech segments, and uses the text information of the sound-producing object with high speech confidence to correct the text information of the sound-producing object with low speech confidence. In this way, the speech recognition method provided in the present application improves the speech recognition accuracy of a standard speech recognition engine by determining the speech confidence of the sound corresponding to the speech segment and using the text information corresponding to the high speech confidence to correct the text information corresponding to the low speech confidence.
[0099] Compared with related technologies, the speech recognition method proposed in the embodiment of the present application does not require setting up different speech recognition devices for different regions. Based on the standard acoustic model, it only needs to use high-confidence text information to correct low-confidence text information, so that the accented speech information can be corrected to the speech information corresponding to the standard accent. Therefore, it is not subject to regional control and is more universal. In addition, since the embodiment of the present application does not need to use accented speech to train the standard acoustic model, it will not reduce the recognition accuracy of standard speech.
[0100] Based on the same inventive concept, the present application also provides a speech recognition device, referring to Figure 4 , Figure 4This is a schematic diagram of the structure of a speech recognition device provided in an embodiment of the present application. The device 40 includes: an acquisition module 401, a determination module 402, a correction module 403 and an output module 404, wherein;
[0101] The acquisition module 401 is used to acquire at least two speech segments belonging to different sounding objects contained in the original speech information;
[0102] The determining module 402 is configured to determine the speech confidence of the sounding object corresponding to each of the at least two speech segments;
[0103] The correction module 403 is configured to identify at least two text information corresponding to the at least two speech segments, and use the text information of the target speech confidence corresponding to the sound object to correct the text information corresponding to the sound object of other speech confidences, wherein the target speech confidence is higher than the other speech confidences;
[0104] The output module 404 is configured to output the at least two pieces of text information after correction.
[0105] In some embodiments, the correction module 403 is also used to determine at least two words in the at least two text information that belong to different sound objects and meet the target similarity, and use the words in the text information corresponding to the sound object corresponding to the target speech confidence to correct the words in the text information corresponding to the sound object corresponding to other speech confidences.
[0106] In some embodiments, the correction module 403 is also used to perform word segmentation processing on each of the at least two text messages to obtain the key words contained in each text message; determine the first similarity between the key words contained in different text messages; obtain at least two keywords whose first similarity meets the target similarity from the key words contained in each text message in the at least two text messages; the at least two keywords are at least two words in the at least two text messages that belong to different sound objects and meet the target similarity.
[0107] In some embodiments, the original voice information includes at least one sub-voice information collected sequentially according to a first time interval, and the acquisition module 401 is also used to pre-process any first sub-voice information in the at least one sub-voice information collected to obtain a second sub-voice information represented by the voice features of the first sub-voice information; segment and cluster the second sub-voice information to obtain at least two sub-voice segments belonging to different sound objects contained in the first sub-voice information; cluster the sub-voice segments belonging to the same sound object in each first sub-voice information to obtain at least two voice segments belonging to different sound objects contained in the original voice information.
[0108] In some embodiments, the acquisition module 401 is also used to segment and cluster the second sub-speech information to obtain at least two intermediate speech segments belonging to different sound objects; when at least one of the intermediate speech segments contains at least two different sound objects, each of the at least two intermediate speech segments is re-segmented and clustered until each of the at least two sub-speech segments belonging to different sound objects obtained contains only one sound object.
[0109] In some embodiments, the determination module 402 is also used to perform sentence division on any first speech segment of the at least two speech segments to obtain multiple sentences contained in the first speech segment; identify any first sentence of the multiple sentences to obtain the sentence confidence corresponding to the first sentence; and determine the speech confidence corresponding to the first speech segment based on each of the sentence confidences.
[0110] In some embodiments, the determination module 402 is also used to determine the type of the first statement; the first statement is any one of the multiple statements; the first statement is input into the recognition model corresponding to the type to obtain the probability that the first statement is correctly recognized; and the confidence of the statement corresponding to the first statement is obtained based on the probability.
[0111] It should be noted that the description of the device embodiment of the present application is similar to the description of the method embodiment described above, and has similar beneficial effects as the method embodiment, so it will not be repeated. For technical details not disclosed in the device embodiment, please refer to the description of the method embodiment of the present application for understanding.
[0112] It should be noted that, in the embodiment of the present application, if the above-mentioned screen sharing method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a terminal to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM, Read Only Memory), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.
[0113] The embodiment of the present application also provides an electronic device, Figure 5 Schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 5 As shown, the electronic device 50 includes at least: a processor 501 and a computer-readable storage medium 502 configured to store executable instructions, wherein the processor 501 generally controls the overall operation of the electronic device. The computer-readable storage medium 502 is configured to store instructions and applications executable by the processor 501, and can also cache data to be processed or processed by the processor 501 and various modules in the electronic device 50. This can be implemented using flash memory or random access memory (RAM).
[0114] The embodiment of the present application provides a storage medium storing executable instructions, wherein the executable instructions are stored. When the executable instructions are executed by a processor, the processor will execute the speech recognition method provided by the embodiment of the present application, for example, Figure 2A The method shown.
[0115] In some embodiments, the storage medium can be a computer-readable storage medium, such as a ferroelectric random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); it can also be various devices including one or any combination of the above memories.
[0116] In some embodiments, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0117] By way of example, executable instructions may, but need not necessarily, correspond to a file in a file system, may be stored as part of a file storing other programs or data, such as one or more scripts in a Hypertext Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions). By way of example, executable instructions may be deployed for execution on one computing device, or on multiple computing devices located at one site, or on multiple computing devices distributed across multiple sites and interconnected by a communication network.
[0118] It should be noted that the various technical features in the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.
[0119] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.
[0120] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed.
[0121] The above is merely an embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A speech recognition method, comprising: Acquire at least two speech segments belonging to different sounding objects contained in the original speech information; Determining a speech confidence of a sounding object corresponding to each speech segment in at least two speech segments; the speech confidence represents a probability that the speech segment is correctly recognized; Identifying at least two text messages corresponding to the at least two speech segments, and using the text information of the target speech confidence corresponding to the sound object to correct the text information corresponding to the sound object of other speech confidences, wherein the target speech confidence is higher than the other speech confidences; Output at least two corrected text messages; The text information corresponding to the sound object corresponding to the target speech confidence level is used to correct the text information corresponding to the sound object corresponding to other speech confidence levels, including: Determine at least two words in the at least two text messages that belong to different sound objects and meet the target similarity, and use the words in the text information corresponding to the sound object corresponding to the target speech confidence to correct the words in the text information corresponding to the sound object corresponding to other speech confidences.
2. The method according to claim 1, wherein determining at least two words in the at least two text messages that belong to different utterance objects and meet a target similarity comprises: Performing word segmentation processing on each of the at least two text messages to obtain key words contained in each text message; Determining a first similarity between the key words contained in different text information; At least two keywords whose first similarity meets the target similarity are obtained from the key words contained in each of the at least two text messages; the at least two keywords are at least two words in the at least two text messages that belong to different voice objects and meet the target similarity.
3. The method according to claim 1, wherein the original voice information includes at least one sub-voice information collected sequentially at a first time interval, and the obtaining of at least two voice segments belonging to different sounding objects contained in the original voice information comprises: Preprocessing any first sub-speech information among the at least one collected sub-speech information to obtain second sub-speech information represented by speech features of the first sub-speech information; performing segmentation and clustering processing on the second sub-speech information to obtain at least two sub-speech segments belonging to different sounding objects contained in the first sub-speech information; The sub-speech segments belonging to the same sounding object in each first sub-speech information are clustered to obtain at least two speech segments belonging to different sounding objects contained in the original speech information.
4. The method according to claim 3, wherein the segmenting and clustering of the second sub-speech information to obtain at least two sub-speech segments belonging to different utterance objects contained in the first sub-speech information comprises: Segmenting and clustering the second sub-speech information to obtain at least two intermediate speech segments belonging to different sounding objects; In the case that at least one of the intermediate speech segments contains at least two different sound objects, each of the at least two intermediate speech segments is re-segmented and clustered until each of the at least two sub-speech segments belonging to different sound objects obtained contains only one sound object.
5. The method according to claim 1, wherein determining the speech confidence of the utterance corresponding to each of the at least two speech segments comprises: Performing sentence division on any first speech segment of the at least two speech segments to obtain a plurality of sentences contained in the first speech segment; Recognize any first sentence among the multiple sentences to obtain a sentence confidence score corresponding to the first sentence; The speech confidence corresponding to the first speech segment is determined based on each of the sentence confidences.
6. The method according to claim 5, wherein the step of identifying any first sentence among the plurality of sentences and obtaining a sentence confidence score corresponding to the first sentence comprises: determining a type of the first statement; The first sentence is any one of the multiple sentences; Inputting the first sentence into a recognition model corresponding to the type to obtain a probability that the first sentence is correctly recognized; A confidence level of a statement corresponding to the first statement is obtained based on the probability.
7. A speech recognition device, comprising: An acquisition module, configured to acquire at least two speech segments belonging to different sounding objects contained in the original speech information; A determination module, configured to determine a speech confidence of a sounding object corresponding to each of at least two speech segments; the speech confidence represents a probability that the speech segment is correctly recognized; a correction module, configured to identify at least two text messages corresponding to the at least two speech segments, and use the text information of the sound object corresponding to the target speech confidence to correct the text information corresponding to the sound object corresponding to other speech confidences, wherein the target speech confidence is higher than the other speech confidences; An output module, configured to output at least two corrected text messages; The correction module is also used to determine at least two words in the at least two text information that belong to different sound objects and meet the target similarity, and use the words in the text information corresponding to the sound object corresponding to the target speech confidence to correct the words in the text information corresponding to the sound object corresponding to other speech confidences.
8. An electronic device, comprising: a memory for storing executable instructions; A processor, configured to implement the speech recognition method according to any one of claims 1 to 6 when executing the executable instructions stored in the memory.
9. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the speech recognition method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Speech recognition method and device
CN112599114A
Speech recognition optimization method and device, computer equipment and storage medium
CN114420123A