Speech recognition method and device, electronic equipment and storage medium
By obtaining the current voice data and recognized voice data in speech recognition, and grouping and screening with the matching degree of hot word databases, the problems of low accuracy and recall of hot word recognition are solved, and more efficient voice recognition effect is achieved.
Patent Information
- Application Number
- CN202411900034.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-20
AI Technical Summary
During the speech recognition process, the accuracy of recognition of hot words is affected, resulting in a decrease in recall rate and affecting the overall recognition effect. There is a local optimal strategy in prior art such as top-k sampling methods, which may result in the correct path prefix being discarded in advance.
By obtaining the current voice data, identified voice data and hot vocabulary database matching the current context, the current voice data and model score corresponding to the current voice data are determined, the first voice data is determined from the current voice data based on the model score, and it is combined with the identified voice data to obtain the second voice data, and the second voice data is grouped according to the matching degree of preset hot words in the hot vocabulary database to filter out the target voice results.
It effectively reduces the error rate of speech recognition when the voice data contains hot words, improves the recall rate of hot words, and improves the overall effect of speech recognition.
Smart Images

Figure CN119943030A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech recognition, and in particular to speech recognition methods, devices, electronic devices and storage media. Background Art
[0002] In the field of speech recognition, the recognition of hot words is of great significance, especially in applications in specific fields. Hot words usually refer to the proprietary vocabulary in a certain field, which can represent the semantic core of a speech. Therefore, the recognition accuracy of hot words is often more critical than the recognition accuracy of the entire speech. Hot word recognition errors will directly affect the understanding of the entire speech. Therefore, hot words need to be paid special attention to in the speech recognition process.
[0003] In the relevant implementation schemes of speech recognition, the path selection method is usually used to determine the optimal recognition result. Due to the diversity of paths and the limitation of computing resources, it is impossible to traverse all possible paths. Therefore, in the process of frame-by-frame recognition, the top-k sampling method is usually used. This method only retains the k paths with the highest scores and cuts off other paths each time a new path is obtained. Although the top-k sampling method can effectively reduce the amount of calculation, it is a local optimal strategy rather than a global optimal strategy. This means that in some cases, the correct path prefix may be discarded in advance because the score is too low midway, resulting in the failure to obtain the result containing the hot word in the end. This phenomenon will significantly reduce the recall rate of hot words and affect the overall effect of speech recognition. Summary of the invention
[0004] In view of the above problems, a speech recognition method, device, electronic device and storage medium are proposed to overcome the above problems or at least partially solve the above problems, including:
[0005] A speech recognition method, the method comprising:
[0006] Obtain current voice data, recognized voice data, and a hot word library that matches the current context;
[0007] Determine, based on the recognized vocalization data and the current voice data, the current vocalization data corresponding to the current voice data and the model score corresponding to the current vocalization data;
[0008] Determining first vocal data from the current vocal data according to the model score, and combining the first vocal data with the recognized vocal data to obtain second vocal data;
[0009] Determine the matching degree between the second voice data and the preset hot words in the hot word library, and group the second voice data according to the matching degree to obtain second voice data groups;
[0010] For each of the second sound data groups, determine a first path score corresponding to each of the second sound data in the second sound data group, and determine a third sound data group from the second sound data group according to the first path score;
[0011] Screening the third vocalization data group to obtain a target vocalization result;
[0012] The target utterance result is determined as the recognition result of the current voice data.
[0013] In an optional embodiment of the present application, determining the current voice data corresponding to the current voice data and the model score corresponding to the current voice data according to the recognized voice data and the current voice data includes:
[0014] Extracting features from the current voice data to obtain audio features;
[0015] Inputting the audio features into a preset encoder to obtain a first feature matrix;
[0016] Inputting the recognized vocalization data into a preset decoder to obtain a second feature matrix;
[0017] The first feature matrix and the second feature matrix are input into a preset connector to obtain current voice data corresponding to the current voice data and a model score corresponding to the current voice data.
[0018] In an optional embodiment of the present application, determining the first vocalization data from the current vocalization data according to the model score includes:
[0019] Sorting the current vocalization data according to the model score to obtain sorted current vocalization data;
[0020] According to a preset first target quantity, a first target quantity of first sound data is determined from the sorted current sound data.
[0021] In an optional embodiment of the present application, after determining the first target number of first sound data from the sorted current sound data according to the preset first target number, the method further comprises:
[0022] The bias scores of the first utterance data and the recognized utterance data are calculated by using a preset hot word detection algorithm.
[0023] In an optional embodiment of the present application, for each second sound production data group, determining a first path score corresponding to each second sound production data in the second sound production data group includes:
[0024] Obtaining a second path score of the identified vocal data in the second vocal data;
[0025] The second path score, the model score and the bias score are summed to obtain a first path score corresponding to the second vocalization data.
[0026] In an optional embodiment of the present application, determining a third sound data group from the second sound data group according to the first path score includes:
[0027] sorting the second sound data included in the second sound data group according to the first path score to obtain sorted second sound data group;
[0028] A second target number of second sound data is selected from the sorted second sound data groups according to a preset second target number to obtain a third sound data group.
[0029] In an optional embodiment of the present application, the third sound production data group includes fourth sound production data and fifth sound production data, and the screening of the third sound production data group to obtain a target sound production result includes:
[0030] Calculating a difference between the first path score of the fourth sound production data and the first path score of the fifth sound production data to obtain a path score difference;
[0031] Calculating a difference between a matching degree of the fourth voice data and a matching degree of the fifth voice data to obtain a matching degree difference;
[0032] The relationship between the path score difference and the matching degree difference is determined, and the third voice data group is screened according to the relationship between the path score difference and the matching degree difference to obtain a target voice result.
[0033] The present application also discloses a speech recognition device, which includes:
[0034] A data acquisition module is configured to acquire current voice data, recognized voice data, and a hot word library matching the current context;
[0035] a model score determination module, configured to determine, based on the recognized vocalization data and the current voice data, current vocalization data corresponding to the current voice data and a model score corresponding to the current voice data;
[0036] A first sampling module is configured to determine first vocalization data from the current vocalization data according to the model score, and combine the first vocalization data with the recognized vocalization data to obtain second vocalization data;
[0037] A data bucketing module is configured to determine a matching degree between the second voice data and a preset hot word in a hot word library, and group the second voice data according to the matching degree to obtain a second voice data group;
[0038] A second sampling module is configured to determine, for each of the second sound data groups, a first path score corresponding to each of the second sound data in the second sound data group, and determine a third sound data group from the second sound data group based on the first path score;
[0039] A third sampling module is configured to filter the third sound data group to obtain a target sound result;
[0040] The result determination module is configured to determine the target utterance result as the recognition result of the current voice data.
[0041] An embodiment of the present application also discloses an electronic device, including a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program implements the speech recognition method as described above when executed by the processor.
[0042] The embodiment of the present application further discloses a computer-readable storage medium, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the speech recognition method as described above is implemented.
[0043] The embodiments of the present application have the following advantages:
[0044] By acquiring the current voice data, the recognized voice data and the hot word library matching the current context, the current voice data corresponding to the current voice data and the model score corresponding to the current voice data are determined according to the recognized voice data and the current voice data, and the first voice data is determined from the current voice data according to the model score, and the first voice data is combined with the recognized voice data to obtain the second voice data, and the matching degree between the second voice data and the preset hot words in the hot word library is determined, and the second voice data is grouped according to the matching degree to obtain the second voice data group, for each second voice data group, the first path score corresponding to each second voice data in the second voice data group is determined, and according to the first path score, a third voice data group is determined from the second voice data group, and the third voice data group is screened to obtain the target voice result, and the target voice result is determined as the recognition result of the current voice data, which effectively reduces the voice recognition error rate when the voice data contains hot words. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the description of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0046] Figure 1 is a flowchart of a method for speech recognition provided by an embodiment of the present application;
[0047] Figure 2 It is a structural block diagram of a speech recognition device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0048] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0049] Reference Figure 1 , shows a flowchart of a speech recognition method provided by an embodiment of the present application, which may specifically include the following steps:
[0050] Step 101: Obtain current voice data, recognized speech data, and a hot word library matching the current context.
[0051] Among them, the current voice data refers to all data contained in the current audio frame (or the audio clip in the current time window) during the speech recognition process. The current voice data may include acoustic features, wherein the acoustic features include Mel-frequency cepstral coefficients, spectrograms, linear prediction coefficients, etc. In the speech recognition process, the recognized sound data refers to multiple possible paths that have been recognized at the current moment. These paths are usually represented in the form of a set of tokens, each token can be a word, and the set of these tokens constitutes the recognized sound data. The hot word library that matches the current context refers to the medical terminology library, the vehicle inspection terminology library, or other terminology libraries with strong professional fields.
[0052] Before extracting the first frame, the recognized vocal data can be an initial node representing silence, and then the voice data of the first frame is used as the current voice data and recognized. Assuming that the current context is a medical environment, it is also necessary to obtain a medical hot word library and start executing the specific steps of subsequent voice recognition.
[0053] Step 102: Determine the current voice data corresponding to the current voice data and the model score corresponding to the current voice data based on the recognized voice data and the current voice data.
[0054] In speech recognition tasks, speech data is a time series data, and the sound data corresponding to the current speech is usually related to the recognized sound data. Among them, the model score can refer to the probability distribution from the previous frame token to the current frame token.
[0055] In the embodiment of the present application, the current vocalization data corresponding to the current voice data refers to the vocalization data that the current voice data may correspond to under the influence of the recognized voice data.
[0056] In some embodiments of the present application, step 102 may include the following sub-steps:
[0057] Sub-step 11: Extract features from the current speech data to obtain audio features.
[0058] Sub-step 12: inputting the audio feature into a preset encoder to obtain first feature information;
[0059] Sub-step 13: inputting the recognized voice data into a preset decoder to obtain a second feature matrix;
[0060] Sub-step 14: Input the first feature information and the second feature information into a preset connector to obtain the current voice data corresponding to the current voice data and the model score corresponding to the current voice data.
[0061] Among them, feature extraction of the current speech data to obtain audio features may specifically include steps such as preprocessing, time-frequency conversion, feature calculation and feature selection. The preprocessing includes denoising, normalization, and framing of the audio signal of the current speech data; time-frequency conversion refers to converting the audio signal from the time domain to the frequency domain (such as STFT, wavelet transform, etc.); feature calculation refers to calculating the required features (such as MFCC, energy, zero-crossing rate, etc.); feature selection refers to selecting a suitable feature subset according to task requirements; feature normalization is to normalize the features for subsequent classification or recognition tasks.
[0062] In some embodiments of the present application, the Zipformer model can be used to extract features from the collected speech data. Zipformer is an efficient speech recognition model architecture that combines three core components: encoder, decoder, and joiner. Among them, the encoder of Zipformer is responsible for extracting the features of speech data. The input of the encoder is raw data, such as audio waveform or feature sequence after feature extraction; the output of the encoder is a feature matrix after feature sequence processing.
[0063] In some implementations of this embodiment, first, feature extraction is performed on the current voice data, and the extracted audio features are input into the encoder in the Zipformer model. The encoder processes the features to obtain a first feature matrix. This first feature matrix is also the encoder feature, and the first feature matrix can be a high-dimensional representation of the audio feature. The recognized sound data is input into the preset decoder Decoder to obtain the corresponding second feature matrix, which is a high-dimensional representation of the recognized sound data. The first feature matrix (i.e., the feature representation of the current sound data) and the second feature matrix (i.e., the feature representation of the recognized sound data) are input into the Joiner connector in the Zipformer model. The Joiner connector calculates the current sound data that the current voice data may correspond to, and the model score corresponding to the current sound data. The model score reflects the degree of match between the current sound data and the recognized sound data. In a specific implementation, the model score refers to the probability distribution of the recognized sound data to the current voice data. Assuming there are only three Chinese characters, namely, ammonia, ethyl, and base, the obtained model score (i.e., probability distribution) can be a 3×3 matrix M
[0064]
[0065] Among them, Mij represents the probability of going from row i to row j, for example, M 12 The probability of "NH3" to "B" is 0.3, M 23The probability of representing "B" to "group" is 0.3.
[0066] Step 103: Determine first voice data from the current voice data according to the model score, and combine the first voice data with the recognized voice data to obtain second voice data;
[0067] In some embodiments of this embodiment, on the basis of the recognized voice data, adding a new corresponding token can obtain a new path, which is the second voice data of the embodiment of the present application. Since the recognized voice data (the current path set) is not a single path, but a current path set composed of multiple paths, and multiple different tokens may be added to each current path, thus generating more branches. Therefore, the obtained new paths are also multiple, that is, multiple second voice data can be obtained.
[0068] In some embodiments of the present application, "determining first voice data from the current voice data according to the model score" in step 103 may further include the following sub-steps:
[0069] Sub-step 21: Sort the current voice data according to the model score to obtain the sorted current voice data;
[0070] Sub-step 22: Determine a first target number of first voice data from the sorted current voice data according to a preset first target number.
[0071] In some embodiments of this embodiment, the model scores of each token calculated in sub-step 13 are sorted in descending order, that is, the current voice data is sorted in descending order according to the model score, and k1 of the largest scores are selected and the corresponding tokens are determined. That is, according to a preset first target number k1, k1 first voice data of the first target number are determined from the sorted current voice data. Among them, k1 can be adjusted according to actual needs. For example, as the computer hardware performance continues to improve, k1 can be set larger, leaving more data, making the speech recognition result more accurate; for another example, in the case where the hardware performance cannot be guaranteed, k1 can be set smaller, leaving less data, reducing the calculation amount of subsequent steps.
[0072] In some embodiments of the present application, after sub-step 22, the following sub-steps may further be included:
[0073] Sub-step 31: Calculate the bias score of the first voice data and the recognized voice data through a preset hot word detection algorithm.
[0074] Among them, the preset hot word detection algorithm refers to AC-automaton. AC-automaton (Aho-CorasickAutomaton) is an efficient multi-pattern matching algorithm. This application uses AC-automaton for hot word detection. For example, multiple pattern strings (such as hot words or keywords) are constructed into an automaton, so as to match multiple patterns at the same time in one scan.
[0075] In some implementations of this embodiment, the recognized speech data and the corresponding token (ie, the first utterance data) are subjected to hot word detection calculation by an AC-automaton, and bias scores corresponding to the first utterance data and the recognized utterance data can be obtained.
[0076] After executing the above sub-steps 21-22, k1 first sound data can be determined, and through the execution of sub-step 31, the bias scores of k1 first sound data and the identified sound data can be obtained, and then the identified sound data and k1 first sound data can be combined to obtain multiple new paths, that is, multiple second sound data.
[0077] Step 104: Determine the matching degree between the second voice data and the preset hot words in the hot word library, and group the second voice data according to the matching degree to obtain second voice data groups.
[0078] Among them, the hot word library is a collection of multiple preset hot words. For example, the preset hot words in the medical hot word library can include commonly used drug names, disease names, symptom descriptions, medical terms, etc. These hot words play an important role in scenarios such as medical speech recognition, electronic medical records, and intelligent consultation systems.
[0079] In some implementations of this embodiment, the obtained new path (i.e., the second voice data) is grouped according to the prefix length of the hot word matched with the hot word library. Assuming that there is a preset hot word in the hot word library in the medical context: aminoethyl indole, the second voice data composed of the recognized voice data and the second voice data corresponding to the current voice data (determined by the first voice data corresponding to the current voice data) includes "want some catering", "some catering", "catering", "want some ammonia", "some ammonia", "ammonia", "want some aminoethyl", "some aminoethyl", "aminoethyl", then the matching degree of the second voice data with the preset hot word can be determined, and the multiple second voice data can be grouped according to the matching degree, which can also be called bucketing, to obtain different second voice data groups. For example, the matching degree of "want some aminoethyl", "some aminoethyl", "aminoethyl" with the preset hot word "aminoethyl indole" is 2, the matching degree of "want some ammonia", "some ammonia", "ammonia" with the preset hot word "aminoethyl indole" is 1, and the matching degree of "want some catering", "some catering", "catering" with the preset hot word is 0. These 9 second voice data can be divided into 3 groups. The second voice data group with a matching degree of 2 includes "want some ethyl ammonia", "some ethyl ammonia", "ethyl ammonia"; the second voice data group with a matching degree of 1 includes "want some ammonia", "some ammonia", "ammonia"; the second voice data group with a matching degree of 0 includes "want some food", "some food", "food".
[0080] Step 105: for each second sound data group, determine the first path score corresponding to each second sound data in the second sound data group, and determine a third sound data group from the second sound data group based on the first path score.
[0081] After determining multiple second voice data groups, in order to reduce the subsequent calculation amount, the embodiment of the present application performs a second sampling on each group, and can determine a third voice data group that is more in line with the actual situation from the second voice data group based on the first path score corresponding to the second voice data.
[0082] Among them, the first path score may refer to the sum of the path score of the identified vocal data, the model score calculated by the Joiner model in the above step, and the bias score calculated by the AC-automaton in the above step, and then the second vocal data group is screened according to the first path score. In some implementations of this embodiment, the use of the bias score can be illustrated by the following example. Assuming that the original path score of "want some ammonia" is 7, since the AC-automaton calculates its matching degree as 1 and the bias score is 3, then the biased score of this path can be 7+3×1. Similarly, the original path score of "want some ammonia" is 9. Since the matching degree is 2, the biased score is 9+2×6=15.
[0083] In some embodiments of the present application, in step 105, “for each second sound data group, determining the first path score corresponding to each second sound data in the second sound data group” may include the following sub-steps:
[0084] Sub-step 41: Obtain a second path score of the recognized vocal data in the second vocal data.
[0085] Sub-step 42: summing up the second path score, the model score and the bias score to obtain the first path score corresponding to the second vocalization data.
[0086] Among them, the second path score of the recognized voice data refers to the path score calculated by the recognized voice data in the recognition stage. In order to distinguish it from the path score corresponding to the second voice data, the path score of the recognized voice data is called the second path score in this application. The model score and bias score obtained in the above steps are summed with the second path score of the recognized voice data to obtain the first path score of the second voice data, that is, the path score of the newly generated path.
[0087] In some embodiments of the present application, the step 105 of “determining a third sound data group from the second sound data group according to the first path score” may include the following sub-steps:
[0088] Sub-step 51: sorting the second sound data included in the second sound data group according to the first path score to obtain sorted second sound data group.
[0089] Sub-step 52: selecting a second target number of second sound data from the sorted second sound data groups according to a preset second target number to obtain a third sound data group.
[0090] In some implementations of this embodiment, the second voice data in each group is sorted by scores, and the k2 paths with the highest scores are retained, wherein the second target number is k2. Specifically, the second target number k2 can be 4. Assuming k2 is 4, all data in the second voice data groups with matching degrees of 2, 1, and 0 can be retained. Of course, the value of k2 can also be adjusted according to actual needs. Assuming k2 is adjusted to 2, in this example, "aminoethyl" in the second voice data with a matching degree of 2 can be retained, and the second voice data group with a matching degree of 2 may retain "want some aminoethyl" and "some aminoethyl", and the second voice data group with a matching degree of 1 may retain "want some ammonia" and "some ammonia", and the second voice data group with a matching degree of 0 may retain "want some catering" and "some catering", that is, the third voice data groups obtained are respectively a matching degree 2 group, a matching degree 1 group, and a matching degree 0 group.
[0091] Step 106: Screen the third vocalization data group to obtain a target vocalization result.
[0092] After obtaining the third sound data group, the present application further screens the sound data in each third sound data group to further reduce the amount of sound data and further reduce the subsequent calculation amount.
[0093] In some embodiments of the present application, the third sound data group includes fourth sound data and fifth sound data, and step 106 may include the following sub-steps:
[0094] Sub-step 41: Calculate the difference between the first path score of the fourth sound data and the first path score of the fifth sound data to obtain a path score difference;
[0095] Sub-step 42: Calculate the difference between the matching degree of the fourth voice data and the matching degree of the fifth voice data to obtain a matching degree difference;
[0096] Sub-step 43: determining the relationship between the path score difference and the matching degree difference, and screening the third voice data grouping according to the relationship between the path score difference and the matching degree difference to obtain a target voice result.
[0097] The fourth sound data refers to a high matching path between two adjacent groups in the third sound data group, and the fifth sound data refers to a low matching path between two adjacent groups in the third sound data group.
[0098] In some implementations of this embodiment, the third sound data grouping can be screened by combining reverse sampling with forward sampling. Wherein, reverse sampling refers to sampling the matching length from high to low, and forward sampling refers to sampling the matching length from low to high. Assuming that after grouping according to the matching degree, group 0, group 1, group 2, group 3...group 8 can be obtained, where group 0 represents a matching degree of 1, and group 8 represents a matching degree of 8; from low to high, the path score difference and the matching degree difference between group 0 and group 1 are first compared for cutting; then group 1 is compared with group 2 for cutting, and so on. If group 2 is empty when comparing group 1 with group 2, group 1 is compared with group 3. Cutting from high to low is the opposite. In some implementations, the third sound data grouping can be screened by first reverse sampling and then forward sampling; in other implementations, the third sound data grouping can be screened by first forward sampling and then reverse sampling.
[0099] If the following conditions are met, the high matching path is retained and the low matching path is discarded;
[0100] High matching path score - low matching path score > poor matching * bias score.
[0101] For "want some ammonia" in the group with a matching degree of 1 in the above example, assuming the calculated path score is 10, for "want some ammonia ethyl" in the group with a matching degree of 2 in the above example, assuming the calculated path score is 14, if the bias score is 3, then in the implementation method of this embodiment, sampling from high to low: 14-10>(2-1)*3, then the path of "want some ammonia" with a matching degree of 1 can be discarded.
[0102] When the following conditions are met, the low matching path is retained and the high matching path is discarded;
[0103] High matching path score - low matching path score < matching difference * bias score - beam
[0104] Considering that there may be a length difference of 1 to 2 tokens between different target speech data, in some implementations, the threshold beam can be set to 3 times the bias score.
[0105] For the "want some ammonia" in the group with a matching degree of 1 in the above example, assuming the calculated path score is 14, and for the "want some ammonia ethyl" in the group with a matching degree of 2, assuming the calculated path score is 7, if the bias score is 3 and the beam is 3×3=9, then in the implementation method of this embodiment, when sampling from low to high: 7-14<(2-1)*3-9, then the path of "want some ammonia ethyl" with a matching degree of 2 can be discarded.
[0106] Step 107: Determine the target utterance result as the recognition result of the current voice data.
[0107] As can be seen from the above steps, after screening the grouped third vocalization data, the target vocalization result is obtained, and the finally retained target vocalization result is used as the recognition result of the current voice data, and the processing of the next frame is continued until all frames are traversed, and the path with the highest score is selected as the final result.
[0108] In the prior art, top-k sampling is often used for speech recognition. When top-k obtains a new path each time, only the top k paths with the highest scores are retained, and other paths are trimmed. However, since the top-k sampling method is locally optimal rather than globally optimal, it is possible that the correct path prefix may be discarded prematurely due to too low a score in the middle. The embodiments of the present application effectively avoid the following two most common problems in top-k sampling in the prior art through the above speech recognition steps.
[0109] Problem 1: Since the bias score accumulates gradually with the degree of matching. When inferring the first few words at the beginning of the inference, even after adding the current bias score, the path score of the candidate vocalization data containing the hot word prefix is still not high enough. As a result, in subsequent top-k sampling, the candidate vocalization data containing the hot word prefix is trimmed. For example:
[0110] Hot word: aminoethylindole
[0111] Correct result: Aminoethylindole is a commonly used anti-inflammatory drug
[0112] Incorrect result: Catering machine drink flower is a commonly used anti-inflammatory drug
[0113] In top-k sampling in the prior art, when inferring the first word 'ammonia', since the bias score has not been accumulated yet, the path score is not high, and it may be trimmed prematurely. Assuming k = 4 in top-k, the possible current paths and their corresponding path scores may be 'catering': 10 points, 'catering and drinking': 12 points, 'ginseng': 7 points, 'peace': 6 points, 'ammonia': 2 points. Among them, since 'ammonia' matches the hot word and the matching degree is 1, the path score after being biased is 2 + 3 = 5. Since k = 4, 4 paths are retained, resulting in the correct path 'ammonia' being trimmed. In the embodiments of the present application, since multiple vocalization data are grouped according to the matching degree between the candidate vocalization data and the preset hot words in the hot word library, the matching degree of the first word 'ammonia' is 1, and it is not in the same group as other high-score paths with a matching degree of 0, so it is retained.
[0114] Problem 2: In fact, the voice data does not contain hot words, but due to the similarity of the prefix part to the hot words, affected by the bias score, the path score of the candidate vocalization data that matches the hot words is relatively high. When top-k trims, the correct candidate vocalization data is trimmed prematurely. For example:
[0115] Hot words: Children's four-dimensional calcium dry suspension
[0116] Correct identification result: Children's thinking training is beneficial to brain development
[0117] Error identification results: Four-dimensional training for children is beneficial to brain development
[0118] Cause of the error: Assuming k=4 in topk, the current path and path scores may be "Children's Thinking": 8 points, "Two Children's Four": 2 points, and the score after bias is 2+3*3=11, "Children's Four": 4 points, and the score after bias is 4+3*3=13 points, "Two Children's Four Dimensions" score 1, after bias: 1+4*3=13, "Children's Four Dimensions": 2 points, after bias: 2+4*3=14, 4 paths are retained, therefore, the correct path "Children's Thinking" is cut off, and finally the wrong result "Children's Four Dimensional Training is Good for Brain Development" is obtained.
[0119] However, in the embodiment of the present application, since the candidate utterance data are grouped according to the matching degree between the candidate utterance data and the preset hot words in the hot word library, the path containing "children's thinking" (matching degree is 0) and the path containing "children's four dimensions" (matching degree is 4) are not in the same group, so they are retained. Moreover, since children's thinking obtains a higher model score when the next token "training" is added, it is finally retained, and the final correct result "children's thinking training is beneficial to brain development" is obtained. In addition, since the candidate utterance data are grouped, and the candidate utterance data in the group are screened within the group according to their respective first path scores, and then positive and reverse sampling is performed between different groups according to the matching degree, the retained screening results are effectively reduced without having a significant impact on the results, and the amount of calculation is reduced.
[0120] The embodiment of the present application performs streaming and non-streaming performance tests on the medical terminology test set. Streaming refers to real-time processing frame by frame or segment by segment when the voice data arrives, without waiting for the entire voice data to be fully input before processing. It is suitable for scenarios that require real-time response, such as real-time voice recognition, voice assistant, voice translation, etc. Non-streaming refers to the need to wait for the entire voice data to be fully input before performing a one-time processing and outputting the final recognition result. It is suitable for scenarios that do not require real-time response, such as offline voice recognition, voice transcription, etc.
[0121] Test data: medical terms, a total of 1,869 audios, a total duration of 11,330 seconds, and a hot word library with 1,495 hot words; test environment: CPU (Central Processing Unit) single thread, including vad (Voice Activity Detection) module, using streaming and non-streaming models; test parameters: single token bias score 3.5, k in topk sampling is 4. In the present application, when determining the current voice data of the current voice data, k1 is set to 20; when determining the third voice data from the second voice data contained in the second voice data group, k2 is set to 4, that is, there are 4 voice data in each third voice data group; the specific test results are shown in Tables 1 and 2 below. As can be calculated from the data in the table, when the time consumption is not significantly increased, the speech recognition method of the embodiment of the present application is compared with the topk sampling method of the related technology for speech recognition, and effectively reduces the word error rate, among which the non-streaming method is relatively reduced by 29.5%, the streaming method is relatively reduced by 38.2%, and the non-streaming method is 29.5%. At the same time, the speech recognition method of the embodiment of the present application is equivalent to the topk method of the related technology for speech recognition, which improves the recall rate of hot words, among which the non-streaming method is relatively improved by 24.3%, and the streaming method is relatively improved by 40%.
[0122] Table 1: Non-streaming
[0123] Word Error Rate recall time consuming Don’t use buzzwords 5.24% 44.1% 626s Related technology topk 4.34% 63.8% 1601s Embodiments of the present application 3.06% 79.3% 1914s
[0124] Table 2: Flow
[0125] Word Error Rate recall time consuming Don’t use buzzwords 7.77% 33.2% 1215s Related technology topk 7.90% 55.7% 2092s Embodiments of the present application 4.88% 78.0% 2385s
[0126] The embodiment of the present application obtains current voice data, recognized voice data, and a hot word library matching the current context, determines the current voice data corresponding to the current voice data and the model score corresponding to the current voice data according to the recognized voice data and the current voice data, determines the first voice data from the current voice data according to the model score, combines the first voice data with the recognized voice data to obtain the second voice data, determines the matching degree between the second voice data and the preset hot words in the hot word library, and groups the second voice data according to the matching degree to obtain second voice data groups, determines the first path score corresponding to each second voice data in the second voice data group for each second voice data group, determines the third voice data group from the second voice data group according to the first path score, screens the third voice data group to obtain a target voice result, and determines the target voice result as the recognition result of the current voice data, thereby effectively reducing the voice recognition error rate when the voice data contains hot words.
[0127] It should be noted that, for the method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present application are not limited by the described order of actions, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present application.
[0128] Reference Figure 2 , shows a schematic diagram of the structure of a speech recognition device provided by an embodiment of the present application, which may specifically include the following modules:
[0129] The data acquisition module 201 is configured to acquire current voice data, recognized voice data, and a hot word library matching the current context;
[0130] A model score determination module 202 is configured to determine, based on the recognized vocalization data and the current voice data, current vocalization data corresponding to the current voice data and a model score corresponding to the current voice data;
[0131] A first sampling module 203 is configured to determine first vocalization data from the current vocalization data according to the model score, and combine the first vocalization data with the recognized vocalization data to obtain second vocalization data;
[0132] The data bucketing module 204 is configured to determine the matching degree between the second voice data and the preset hot words in the hot word library, and group the second voice data according to the matching degree to obtain second voice data groups;
[0133] The second sampling module 205 is configured to determine, for each of the second sound data groups, a first path score corresponding to each of the second sound data in the second sound data group, and determine a third sound data group from the second sound data group according to the first path score;
[0134] The third sampling module 206 is configured to filter the third sound data group to obtain a target sound result;
[0135] The result determination module 207 is configured to determine the target utterance result as the recognition result of the current voice data.
[0136] In an optional embodiment of the present application, the voice data determination module 202 includes:
[0137] A first feature extraction submodule is configured to extract features from the current voice data to obtain audio features;
[0138] An encoding submodule, configured to input the audio features into a preset encoder to obtain a first feature matrix;
[0139] A decoding submodule, configured to input the recognized sound data into a preset decoder to obtain a second feature matrix;
[0140] The score determination submodule is configured to input the first feature matrix and the second feature matrix into a preset connector to obtain the current voice data corresponding to the current voice data and the model score corresponding to the current voice data.
[0141] In an optional embodiment of the present application, the first sampling module 203 includes:
[0142] A first determining subunit is configured to sort the current sound data according to the model score to obtain sorted current sound data;
[0143] The second determining subunit is configured to determine the first target quantity of first sound data from the sorted current sound data according to a preset first target quantity.
[0144] In an optional embodiment of the present application, the device further includes:
[0145] The hot word detection module is configured to calculate the bias scores of the first utterance data and the recognized utterance data by using a preset hot word detection algorithm.
[0146] In an optional embodiment of the present application, the second sampling module 205 includes:
[0147] a path score acquisition submodule, configured to acquire a second path score of the identified vocal data in the second vocal data;
[0148] The path score calculation submodule is configured to sum the second path score, the model score and the bias score to obtain a first path score corresponding to the second vocalization data.
[0149] In an optional embodiment of the present application, the second sampling module 205 includes:
[0150] A first grouping submodule is configured to sort the second sound data included in the second sound data group according to the first path score to obtain sorted second sound data group;
[0151] The second grouping submodule is configured to select a second target number of second sound data from the sorted second sound data groups according to a preset second target number to obtain a third sound data group.
[0152] In an optional embodiment of the present application, the third sound data group includes fourth sound data and fifth sound data, and the third sampling module 206 includes:
[0153] A first difference calculation module is configured to calculate a difference between a first path score of the fourth sound data and a first path score of the fifth sound data to obtain a path score difference;
[0154] A second difference calculation module is configured to calculate the difference between the matching degree of the fourth voice data and the matching degree of the fifth voice data to obtain a matching degree difference;
[0155] The first screening submodule is configured to determine the relationship between the path score difference and the matching degree difference, and screen the third voice data group according to the relationship between the path score difference and the matching degree difference to obtain a target voice result.
[0156] An embodiment of the present application also provides an electronic device, which may include a processor, a memory, and a computer program stored in the memory and capable of running on the processor, and when the computer program is executed by the processor, the above-mentioned method of speech recognition is implemented.
[0157] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for generating game voice as described above is implemented.
[0158] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0159] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0160] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0161] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the embodiments of the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0162] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0163] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0164] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0165] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.
[0166] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the above elements.
[0167] The provided speech recognition method, device, electronic device and storage medium are introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A speech recognition method, characterized in that: The method comprises: Obtain current voice data, recognized voice data, and a hot word library that matches the current context; Determine, based on the recognized vocalization data and the current voice data, the current vocalization data corresponding to the current voice data and the model score corresponding to the current vocalization data; Determining first vocal data from the current vocal data according to the model score, and combining the first vocal data with the recognized vocal data to obtain second vocal data; Determine the matching degree between the second voice data and the preset hot words in the hot word library, and group the second voice data according to the matching degree to obtain second voice data groups; For each of the second sound data groups, determine a first path score corresponding to each of the second sound data in the second sound data group, and determine a third sound data group from the second sound data group according to the first path score; Screening the third vocalization data group to obtain a target vocalization result; The target utterance result is determined as the recognition result of the current voice data.
2. The method according to claim 1, characterized in that The step of determining, based on the recognized vocalization data and the current voice data, the current vocalization data corresponding to the current voice data and the model score corresponding to the current vocalization data comprises: Extracting features from the current voice data to obtain audio features; Inputting the audio features into a preset encoder to obtain a first feature matrix; Inputting the recognized vocalization data into a preset decoder to obtain a second feature matrix; The first feature matrix and the second feature matrix are input into a preset connector to obtain current voice data corresponding to the current voice data and a model score corresponding to the current voice data.
3. The method according to claim 2, characterized in that Determining first vocalization data from the current vocalization data according to the model score includes: Sorting the current vocalization data according to the model score to obtain sorted current vocalization data; According to a preset first target quantity, a first target quantity of first sound data is determined from the sorted current sound data.
4. The method according to claim 3, characterized in that After determining the first target number of first sound data from the sorted current sound data according to the preset first target number, the method includes: The bias scores of the first utterance data and the recognized utterance data are calculated by using a preset hot word detection algorithm.
5. The method according to claim 4, characterized in that The step of determining, for each second sound data group, a first path score corresponding to each second sound data in the second sound data group includes: Obtaining a second path score of the identified vocal data in the second vocal data; The second path score, the model score and the bias score are summed to obtain a first path score corresponding to the second vocalization data.
6. The method according to any one of claims 1 to 5, characterized in that: The step of determining a third sound data group from the second sound data group according to the first path score includes: sorting the second sound data included in the second sound data group according to the first path score to obtain sorted second sound data group; A second target number of second sound data is selected from the sorted second sound data groups according to a preset second target number to obtain a third sound data group.
7. The method according to any one of claims 1 to 5, characterized in that: The third sound data group includes fourth sound data and fifth sound data, and the third sound data group is screened to obtain a target sound result, including: Calculating a difference between the first path score of the fourth sound production data and the first path score of the fifth sound production data to obtain a path score difference; Calculating a difference between a matching degree of the fourth voice data and a matching degree of the fifth voice data to obtain a matching degree difference; The relationship between the path score difference and the matching degree difference is determined, and the third voice data group is screened according to the relationship between the path score difference and the matching degree difference to obtain a target voice result.
8. A speech recognition device, characterized in that: The device comprises: A data acquisition module is configured to acquire current voice data, recognized voice data, and a hot word library matching the current context; A model score determination module is configured to determine, based on the recognized vocalization data and the current voice data, current vocalization data corresponding to the current voice data and a model score corresponding to the current voice data; A first sampling module is configured to determine first vocalization data from the current vocalization data according to the model score, and combine the first vocalization data with the identified vocalization data to obtain second vocalization data; A data bucketing module is configured to determine a matching degree between the second voice data and a preset hot word in a hot word library, and group the second voice data according to the matching degree to obtain a second voice data group; a second sampling module configured to determine, for each of the second sound data groups, a first path score corresponding to each of the second sound data in the second sound data group, and determine a third sound data group from the second sound data group based on the first path score; A third sampling module is configured to filter the third sound data group to obtain a target sound result; The result determination module is configured to determine the target utterance result as the recognition result of the current voice data.
9. An electronic device, characterized in that: The invention comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the speech recognition method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the speech recognition method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Speech recognition method, device and equipment and storage medium
CN110164416A
Speech recognition method and device, equipment and storage medium
CN114360499A
Speech recognition method and related product thereof
CN115312041A
Cross-lingual speech recognition
US20200111484A1
Cited By
Medicine name identification method and device based on voice information and electronic equipment
CN120783752A