Speech recognition method and device, electronic equipment and storage medium
By decoding and matching the target decoding path of the speech frame, combining the path score and hot word score, the target path of the speech frame is determined, which solves the problem of poor recognition accuracy in traditional speech recognition hot word enhancement technology, and achieves higher speech signal recognition accuracy and scene customization capabilities.
Patent Information
- Application Number
- CN202311824194.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
AI Technical Summary
In traditional pronunciation hot word enhancement technology, the hot word recognition effect is limited, resulting in poor recognition accuracy.
By decoding the target decoding path of the speech frame, multiple candidate paths and path scores are obtained. Combined with the target hot words, the reserved path is determined from the candidate path, including the score matching path of the top N of the path score and the hot word matching path matching path matching the target hot word, and the path score is updated based on the hot word score of the target hot word in the preset hot word library to finally determine the target path of the speech frame.
By combining the path score and the hot word score, the determination of the hot word matching path is optimized, the performance of hot word recognition is improved, the scene customization ability in speech recognition is enhanced, and the recognition accuracy of the voice signal is improved.
Smart Images

Figure CN120220685A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition technology, and particularly to a speech recognition method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] Speech recognition technology is a technology that can convert audio into text or commands. Recognition accuracy is an important dimension for evaluating the quality of speech recognition technology. Usually, there are differences in the content field preferences for recognition in different business scenarios. Words that often appear in the application field or customized preference words, which can also be called hot words. In actual application scenarios, the influence of hot words can be considered during the speech recognition process to improve the recognition accuracy.
[0003] However, in traditional speech recognition hot word enhancement technologies, due to the limited hot word recognition effect, the recognition accuracy is relatively poor. Summary of the Invention
[0004] Embodiments of this application provide a speech recognition method, apparatus, electronic device, and computer-readable storage medium, which can improve the recognition accuracy.
[0005] In a first aspect, this application provides a speech recognition method. The method is used to recognize speech frames in a speech signal, and includes:
[0006] Decoding the speech frame according to the target decoding path of the speech frame to obtain multiple candidate paths of the speech frame and corresponding path scores. Each candidate path corresponds to a path score, and the target decoding path is any target path of the previous speech frame adjacent to the speech frame;
[0007] Determining a retained path from the multiple candidate paths according to the path score and the target hot word. The retained path includes the score matching paths with the top N path scores in terms of ranking and the hot word matching paths that match the target hot word. N is a positive integer, and the target hot word is determined from a preset hot word library according to the target decoding path;
[0008] Updating the path score of the retained path according to the hot word score of the target hot word in the preset hot word library to obtain an updated path score;
[0009] Determining the target path corresponding to the target decoding path of the speech frame from the retained path according to the updated path score.
[0010] In one of the embodiments, the determining method for the hot word matching path includes:
[0011] Determine an initial matching path that matches the target hot word from the multiple candidate paths;
[0012] Update the path score of the initial matching path according to the hot word score of the target hot word to obtain the path score of the updated initial matching path;
[0013] Determine the initial matching paths among the top M in the path score ranking of the updated initial matching path as the hot word matching paths, where M is a positive integer.
[0014] In one embodiment, the target hot word includes at least one of a potential hot word and a homophone hot word, and the method for determining the hot word matching path includes:
[0015] Determine the candidate paths that contain the potential hot word from the multiple candidate paths as potential matching paths;
[0016] Determine the candidate paths that match the homophone hot word from the multiple candidate paths as homophone matching paths;
[0017] Obtain the hot word matching path according to at least one of the potential matching path and the homophone matching path.
[0018] In one embodiment, the method for determining the homophone hot word includes:
[0019] For the target decoding path, determine, from the other decoding paths of the speech frames except the target decoding path, the path whose penultimate word unit is the same as the last word unit of the target decoding path as the homophone path;
[0020] Determine, from the preset hot word library, the target word unit that is homophonous with the last word unit of the homophone path, and use the hot word corresponding to the target word unit as the homophone hot word.
[0021] In one embodiment, the method for determining the potential hot word includes:
[0022] Determine, from the preset hot word library, the hot word that matches the last word unit of the target decoding path as the potential hot word.
[0023] In one embodiment, the step of determining, from the preset hot word library, the hot word that matches the last word unit as the potential hot word includes:
[0024] Determine the hot word that matches the last word unit from the preset hot word library as the candidate hot word;
[0025] From the candidate hot words, determine the hot words with hot word scores greater than the preset score as potential hot words.
[0026] In a second aspect, the present application further provides a voice recognition device. The device is used to recognize speech frames in a speech signal, and includes:
[0027] A decoding module, configured to decode the speech frame according to the target decoding path of the speech frame, to obtain multiple candidate paths of the speech frame and corresponding path scores, each candidate path corresponding to a path score, and the target decoding path being any target path of the previous speech frame adjacent to the speech frame;
[0028] A first determination module, configured to determine a reserved path from the multiple candidate paths according to the path scores and the target hot word, the reserved path including score matching paths ranked top N in terms of path scores and hot word matching paths matching the target hot word, where N is a positive integer, and the target hot word is determined from a preset hot word library according to the target decoding path;
[0029] An update module, configured to update the path scores of the reserved paths according to the hot word scores of the target hot words in the preset hot word library, to obtain updated path scores;
[0030] A second determination module, configured to determine the target path corresponding to the target decoding path of the speech frame from the reserved paths according to the updated path scores.
[0031] In one embodiment, the device includes a determination module for hot word matching paths. The determination module for hot word matching paths includes:
[0032] A first determination sub-module, configured to determine an initial matching path matching the target hot word from the multiple candidate paths according to the target hot word;
[0033] A first update sub-module, configured to update the path score of the initial matching path according to the hot word score of the target hot word, to obtain the path score of the updated initial matching path;
[0034] A second determination sub-module, configured to determine the initial matching paths ranked top M in terms of the path scores of the updated initial matching paths as hot word matching paths, where M is a positive integer.
[0035] In one embodiment, the target hot word includes at least one of a potential hot word and a homophone hot word. The device includes a determination module for hot word matching paths. The determination module for hot word matching paths includes:
[0036] A third determination sub-module, configured to determine, from the multiple candidate paths, the candidate path that contains the potential hot word as the potential matching path;
[0037] A fourth determination sub-module, configured to determine, from the multiple candidate paths, the candidate path that matches the homophone hot word as the homophone matching path;
[0038] A fifth determination sub-module, configured to obtain a hot word matching path according to at least one of the potential matching path and the homophone matching path.
[0039] In one embodiment, the apparatus further includes a determination module for homophone hot words, and the determination module for homophone hot words includes:
[0040] A sixth determination sub-module, configured to, for the target decoding path, determine, from other decoding paths of the speech frames except the target decoding path, the path whose penultimate word unit is the same as the last word unit of the target decoding path as the homophone path;
[0041] A seventh determination sub-module, configured to determine, from the preset hot word library, a target word unit that is homophonous with the last word unit of the homophone path, and use the hot word corresponding to the target word unit as the homophone hot word.
[0042] In one embodiment, the apparatus further includes a determination module for potential hot words, and the determination module for potential hot words includes:
[0043] An eighth determination sub-module, configured to determine, according to the last word unit of the target decoding path, a hot word that matches the last word unit from the preset hot word library as the potential hot word.
[0044] In one embodiment, the eighth determination sub-module includes:
[0045] A first determination unit, configured to determine a hot word that matches the last word unit from the preset hot word library as a candidate hot word;
[0046] A second determination unit, configured to determine, from the candidate hot words, a hot word whose hot word score is greater than a preset score as the potential hot word.
[0047] In a third aspect, the present application further provides an electronic device. The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the embodiments of the present disclosure are implemented.
[0048] Fourthly, the present application also provides a computer-readable storage medium. On the computer-readable storage medium, there is a computer program stored, and when the computer program is executed by a processor, the steps of the method described in any one of the embodiments of the present disclosure are implemented.
[0049] Fifthly, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of the embodiments of the present disclosure are implemented.
[0050] The above speech recognition method, device, electronic device, computer-readable storage medium, and computer program product are used to recognize speech frames of a speech signal. For the target decoding path of a speech frame, the speech frame is decoded to obtain multiple candidate paths and corresponding path scores, and a retained path is determined according to the path scores and the target hotword. The retained path includes the score matching paths with the top N path scores and the hotword matching paths that match the target hotword. According to the hotword score of the target hotword in the preset hotword library, the path scores of the retained path are updated to obtain the updated path scores, and according to the updated path scores, the target path corresponding to the speech frame and the target decoding path is determined from the retained path. Since in this solution, when selecting the target path from the candidate paths, the retained path is first determined according to the path scores and the target hotword, which can take into account the influence of both the path scores and the hotword on the recognition effect, retain the score matching paths and the hotword matching paths before the path scores are updated, and reduce the probability that the hotword matching paths with higher hotword scores are missed during the screening process of the retained path; then the path scores are updated according to the hotword score of the target hotword, and the target path is determined from the retained path according to the updated path scores, which can enhance the hotword matching paths through the hotword scores, improve the hit probability of the hotword matching paths in the target path, optimize the recognition performance of the hotword, effectively enhance the scene customization ability in speech recognition, ensure the recognition accuracy of each speech frame, and thus effectively improve the recognition accuracy of the speech signal; and during the recognition process of the speech signal, the selection method for determining the target path from the candidate paths is optimized and adjusted, without adjusting the recognition result search space composed of the target paths of each speech frame, balancing the hotword recognition performance and the decoding efficiency, not adding extra decoding time-consuming to ensure the efficiency of the decoding process, effectively improving the hotword recognition effect, and thus improving the accuracy of the speech signal recognition result, effectively enhancing the scene customization ability in speech recognition. Description of the Drawings
[0051] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0052] Figure 1 It is a schematic flowchart of a speech recognition method in an embodiment;
[0053] Figure 2 It is a schematic diagram of a hot word path and hot word scores in an embodiment;
[0054] Figure 3 It is a schematic flowchart of a method for determining a hot word matching path in an embodiment;
[0055] Figure 4 It is a schematic flowchart of a method for determining a hot word matching path in another embodiment;
[0056] Figure 5 It is a schematic flowchart of a method for determining homophonic hot words in an embodiment;
[0057] Figure 6 It is a schematic flowchart of a method for determining potential hot words in an embodiment;
[0058] Figure 7 It is a schematic flowchart of a speech recognition method in another embodiment;
[0059] Figure 8 It is a structural block diagram of a speech recognition device in an embodiment;
[0060] Figure 9 It is an internal structure diagram of an electronic device in an embodiment. Detailed implementation manners
[0061] To make the objectives, technical solutions and advantages of the present application more clear, the following further details the present application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0062] In one embodiment, as Figure 1 shown, a speech recognition method is provided. This method is illustrated by taking a terminal as an example. It can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, this method is used to recognize speech frames in a speech signal, and includes steps S110 to S150.
[0063] Step S110: Decode the speech frame according to the target decoding path of the speech frame to obtain multiple candidate paths of the speech frame and corresponding path scores. Each candidate path corresponds to a path score, and the target decoding path is any target path of the previous speech frame adjacent to the speech frame.
[0064] Among them, the speech signal may include the directly collected original speech signal, or may include the speech signal obtained by processing the original speech signal through signal processing techniques. The signal processing techniques include, but are not limited to, speech signal processing methods such as filtering processing and enhancement processing. Usually, the speech signal can be a continuous time signal. When performing speech recognition, it is necessary to divide the speech signal into speech frames. The method of dividing the speech signal into speech frames can be determined according to the actual application scenario. For example, the speech signal can be directly divided into multiple signal segments of a fixed time length according to the time length to obtain speech frames; further, speech frames can also be obtained by combining methods such as windowing, Fourier transform, and feature extraction for subsequent analysis and processing.
[0065] In this embodiment, the speech frame is decoded according to the target decoding path of the speech frame to obtain multiple candidate paths of the speech frame and the corresponding path scores. Exemplarily, the target decoding path is any target path of the previous speech frame adjacent to the speech frame. When recognizing a speech signal, the speech frames of the speech signal are decoded in sequence to obtain the target path of each speech frame. Each time, decoding can be performed according to any target path of the previous adjacent speech frame. In one example, each decoding path of the speech frame can be regarded as a beam, that is, a decoding score. The speech frame is decoded according to the decoding path to obtain multiple tokens, that is, the corresponding probability distribution. Combining the multiple tokens and the corresponding beams obtains multiple candidate paths, and combining the probability distribution and the path scores of the beams obtains the path scores corresponding to the candidate paths. In one example, when the speech frame to be recognized is the first frame of the speech signal, the target decoding path can be set to be empty, and multiple candidate paths and the corresponding path scores are directly determined according to the feature information of the speech frame. Optionally, the path scores corresponding to the candidate paths match the probabilities of the candidate paths. Among them, the probability distributions of multiple candidate paths can be determined through decoding, and the path scores corresponding to each candidate path can be obtained directly or indirectly according to the probability distributions of the multiple candidate paths. Generally, the greater the probability of the candidate path, the closer the candidate path is considered to be to the correct recognition result of the speech frame under the current target decoding path. In one example, the candidate path is determined based on the target decoding path and the possible recognition results of the currently decoded speech frame. For example, if the target decoding path is "Beijing", the possible recognition results of the currently decoded speech frame can include "city", "people", "drama", etc., and the corresponding multiple candidate paths can include "Beijing City", "Beijing People", "Beijing Drama", etc. In one example, the possible recognition results of the speech frame can be determined based on a preset dictionary library. The preset dictionary library is set according to the actual application scenario. For example, the possible recognition results can be determined from the preset dictionary library according to the feature data of the speech frame to obtain multiple candidate paths; the possible recognition results can be determined from the preset dictionary library according to the target decoding path to obtain multiple candidate paths; or the possible recognition results can be determined from the preset dictionary library by combining the feature data of the speech frame and the target decoding path to obtain multiple candidate paths. When determining the path scores of the candidate paths, the probability distributions of multiple candidate paths can be determined through decoding, so as to determine the path scores of each candidate path. Among them, the decoding process can be implemented based on a preset decoding algorithm. The preset decoding algorithm can include one or more of speech decoding algorithms such as, but not limited to, the greedy search algorithm, the beam search algorithm, and the Viterbi algorithm, and can be specifically set according to the actual application scenario. The present disclosure does not limit this. In one example, the speech frame can be decoded by constructing a speech recognition model.
[0066] Step S120: Determine the retained paths from multiple candidate paths according to the path scores and the target hot words. The retained paths include the score matching paths with the top N path scores and the hot word matching paths that match the target hot words. N is a positive integer, and the target hot word is determined from a preset hot word library according to the target decoding path.
[0067] Among them, the target hot word can be obtained from the hot word library, and different hot word libraries can correspond to different application scenarios and different recognition contents. In some possible implementation manners, the preset hot word library can be determined according to the signal source of the voice signal to be recognized. For example, if the source of the voice signal is a social platform, the preset hot word library can include the hot word library corresponding to the social field; the preset hot word library can be determined according to the service identifier corresponding to the recognized service scenario. For example, different service identifiers can be set for different service scenarios in actual applications, and the corresponding hot word libraries are marked with the service identifiers. When voice recognition is required, the hot word library corresponding to the service identifier is determined as the preset hot word library according to the service identifier of the service scenario of the voice recognition. In one example, the preset hot word library can also be adaptively adjusted according to the actual application scenario. For example, in the case of strong server performance, strong terminal data processing capabilities, and high requirements for recognition accuracy, the preset hot word library can be expanded and enriched to improve the hot word recognition effect of voice recognition as much as possible, thereby ensuring the recognition accuracy; in the case of poor server performance, limited terminal data processing capabilities, and high requirements for recognition efficiency, the hot words in the preset hot word library can be appropriately screened, and the hot words with a greater association with the scenario are retained. Among them, when screening, the hot word scores can be used for screening, taking into account the recognition efficiency while ensuring the hot word recognition effect. In one example, the preset hot word library can also be updated according to the change of the vocabulary usage preference in the specific application scenario. For example, hot words are mined periodically from the vocabulary library corresponding to the application scenario and the preset hot word library is updated.
[0068] The target hot words may include some or all of the hot words in the preset hot word library. The target hot words are determined from the preset hot word library according to the target decoding path. In some examples, the determination method of the target hot words can be determined according to the actual application scenario. For example, the target hot words can be directly determined according to the words or word units included in the target decoding path; the target hot words can be determined according to the relationships between the words or word units included in the target decoding path; the target hot words can also be determined according to the association relationship between the target decoding path and other decoding paths. In one example, the number of target hot words can be adaptively adjusted according to the actual application scenario. For example, in the case of strong server performance, strong terminal data processing capabilities, and high requirements for recognition accuracy, when determining the target hot words from the preset hot word library, as many target hot words as possible can be retained; in the case of poor server performance, limited terminal data processing capabilities, and high requirements for recognition efficiency, when determining the target hot words from the preset hot word library, the target number or the hot words within the target number can be screened according to the actual needs as the target hot words. Among them, when screening, it can be screened according to the hot word score, the association relationship between the hot word and the candidate path, etc., which can be specifically determined according to the actual application scenario, and the target number can be set according to the actual application scenario. In this embodiment, each hot word corresponds to a hot word score, and different hot words can correspond to different hot word scores. Among them, the hot word score is related to the correlation degree between the hot word and the scenario, and the correlation degree between the hot word and the scenario can be related to many factors such as the frequency of the hot word appearing in the specific application scenario, the correlation degree between the hot word and the application scenario field, etc. In one example, the hot word score corresponding to the hot word can also be updated and modified according to the actual application scenario. In one example, according to the preset hot word library, a finite weighted state transition machine is used to represent all the hot word paths and the corresponding hot word scores, where a finite weighted state transition machine corresponding to the preset hot word library is constructed based on the preset hot word library, and is used to identify the hot word paths and the corresponding hot word scores. Figure 2 FIG. is a schematic diagram of a hot word path and a hot word score shown according to an exemplary embodiment. Refer to Figure 2 As shown, the hot words include "Peking Opera", "Kyoto", "Jingcheng", "strong", "robust", and the number on each edge is the score incentive for matching the corresponding word unit. The hot word score of each hot word can be determined according to the numbers on the path of the hot word. For example, the hot word score of "Jingcheng" is 7.
[0069] In this embodiment, a retained path is determined from multiple candidate paths according to the path score and the target hot word. The retained path includes a score-matching path and a hot-word matching path. Among them, the path score is associated with the probability of the candidate path. Therefore, candidate paths with higher probabilities can be screened out as score-matching paths through the path score. Optionally, the top N candidate paths in terms of path score are determined as score-matching paths, where N is a positive integer. When ranking the score-matching paths according to the path score, the ranking can be performed in descending order of the probability of the candidate path. In one example, the path score is positively correlated with the probability of the candidate path, and multiple candidate paths can be arranged in descending order of the path score, and the top N candidate paths are determined as score-matching paths, where N is a positive integer. N can be determined according to the actual application scenario, and can be a fixed value or a variable value. For example, candidate paths with path scores greater than a preset threshold can be determined as score-matching paths, and the number of candidate paths with path scores greater than the preset threshold is N. It can be understood that in this embodiment, the probability of the screened score-matching path is greater than that of other paths in the candidate paths. The number N of score-matching paths can be determined according to the actual application scenario. For example, the number of score-matching paths can be determined according to the data processing capacity and recognition requirements of the actual application scenario.
[0070] Hot words can reflect the vocabulary preferences in the current speech recognition scenario. Therefore, candidate paths that are more closely associated with the current speech recognition scenario can be screened out as hot-word matching paths through the target hot word. The method for determining whether a candidate path matches the target hot word can include one or more of the following. For example, it can be determined whether there is a match by whether the target hot word is included in the candidate path; it can be determined whether there is a match by the degree of association between the candidate path and the target hot word, such as homophony, near homophony, etc., and a greater degree of association can be considered; it can also be determined whether there is a match by combining different methods. The target hot word can include hot words of different categories, and the matching method may also be different for different categories of hot words. For example, when the target hot word is a potential hot word, if the candidate path contains the target hot word, it can be considered that the candidate path matches the target hot word; when the target hot word is a homophonic hot word, if a preset word unit in the candidate path, such as the last word unit, contains the homophonic word unit in the homophonic hot word, it can be considered that the candidate path matches the target hot word.
[0071] After determining the score matching path and the hot word matching path, according to the score matching path and the hot word matching path, a retained path is obtained. Among them, the union of the score matching path and the hot word matching path can be directly determined as the retained path; the union of the score matching path and the hot word matching path can be obtained first, and then further screened in combination with other factors to obtain the retained path; or the retained path can be determined according to the score matching path and the hot word matching path in other possible ways, and specifically can be determined according to the actual application scenario. In some possible implementation manners, when there is no same path in the score matching path and the hot word matching path, the union of the obtained score matching path and the hot word matching path is all the score matching paths and all the hot word matching paths; when there is a same path in the score matching path and the hot word matching path, only one same path (such as the hot word matching path) needs to be retained. For example, if the score matching path includes "Beijing City" and the hot word matching path includes "Beijing City", when determining the retained path, only the path of "Beijing City" needs to be retained.
[0072] Step S130, update the path score of the retained path according to the hot word score of the target hot word in the preset hot word library to obtain the updated path score.
[0073] In the embodiments of the present disclosure, the path score of the retained path is updated according to the hot word score of the target hot word to obtain the updated path score. In some possible implementation manners, when updating the path score of the retained path according to the hot word score of the target hot word, the path score of the hot word matching path that matches the target hot word in the retained path can be updated. In some examples, the path score can be directly updated according to the sum of the path score and the hot word score; the hot word matching path in the retained path can be weighted according to the hot word score to obtain the updated path score; or the path score can be updated according to other possible calculation methods to obtain the updated path score associated with the hot word score. In one example, when updating the path score of the retained path according to the hot word score, the path score of the hot word matching path is updated, which further improves the effectiveness of path score update and the probability of the subsequent target path hitting the hot word, and thus improves the hot word recognition performance.
[0074] In a possible implementation, the path score of the retained path is updated in combination with the hot word scores of other hot words in the preset hot word library. When updating, the hot word score may include the hot word path score, and each word unit in the hot word corresponds to a unit score. When updating the path score of the retained path, when the last word unit of the retained path corresponds to a word unit in the preset hot word library, the path score is updated according to the unit score of this word unit. Further, the hot word corresponding to this word unit in the hot word library is the first hot word. When the word unit before the last word unit of the retained path matches the word unit before the corresponding word unit in the first hot word, the path score is increased by the unit score of this word unit to obtain the updated path score; when the last word unit of the retained path does not match the word unit in the preset hot word library, but there are other matching word units adjacent to the last word unit in the retained path that match the word unit in the preset hot word library, the path score is decreased by the unit score of the other matching word units to obtain the updated path score. It can be understood that for the same path, both addition and subtraction operations can be performed simultaneously.
[0075] Step S140, according to the updated path score, determine the target path corresponding to the target decoding path of the speech frame from the retained path.
[0076] In the embodiments of the present disclosure, after the path score is updated, according to the updated path score, determine the target path corresponding to the target decoding path from the retained path, where the number of target paths can be set according to the actual application scenario, and different application scenarios may correspond to different numbers of target paths. In one example, all the retained paths of the speech frame can be obtained by combining the retained paths of other decoding paths, and all the retained paths are sorted according to the updated path score, and a preset number of retained paths are selected from the sorted retained paths as the target paths, where the target paths selected from the retained paths of the target decoding path are the target paths corresponding to the target decoding path. In some application scenarios, the target path corresponding to the speech frame can participate in the decoding process of the next speech frame, and the number of target paths will affect the efficiency of the decoding process. Therefore, the preset number can be determined by comprehensively considering the device performance, efficiency requirements, and accuracy requirements in the actual application scenario as the number of target paths.
[0077] In a possible implementation, the recognition result of the speech signal is determined according to the target path of each speech frame in the speech signal, where the target path of each speech frame in the speech signal is the path with a relatively high probability and a relatively high hot word hit probability obtained after the above steps. Through the target path of each speech frame, the path tree corresponding to the speech signal can be obtained. In a possible implementation, after obtaining the target path of each speech frame in the speech signal, the path score corresponding to each total path of the speech signal can be directly determined according to the updated path score in the target path determination process, and the target total path is obtained according to the path score of each total path to determine the recognition result of the speech signal; further, after obtaining the target path of each speech frame in the speech signal, other dimensions such as hot word scores can also be combined to further update and adjust the path score, and the target total path is determined from the path tree corresponding to the speech signal according to the adjusted score to obtain the recognition result of the speech signal. Exemplarily, according to the target path of each speech frame, the target path of the last speech frame in the speech signal can be determined, and the target path with the highest updated path score is selected from the target path of the last speech frame as the recognition result of the speech signal; further, other factors that may affect the recognition result, such as noise interference, can also be combined to adjust the updated path score of the target path of the last speech frame, and the target path with the highest adjusted path score is selected as the recognition result of the speech signal. In one embodiment, the method described in this embodiment can be independently deployed on a terminal or a server, ensuring the flexibility of speech recognition.
[0078] Embodiments of the present disclosure are used to recognize speech frames of a speech signal. For a target decoding path of a speech frame, the speech frame is decoded to obtain multiple candidate paths and corresponding path scores, and a retained path is determined according to the path scores and a target hotword. The retained path includes score matching paths ranked top N in terms of path scores and hotword matching paths that match the target hotword. According to the hotword score of the target hotword in a preset hotword library, the path scores of the retained paths are updated to obtain updated path scores, and according to the updated path scores, a target path corresponding to the speech frame and the target decoding path is determined from the retained paths. Since in this solution, when selecting a target path from candidate paths, the retained path is first determined according to the path scores and the target hotword, which can take into account the influence of both path scores and hotwords on the recognition effect, retain score matching paths and hotword matching paths before updating the path scores, and reduce the probability of missing hotword matching paths with higher hotword scores during the screening process of the retained paths; then the path scores are updated according to the hotword scores of the target hotword, and the target path is determined from the retained paths according to the updated path scores, which can enhance the hotword matching paths through the hotword scores, improve the hit probability of the hotword matching paths in the target path, optimize the recognition performance of hotwords, effectively enhance the scene customization ability in speech recognition, ensure the recognition accuracy of each speech frame, and thus effectively improve the recognition accuracy of the speech signal; and during the recognition process of the speech signal, the selection method for determining the target path from candidate paths is optimized and adjusted, without adjusting the recognition result search space composed of the target paths of each speech frame, balancing the hotword recognition performance and decoding efficiency, not adding extra decoding time-consuming to ensure the decoding process efficiency, while effectively improving the hotword recognition effect, and thus improving the accuracy of the speech signal recognition result and effectively enhancing the scene customization ability in speech recognition.
[0079] In one embodiment, the hotword matching path can be determined by combining the path score and the hotword score. As Figure 3 shown, the determination method of the hotword matching path includes:
[0080] Step S310, according to the target hotword, determine an initial matching path that matches the target hotword from multiple candidate paths;
[0081] Step S320, according to the hotword score of the target hotword, update the path score of the initial matching path to obtain the updated path score of the initial matching path;
[0082] Step S330, determine the first M initial matching paths with the highest path scores among the updated initial matching paths as the hotword matching paths, where M is a positive integer.
[0083] In the embodiments of the present disclosure, when determining a hot word matching path from multiple candidate paths, screening can be performed according to the path score and the hot word score to further optimize the hot word matching path. Exemplarily, an initial matching path matching the target hot word is determined from multiple candidate paths. When performing path matching according to the target hot word, the hot word matching method described in the foregoing embodiments can be referred to, which will not be elaborated herein. After the initial matching path is determined, further screening is performed according to the path score and the hot word score, and a hot word matching path is determined from the initial matching paths. In one example, the path score of the initial matching path can be updated according to the hot word score to perform hot word enhancement, and an updated path score is obtained. The update method includes, but is not limited to, adding points and weighting based on the hot word score. The hot word matching path is determined from the initial matching paths according to the updated path score. Optionally, the initial matching paths with the top M path scores among the updated initial matching paths are determined as the hot word matching paths, where M is a positive integer and M can be determined according to the actual application scenario. When ranking the initial matching paths, they are sorted from high to low according to the probability of the path. In one possible implementation, the path score is positively correlated with the probability of the path, and the hot word score is positively correlated with the relevance between the hot word and the current scenario. The top M paths can be determined from the initial matching paths in the order of the updated path scores from high to low as the hot word matching paths, and M can be determined according to the actual application scenario. It can be understood that in this embodiment, the determined hot word matching path is screened from the initial matching paths after comprehensively considering the probability of the path and the relevance between the hot word and the scenario.
[0084] In the embodiments of the present disclosure, when determining the hot word matching path, the hot word matching path is screened from all the initial matching paths matching the target hot word according to the hot word score and the path score, which optimizes the determination process of the hot word matching path, screens out some initial matching paths whose path scores and hot word scores do not meet the conditions, reduces the subsequent data processing workload, and improves the data processing efficiency. Therefore, in the process of retaining the path determination, while retaining the hot word matching path to optimize the hot word recognition performance, it also ensures the efficiency of subsequent target path and recognition result determination, taking into account both recognition accuracy and recognition efficiency. Voice recognition can be performed according to actual needs in different application scenarios, and it is applicable to more application scenarios.
[0085] In one embodiment, in order to be able to more flexibly adapt to various speech recognition scenarios, multiple types of hot words can be set, such as Figure 4 shown, the target hot word includes at least one of a potential hot word and a homophone hot word, and the method for determining the hot word matching path includes:
[0086] Step S410, determining a candidate path including a potential hot word as a potential matching path from multiple candidate paths;
[0087] Step S420: Determine, from multiple candidate paths, the candidate path that matches the homophone hot word as the homophone matching path;
[0088] Step S430: Obtain a hot word matching path according to at least one of the potential matching path and the homophone matching path.
[0089] In the embodiments of the present disclosure, the target hot word includes at least one of the potential hot word and the homophone hot word. When determining the hot word matching path, the hot word matching path is obtained according to at least one of the potential matching path matched by the potential hot word and the homophone matching path matched by the homophone hot word. Optionally, determine, from multiple candidate paths, the candidate path that includes the potential hot word as the potential matching path. Wherein, when the hot word is a potential hot word, the candidate path hits the hot word (the candidate path includes the hot word itself), and then it can be considered that the candidate path matches the hot word. In one example, the potential hot word can be determined based on multiple candidate paths of the current speech frame; in another example, the potential hot word can be determined based on the last word unit of the target decoding path. For example, if the last word unit of the target decoding path matches the word unit in the preset hot word library, determine the hot word including the word unit in the preset hot word library as the potential hot word. Further, when the last word unit of the potential hot word does not match the last word unit of the target decoding path, when the word unit in the potential hot word is not the first word unit, the other word units before the word unit also match the other word units adjacent to the last word unit. In one possible implementation manner, it is also possible to determine the candidate path whose last word is the potential hot word as the potential matching path to further optimize the potential matching path.
[0090] Determine, from multiple candidate paths, the candidate path that matches the homophone hot word as the homophone matching path. Wherein, the homophone hot word can be a hot word that the candidate path does not directly hit but hits some of its word units. When the candidate path includes a specific word unit in the homophone hot word, it can be considered that the candidate path matches the homophone hot word. In one example, the homophone hot word can be determined based on the potential hot word. For example, the homophone hot word includes the homophone of the potential hot word; in another example, the homophone hot word can be determined based on the word units in the target decoding path and other decoding paths. In one possible implementation manner, it is also possible to determine the candidate path whose last word unit is the homophone unit in the homophone hot word as the homophone matching path. In an exemplary embodiment, the homophone unit in the homophone hot word can be determined according to the relationship between the word unit in the hot word and the current multiple candidate paths.
[0091] Determine a hot word matching path according to at least one of a potential matching path and a homophone matching path, where the hot word matching path can be determined separately according to the potential matching path or the homophone matching path; the union of the potential matching path and the homophone matching path can be directly used as the hot word matching path; or after taking the union of the potential matching path and the homophone matching path, the paths in the union are further processed and filtered to obtain the hot word matching path, which can be specifically determined according to the actual application scenario.
[0092] In the embodiments of the present disclosure, the target hot words include at least one of potential hot words and homophone hot words. The potential matching path including the potential hot words and the homophone matching path matching the homophone hot words are respectively determined, and a hot word matching path is obtained according to at least one of the potential matching path and the homophone matching path; by considering different hot word matching situations, the accuracy and comprehensiveness of the obtained hot word matching path are ensured, the subsequent hot word recognition performance is improved, and by enriching the determination method of the hot word matching path, the hot word hit probability in the target path is further increased, thereby effectively improving the hot word recognition accuracy rate when finally recognizing a speech signal, and improving the scene-based customized recognition ability of speech recognition in different scenarios, and being applicable to more application scenarios.
[0093] In one embodiment, the homophone hot words are determined according to the target decoding path, as Figure 5 shown. The determination method of the homophone hot words includes:
[0094] Step S510, for the target decoding path, determine, from the other decoding paths of the speech frame except the target decoding path, a path where the penultimate word unit is the same as the last word unit of the target decoding path as the homophone path;
[0095] Step S520, determine, from the preset hot word library, a target word unit that is homophonous with the last word unit of the homophone path, and use the hot word corresponding to the target word unit as the homophone hot word.
[0096] In the embodiments of the present disclosure, the homophonic hot words can be determined based on the target decoding path. Optionally, for the target decoding path, a path in which the penultimate word unit is the same as the last word unit of the target decoding path is determined as the homophonic path from other decoding paths of the speech frame, where the target decoding path is any target path of the previous speech frame, and other decoding paths are other target paths in the target paths of the previous speech frame except this target decoding path. For example, when the target decoding path is "Beijing", a path in which the penultimate word unit is "Jing" is determined as the homophonic path from other paths in the decoding paths of the speech frame except "Beijing", such as "Jingqiang". According to the last word unit of the homophonic path, the hot word corresponding to the target word unit is determined as the homophonic hot word from the preset hot word library, where the target word unit and the last word unit of the homophonic path are homophones. In one possible implementation, all word units that are homophones of the last word unit of the homophonic path are determined as the target word units from the preset dictionary library, and the hot words containing the target word units are determined as the homophonic hot words from the preset hot word library; in another possible implementation, combining the preset dictionary library and the preset hot word library, word units that can match the word units in the preset hot word library and are homophones of the last word unit of the homophonic path are determined as the target word units from the preset dictionary library, and the hot words containing the target word units are determined as the homophonic hot words from the preset hot word library. Further, in one embodiment, the word unit in the homophonic hot word that hits the target word unit is other word units except the last word unit, so as to filter out the homophonic matching paths that may hit the hot words in the subsequent decoding process as the reserved paths. In one possible implementation, the speech frame corresponds to multiple decoding paths, and each decoding path corresponds to one of the multiple target paths of the adjacent previous speech frame.
[0097] Taking the homophonic path determined in the above steps as "Jingqiang" as an example, the last word unit of the homophonic path is "qiang", and its homophones can include but are not limited to "qiang", "qiang", etc. The target word units can be determined as the above homophones, and the hot words corresponding to the target word units are searched from the preset hot word library, such as "qiangdiao", "qiangzhuang", etc. The hot words are determined as the homophonic hot words. When there is no hot word corresponding to "qiang" in the preset hot word library, "qiang" does not match the hot word as the homophonic hot word; further, if the preset hot word library includes hot words with the last word unit as the target word unit such as "gengqiang", hot words such as "gengqiang" are not used as the homophonic hot words.
[0098] In the embodiments of the present disclosure, by optimizing the process of determining homophone hotwords, when determining the hotword matching path, associated homophones can be determined according to the target decoding path and other decoding paths, and the corresponding hotwords can be determined from the preset hotword library, so as to retain the candidate path of homophone matching as the hotword matching path and participate in the subsequent screening and determination of the target path; when determining the retention path of homophone hotwords according to this embodiment, the influence of the decoding path and homophone hotwords on the hotword recognition performance and speech recognition accuracy is considered, and the hotword matching path associated with the homophone hotwords is obtained, effectively improving the hotword recognition performance, reducing the probability of omission of the homophone hotword path, optimizing the process of determining the target path of the speech frame, and further improving the accuracy of the subsequent speech signal recognition result and the recognition accuracy of hotwords.
[0099] In one embodiment, the method for determining potential hotwords includes:
[0100] According to the last word unit of the target decoding path, determine the hotword that matches the last word unit from the preset hotword library as the potential hotword.
[0101] In the embodiments of the present disclosure, when determining potential hotwords, according to the last word unit of the target decoding path, determine the potential hotword that matches the last word unit from the preset hotword library. In a possible implementation manner, the hotwords in the preset hotword library include multiple word units. When determining potential hotwords, the last word unit of the target decoding path can be first searched in the preset hotword library. If the last word unit exists in the preset hotword library, the potential hotword is determined according to the hotword corresponding to the last word unit in the hotword library; further, the last word unit of the potential hotword is different from the last word unit of the target decoding path. In one embodiment, when determining potential hotwords from the preset hotword library, all the hotwords that match the last word unit of the target decoding path can be used as potential hotwords; or the potential hotwords can be obtained by further processing and screening of all the hotwords, which can be specifically set according to the actual application scenario.
[0102] In an example, the target decoding path of the first speech frame in the speech signal can be set to be empty. In an exemplary embodiment, when the target decoding path is empty or the last word unit does not hit the hotword library, some hotwords can be directly screened out from the preset hotword library according to the preset screening conditions, and the preset screening conditions can be determined according to the actual application scenario and recognition requirements; or potential hotwords can be screened from the preset hotword library according to multiple candidate paths; or all the hotwords in the preset hotword library can be directly used as potential hotwords; or potential hotwords can be obtained through other possible screening methods. Further, when the target decoding path is empty or the last word unit does not hit the preset hotword library, the potential hotwords can also be directly set to be empty.
[0103] Taking the target decoding path as "Beijing" as an example, the last word unit is "京", and the hot words that hit "京" are searched from the preset hot word library, such as "京剧", "京都", "京城", "北京人", etc., and potential hot words are determined based on the above hot words. In one example, all the above hot words can be directly determined as potential hot words; in another example, the hot words whose last word unit is "京" are screened out from the above hot words, such as "南京", etc., and the hot words obtained after screening are determined as potential hot words; in another example, the above hot words can also be screened in combination with other influencing factors such as hot word scores to obtain potential hot words.
[0104] The disclosed embodiment determines potential hot words according to the last word unit of the target decoding path. The implementation is simple and can quickly determine potential hot words from a preset hot word library. The embodiment can accurately locate the target hot words required in the current candidate path screening and determination process, so that when determining the hot word matching path, the potential matching path can be determined in a targeted manner according to the potential hot words, thereby improving the efficiency and accuracy of hot word matching path determination, reducing the data processing workload in the matching process of the hot word matching path, improving speech recognition efficiency, and being suitable for more application scenarios.
[0105] In one embodiment, the potential hot words can be determined in combination with the hot word scores, such as Figure 6 As shown, the hot words related to the last word unit are determined from the preset hot word library as potential hot words, including:
[0106] Step S610, determining the hot word matching the last word unit from the preset hot word library as a candidate hot word;
[0107] Step S620: Determine, from the candidate hot words, the hot words whose scores are greater than a preset score as potential hot words.
[0108] In the embodiments of the present disclosure, when determining potential hot words from a preset hot word library according to the last word unit of the target decoding path, screening can be performed according to the hot word scores. Optionally, according to the last word unit of the target decoding path, candidate hot words matching the last word unit are determined from the preset hot word library, where the candidate hot words may include hot words containing the last word unit; further, in one example, the last word unit of the candidate hot word is different from the last word unit of the target decoding path. After determining the candidate hot words, potential hot words are determined from the candidate hot words according to the hot word scores corresponding to the candidate hot words. Different candidate hot words may correspond to different hot word scores, and the hot word scores can be used to reflect the relevance of the hot words to the current scenario. In this embodiment, the candidate hot words are further screened by the hot word scores to determine, as potential hot words, the hot words with a higher relevance to the current scenario from the candidate hot words. The number of hot words of the potential hot words can be determined according to the data processing capabilities and recognition requirements in the actual application scenario.
[0109] In some possible implementation manners, the higher the hot word score is, the higher the relevance of the hot word to the current scenario is. When screening potential hot words, hot words with hot word scores greater than a preset score are screened as potential hot words, where the preset score can be determined according to the actual application scenario. For example, it can be set as a fixed value. When the hot word score is less than or equal to the preset score, it indicates that the hot word score is low and the candidate hot word does not need to be retained; the preset number of hot words can be taken as potential hot words from the candidate hot words in descending order of the hot word scores, and the preset score can be determined according to the candidate hot word with the lowest hot word score among the candidate hot words corresponding to the preset number; the preset score can also be determined from the candidate hot words through other possible screening methods.
[0110] In the embodiments of the present disclosure, candidate hot words are determined from the hot word library according to the last word unit of the target decoding path, and potential hot words are determined from the candidate hot words according to the hot word scores, which streamlines the determined potential hot words. While matching the last word unit of the target decoding path, potential hot words with qualified hot word scores are screened; while ensuring the accuracy and effectiveness of the potential hot words, the data processing workload in the matching process of the hot word matching path can be further reduced, taking into account both the efficiency and accuracy of decoding and recognition, and it can be applicable to scenarios with different data processing performances and different recognition requirements.
[0111] Figure 7 FIG. is a schematic flowchart of a speech recognition method shown according to an exemplary embodiment. This method is used to recognize speech frames of a speech signal. Refer to Figure 7As shown, the target decoding path is the beam of "Beijing". "Beijing" is any target path of the previous speech frame adjacent to this speech frame, and the corresponding path score is 10. In this embodiment, the higher the path score, the greater the probability of the current path. Decode the current speech frame based on "Beijing". After decoding through the speech recognition model, multiple candidate paths corresponding to multiple tokens and their corresponding path scores are output, such as "Beijing City: 14", "People from Beijing: 13", etc. Take N candidate paths with the highest path scores from high to low as the score-matching paths. Taking N = 3 as an example, the score-matching paths are "Beijing City", "People from Beijing", "Beijing Opera". Determine the hot-word matching paths from the candidate paths. Among them, the hot-word matching paths can include potential matching paths and homophone matching paths. The target hot words include potential hot words and homophone hot words. The potential hot words are the hot words corresponding to the last word unit "Jing" of the target path "Beijing", such as "Beijing Opera", "Jingdu", "Jingcheng", then determine the potential matching paths as "Beijing Opera", "Beijingdu", "Beijingcheng"; the homophone hot words are determined based on the target decoding path and other decoding paths (i.e., other beams). The other decoding paths are other target paths in the target paths of the previous speech frame adjacent to this speech frame except this target decoding path. For example, the other decoding paths also include "Beijing accent", then determine the homophone hot words corresponding to the homophones based on "accent", such as "strong" and "robust", and determine the homophone matching paths that match the homophone hot words from multiple candidate paths. The corresponding homophone matching paths include "Beijing Qiang"; determine the hot-word matching paths based on the potential matching paths and homophone matching paths, including "Beijing Opera", "Beijingdu", "Beijingcheng", "Beijing Qiang". Determine the retained paths corresponding to "Beijing" based on the score-matching paths and hot-word matching paths, including "Beijing City", "People from Beijing", "Beijing Opera", "Beijingdu", "Beijingcheng", "Beijing Qiang". Update the path scores of the retained paths according to the hot-word scores. Among them, update the path scores of the hot-word matching paths "Beijing Opera", "Beijingcheng", "Beijingdu" according to the scores of the target hot words, that is, "Beijing Opera: 15", "Beijingcheng: 17", "Beijingdu: 12". In an example, update the scores of other retained paths according to other hot words in the preset hot-word library. For example, if "Beijing City" does not hit the hot word corresponding to "Jing", a score reduction process is performed; the "Qiang" in "Beijing Qiang" may hit "strong", so according to the word unit score corresponding to "Qiang", a score increase process is performed. It can be understood that for the same path, both score increase and score reduction processes can be performed simultaneously. After the above processing, the updated path scores are obtained. Combine the retained paths corresponding to other decoding paths and the updated path scores to determine the target path corresponding to the current speech frame. For example, M retained paths with the highest updated path scores can be taken from high to low as the target path, where M is a positive integer and M can be equal to or different from N.If the target path corresponding to the current speech frame includes "Beijing City", then "Beijing City" is the target path corresponding to the target decoding path of "Beijing" for the speech frame.
[0112] Through this embodiment, it is possible to retain the hot word matching path before hot word enhancement, reduce the probability of missing the path where the hot word is hit, improve the accuracy of hot word recognition, and further improve the accuracy of the speech recognition result, making it applicable to more application scenarios; it balances the conflict between the hot word effect and the decoding efficiency, and improves the recognition effect of the hot word without additionally increasing the decoding time consumption, effectively enhancing the scene customization ability in speech recognition.
[0113] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0114] Based on the same inventive concept, the embodiments of the present application also provide a speech recognition device for implementing the above-mentioned speech recognition method. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the speech recognition device provided below can refer to the limitations on the speech recognition method in the above text, and will not be elaborated here.
[0115] In one embodiment, as Figure 8 shown, a speech recognition device 800 is provided. The device is used to recognize speech frames in a speech signal, and includes:
[0116] A decoding module 810, configured to decode the speech frame according to the target decoding path of the speech frame to obtain multiple candidate paths of the speech frame and corresponding path scores. Each candidate path corresponds to a path score, and the target decoding path is any target path of the previous speech frame adjacent to the speech frame.
[0117] A first determination module 820, configured to determine a reserved path from the multiple candidate paths according to the path score and the target hot word, where the reserved path includes score matching paths with the top N path scores and hot word matching paths matching the target hot word, N is a positive integer, and the target hot word is determined from a preset hot word library according to the target decoding path;
[0118] An update module 830, configured to update the path score of the reserved path according to the hot word score of the target hot word in the preset hot word library to obtain an updated path score;
[0119] A second determination module 840, configured to determine a target path corresponding to the target decoding path of the speech frame from the reserved path according to the updated path score.
[0120] In one embodiment, the device includes a determination module for the hot word matching path, and the determination module for the hot word matching path includes:
[0121] A first determination sub-module, configured to determine an initial matching path matching the target hot word from the multiple candidate paths according to the target hot word;
[0122] A first update sub-module, configured to update the path score of the initial matching path according to the hot word score of the target hot word to obtain an updated path score of the initial matching path;
[0123] A second determination sub-module, configured to determine the initial matching paths with the top M path scores of the updated initial matching path as the hot word matching paths, where M is a positive integer.
[0124] In one embodiment, the target hot word includes at least one of a potential hot word and a homophone hot word, and the device includes a determination module for the hot word matching path, and the determination module for the hot word matching path includes:
[0125] A third determination sub-module, configured to determine, from the multiple candidate paths, a candidate path including the potential hot word as a potential matching path;
[0126] A fourth determination sub-module, configured to determine, from the multiple candidate paths, a candidate path matching the homophone hot word as a homophone matching path;
[0127] A fifth determination sub-module, configured to obtain a hot word matching path according to at least one of the potential matching path and the homophone matching path.
[0128] In one embodiment, the device further includes a determination module for the homophone hot word, and the determination module for the homophone hot word includes:
[0129] The sixth determination sub-module is configured to, for the target decoding path, determine, from other decoding paths of the speech frame except the target decoding path, a path whose penultimate word unit is the same as the last word unit of the target decoding path as the homophone path;
[0130] The seventh determination sub-module is configured to determine, from the preset hot word library, a target word unit that is homophonous with the last word unit of the homophone path, and use the hot word corresponding to the target word unit as the homophone hot word.
[0131] In one embodiment, the apparatus further includes a potential hot word determination module, and the potential hot word determination module includes:
[0132] The eighth determination sub-module is configured to determine, according to the last word unit of the target decoding path, a hot word that matches the last word unit from the preset hot word library as the potential hot word.
[0133] In one embodiment, the eighth determination sub-module includes:
[0134] The first determination unit is configured to determine, from the preset hot word library, a hot word that matches the last word unit as the candidate hot word;
[0135] The second determination unit is configured to determine, from the candidate hot words, a hot word whose hot word score is greater than a preset score as the potential hot word.
[0136] Each module in the above speech recognition apparatus can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the electronic device in hardware form or be independent of the processor, or be stored in the memory in the electronic device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0137] In one embodiment, an electronic device is provided. The electronic device can be a server, and its internal structure diagram can be as Figure 9As shown in the figure. The electronic device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the electronic device is used to store data involved in the method described in this embodiment, such as voice signals. The input / output interface of the electronic device is used to exchange information between the processor and external devices. The communication interface of the electronic device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a voice recognition method.
[0138] Those skilled in the art can understand that Figure 9 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0139] This embodiment of the application also provides a computer-readable storage medium. One or more non-volatile computer-readable storage media containing computer-executable instructions, when the computer-executable instructions are executed by one or more processors, cause the processors to execute the steps of the voice recognition method.
[0140] This embodiment of the application also provides a computer program product containing instructions, which when run on a computer, causes the computer to execute the voice recognition method.
[0141] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards.
[0142] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0143] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0144] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A voice recognition method, characterized in that, The method is used to recognize speech frames in a speech signal, and includes: Decoding the speech frame according to the target decoding path of the speech frame to obtain multiple candidate paths of the speech frame and corresponding path scores, each candidate path corresponding to a path score, and the target decoding path being any target path of the previous speech frame adjacent to the speech frame; Determining a reserved path from the multiple candidate paths according to the path scores and the target hot word, the reserved path including score matching paths ranked top N in terms of path scores and hot word matching paths matching the target hot word, where N is a positive integer, and the target hot word is determined from a preset hot word library according to the target decoding path; Updating the path scores of the reserved path according to the hot word scores of the target hot word in the preset hot word library to obtain updated path scores; Determining the target path corresponding to the target decoding path of the speech frame from the reserved path according to the updated path scores.
2. The method according to claim 1, wherein The determining method of the hot word matching path includes: Determining an initial matching path matching the target hot word from the multiple candidate paths according to the target hot word; Updating the path score of the initial matching path according to the hot word score of the target hot word to obtain the path score of the updated initial matching path; Determining the initial matching paths ranked top M in terms of the path scores of the updated initial matching path as the hot word matching paths, where M is a positive integer.
3. The method according to claim 1, characterized in that, The target hot word includes at least one of a potential hot word and a homophone hot word, and the determining method of the hot word matching path includes: Determining a candidate path including the potential hot word from the multiple candidate paths as a potential matching path; Determining a candidate path matching the homophone hot word from the multiple candidate paths as a homophone matching path; Obtaining a hot word matching path according to at least one of the potential matching path and the homophone matching path.
4. The method according to claim 3, characterized in that The determining method of the homophone hot word includes: For the target decoding path, determining, from other decoding paths of the speech frame except the target decoding path, a path whose penultimate word unit is the same as the last word unit of the target decoding path as a homophone path; Determining a target word unit homophonic to the last word unit of the homophone path from the preset hot word library, and taking the hot word corresponding to the target word unit as the homophone hot word.
5. The method according to claim 3, wherein The determining method of the potential hot word includes: Determining a hot word matching the last word unit from the preset hot word library according to the last word unit of the target decoding path as the potential hot word.
6. The method according to claim 5, characterized in that, The determining, from the preset hot word library, a hot word matching the last word unit as the potential hot word includes: Determining a hot word matching the last word unit from the preset hot word library as a candidate hot word; Determining, from the candidate hot words, a hot word with a hot word score greater than a preset score as the potential hot word.
7. A voice recognition device, characterized in that, The device is used to recognize speech frames in a speech signal, and includes: A decoding module, configured to decode the speech frame according to the target decoding path of the speech frame, to obtain multiple candidate paths of the speech frame and corresponding path scores, each candidate path corresponding to a path score, and the target decoding path being any target path of the previous speech frame adjacent to the speech frame; A first determination module, configured to determine a retained path from the multiple candidate paths according to the path scores and a target hot word, the retained path including score matching paths ranked top N in terms of path scores and hot word matching paths matching the target hot word, where N is a positive integer, and the target hot word is determined from a preset hot word library according to the target decoding path; An updating module, configured to update the path scores of the retained path according to the hot word scores of the target hot word in the preset hot word library, to obtain updated path scores; A second determination module, configured to determine, according to the updated path scores, the target path corresponding to the target decoding path of the speech frame from the retained path.
8. The device according to claim 7, characterized in that, The apparatus further includes a determination module for the hot word matching path, and the determination module for the hot word matching path includes: A first determination sub-module, configured to determine, according to the target hot word, an initial matching path matching the target hot word from the multiple candidate paths; A first updating sub-module, configured to update the path score of the initial matching path according to the hot word score of the target hot word, to obtain the path score of the updated initial matching path; A second determination sub-module, configured to determine the initial matching paths ranked top M in terms of the path scores of the updated initial matching path as the hot word matching paths, where M is a positive integer.
9. An electronic device, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that, When the computer program is executed by the processor, the processor is caused to execute the steps of the speech recognition method according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the speech recognition method according to any one of claims 1 to 6 are implemented.