Provide prompts in real time in speech recognition results
By identifying the vocabulary text in the current voice input and generating a predicted subsequent vocabulary text sequence, the problem in the prior art is solved that it is difficult to provide prediction prompts for subsequent vocabulary content in the real-time speech recognition process, and a more accurate and smooth narration process is achieved.
Patent Information
- Application Number
- CN202010517639.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-04
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-06-04
AI Technical Summary
Existing voice recognition technologies are difficult to provide prediction prompts for subsequent vocabulary content in real-time voice recognition, resulting in possible omissions or errors during the narration.
Prompts are provided in real time by identifying the vocabulary text in the current voice input and generating a predicted sequence of subsequent vocabulary texts based on the text. The prompt generation process can take into account factors such as the current discourse text, previous discourse text, event identification of the target event, and speaker's speaker ID.
It realizes the real-time provision of rich information in speech recognition results, improves the accuracy and fluency of the narration process, and improves the audience's comprehension experience.
Smart Images

Figure CN113763943B_ABST
Abstract
Description
Background Art
[0001] Speech recognition technology aims to convert speech signals into text information. For example, by performing speech recognition, acoustic features can be extracted from the speech waveform, and the acoustic features can be further mapped to text corresponding to the speech waveform. Speech recognition technology has been widely used in a variety of scenarios. Summary of the invention
[0002] This summary is provided to introduce a set of concepts that will be further described in the following detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0003] Embodiments of the present disclosure propose a method and apparatus for providing prompts in real time in speech recognition results. A current speech input in an audio stream for a target event can be obtained. A current utterance text corresponding to the current speech input can be identified. A prompt can be generated based at least on the current utterance text, the prompt including at least one predicted subsequent utterance text sequence. A speech recognition result for the current speech input can be provided, the speech recognition result including the current utterance text and the prompt.
[0004] It should be noted that one or more of the above aspects include the features specifically pointed out in the following detailed description and claims. The following description and drawings set forth in detail certain illustrative features of the one or more aspects. These features are merely indicative of the various ways in which the principles of various aspects may be implemented, and the present disclosure is intended to include all such aspects and their equivalents. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] The disclosed aspects will be described below in conjunction with the accompanying drawings, which are provided to illustrate rather than limit the disclosed aspects.
[0006] Figure 1 An exemplary architecture of a speech recognition service deployment according to an embodiment is shown.
[0007] Figure 2 An exemplary architecture of a speech recognition service deployment according to an embodiment is shown.
[0008] Figure 3 An exemplary scenario of providing prompts in real time in speech recognition results according to an embodiment is shown.
[0009] Figure 4 An exemplary scenario of providing prompts in real time in speech recognition results according to an embodiment is shown.
[0010] Figure 5An exemplary process for establishing a prompt generator according to an embodiment is shown.
[0011] Figure 6 An exemplary process of providing prompts in real time in speech recognition results according to an embodiment is shown.
[0012] Figure 7 An exemplary process of providing prompts in real time in speech recognition results according to an embodiment is shown.
[0013] Figures 8 to 11 An example of providing a prompt in a speech recognition result in real time according to an embodiment is shown.
[0014] Fig.12 The flowchart of an exemplary method for providing prompts in speech recognition results in real time according to an embodiment is shown.
[0015] Fig.13 An exemplary apparatus for providing prompts in real time in speech recognition results according to an embodiment is shown.
[0016] Fig.14 An exemplary apparatus for providing prompts in real time in speech recognition results according to an embodiment is shown. DETAILED DESCRIPTION
[0017] The present disclosure will now be discussed with reference to various exemplary embodiments. It should be understood that the discussion of these embodiments is only for enabling those skilled in the art to better understand and thereby implement the embodiments of the present disclosure, and does not teach any limitation on the scope of the present disclosure.
[0018] Existing speech recognition technology can provide text corresponding to the input speech waveform in the speech recognition result. In the case of performing real-time speech recognition, the speech in the audio stream can be converted into text in real time. Some software applications can support the call of real-time speech recognition function, for example, remote conference system, social network application, slide application, etc. When a user uses these software applications and provides voice input or narrates speech in the form of audio stream, the real-time speech recognition function can convert the speech in the audio stream into a speech text sequence as a speech recognition result in real time, and the speech text sequence can be presented to the user or other users to indicate the content narrated by the speaker in text form. The speech recognition result provided by the real-time speech recognition function can include the speech text recognized from the current speech input and the possible speech text recognized from the previous speech input.
[0019] The embodiment of the present disclosure proposes to provide a prompt in real time in the speech recognition result, and the prompt includes a prediction of the content of the possible subsequent speech. Thus, the speech recognition result can not only present the current speech text recognized from the current speech input, but also present the predicted subsequent speech text. Taking a user using a slide application to give a speech as an example, not only can the speech text corresponding to the speech speech currently spoken by the user be presented in the slide application through speech recognition, but also the text of the speech that the user may speak next can be presented in real time in the form of a prompt. For the user as a speaker, the prompt can help remind the speaker what content should be spoken next, avoid missing content, avoid incorrectly speaking content, etc., so that the entire speaking process is more accurate and smooth. For the audience, the prompt can help them to know in advance what content will be heard next, so that they can better understand the speaker's intention. In addition, the prompt can also include using the predicted subsequent speech text to correct the omissions or errors in the speech actually spoken by the speaker, so as to further facilitate the understanding of the audience. It should be understood that in this article, a speech can include one or more words.
[0020] In one aspect, corresponding prompts can be provided in real time with the recognition of the current voice input in the audio stream. For example, whenever the speaker tells one or more words, prompts in response to the current words are provided. In this article, words can refer to language units in different languages, for example, characters in Chinese, words in English, etc. In addition, in this article, depending on the specific implementation of speech recognition, a word can include one or more characters, one or more words, one or more phrases, etc.
[0021] In one aspect, the prompt may include one or more predicted subsequent utterance text sequences, each predicted subsequent utterance text sequence may include one or more predicted subsequent utterance texts, and each predicted subsequent utterance text may include one or more words. Different numbers of predicted subsequent utterance texts included in each sequence may provide different prediction spans.
[0022] In one aspect, the prompt generation process can be based on the current utterance text recognized for the current voice input. For example, a subsequent utterance text sequence can be predicted based at least on the current utterance text. In addition, the prompt generation process can further consider one or more recognized previous utterance texts. For example, a subsequent utterance text sequence can be predicted based at least on the current utterance text and at least one previous utterance text.
[0023] In one aspect, the prompt generation process can consider various factors that contribute to improving the accuracy of the prompt. For example, the prompt generation process can consider the event identification (ID) of the target event. The target event can refer to the event targeted by the audio stream or voice input. Taking a discussion about improving the efficiency of productivity tools by multiple participants through a teleconference system as an example, the target event can be <discussion about improving the efficiency of productivity tools>, and the voice utterances in the audio streams from these participants can be considered to be targeted at the target event. By using the event ID of the target event in the prompt generation process, the prompt can be generated in a manner more suitable for the target event. In this article, event ID is an identification for distinguishing between different events, which can be used in various ways such as strings, numerical values, vectors, etc. In addition, for example, the prompt generation process can consider the speaker ID of the speaker of the current voice input. By using the speaker ID of the speaker in the prompt generation process, the prompt can be generated in a manner more suitable for the speaker. In this article, speaker ID is an identification for distinguishing between different speakers, which can be used in various ways such as strings, numerical values, vectors, etc. It should be understood that in addition to considering event ID and speaker ID, the prompt generation process may also consider any other factors that help improve the accuracy of the prompt.
[0024] In one aspect, a prompt generator can be pre-established for performing prompt generation. In one case, a document associated with a target event can be used to establish a prompt generator to achieve customization of the prompt generator for the target event. Optionally, event information associated with the target event can be further used to establish the prompt generator. In one case, multiple documents associated with multiple events can be used to establish a prompt generator so that the prompt generator has higher versatility. Optionally, event information associated with each of the multiple events can be further used to establish a prompt generator. In one case, multiple documents associated with multiple speakers can be used to establish a prompt generator to achieve customization of the prompt generator for a specific speaker. Optionally, a speaker ID associated with each speaker can be further used to establish a prompt generator. In addition, all of the above situations can also be combined to establish a prompt generator that is both universal and customizable. The prompt generator can be established based on various text prediction technologies, such as neural network models, correlation matrices, etc. In the application phase, the established prompt generator can predict a subsequent utterance text sequence based on at least the current utterance text and optionally the previous utterance text.
[0025] In one aspect, the current utterance text and the prompt can be provided in different presentation modes in the speech recognition result, so as to visually distinguish the current utterance text and the prompt. In this article, the presentation mode can refer to, for example, font, font size, color, bold, italic, underline, layout position, etc.
[0026] By providing prompts in real time in speech recognition results according to an embodiment of the present disclosure, an improved speech recognition service can be achieved so that the speech recognition results contain richer information, thereby improving the usage efficiency and product effect of related software applications, improving the user experience of both speakers and listeners, etc.
[0027] Figure 1 An exemplary architecture 100 for speech recognition service deployment according to an embodiment is shown. In the architecture 100, the speech recognition results including prompts generated according to the embodiment of the present disclosure are directly provided to the target application operated by the end user, so that the end user can obtain improved speech recognition services.
[0028] Assume that user 102 is an end user of target application 110. The target application 110 may be various software applications that support the call of speech recognition function, such as a teleconferencing system, a social networking application, a slide application, etc. The target application 110 may include a speech recognition application program interface (API) 112, and the speech recognition API 112 may call a speech recognition service 120 to obtain the speech recognition function. For example, during the period when user 102 uses the target application 110, if user 102 wants to enable the speech recognition function, the speech recognition API 112 may be triggered by, for example, a control button in a user interface (UI). In addition, optionally, the speech recognition API 112 may also be enabled by default or preset to be enabled by the provider of the target application 110.
[0029] The speech recognition service 120 may generally refer to various functional entities capable of implementing speech recognition functions, such as a speech recognition platform or server independent of the target application 110, a speech recognition module included in the target application 110, etc. Therefore, although for ease of explanation, Figure 1 1 and 110 are shown in different blocks, but it should be understood that the speech recognition service 120 can be independent of the target application 110 or included in the target application 110. In addition, the speech recognition service 120 is an improved speech recognition service that can provide prompts in real time in the speech recognition results.
[0030] When the user 102 provides a series of voice inputs in the form of an audio stream, for example, when the user 102 speaks an utterance in voice, the speech recognition API 112 can provide the current voice input 104 in the audio stream to the speech recognizer 122 in the speech recognition service 120. The speech recognizer 122 can use any known speech recognition technology to convert the current voice input 104 into the corresponding current utterance text.
[0031] According to an embodiment of the present disclosure, the prompt generator 124 in the speech recognition service 120 may generate a prompt based at least on the current utterance text output by the speech recognizer 122. The prompt may include at least one predicted subsequent utterance text sequence.
[0032] The combination module 126 in the speech recognition service 120 can combine the current utterance text output by the speech recognizer 122 and the prompt output by the prompt generator 124 to form a speech recognition result 106 corresponding to the current speech input 104. Optionally, the combination module 126 can combine the current utterance text and the prompt in different presentation modes.
[0033] The speech recognition result 106 including the prompt may be returned to the speech recognition API 112. Thus, the target application 110 may present the speech recognition result 106 in the UI to the user 102 and other persons who may see the UI. It should be understood that, in the case where the target application 110 is a multi-party application that allows different users to run simultaneously on their respective terminal devices, for example, in the case where the target application 110 is a remote conference system that allows multiple participants to run clients on their respective terminal devices, the speech recognition result 106 may be provided simultaneously to different terminal devices running the target application 110, so that different users may view the speech recognition result.
[0034] Figure 2 An exemplary architecture 200 for the deployment of a speech recognition service according to an embodiment is shown. The architecture 200 is in Figure 1 A modification of the architecture 100, and Figure 1 and Figure 2 The same reference numerals in the structure 200 may indicate the same unit or function. In the structure 200, the speech recognition result including the prompt generated according to the embodiment of the present disclosure is first provided to the application provider platform, and after being processed by the application provider platform, the processed speech recognition result is provided to the target application operated by the end user.
[0035] like Figure 2As shown, the speech recognition service 120 can provide the speech recognition result 106 containing the prompt to the application provider platform 210. The application provider platform 210 can refer to a network entity that can provide, for example, operation control, business management, data processing, development, etc. for the target application 110. Assuming that the target application 110 is a teleconference system, the application provider platform 210 can be, for example, a server of the teleconference system. The application provider platform 210 can process the speech recognition result 106 based on a predefined prompt presentation strategy 212 to generate a processed speech recognition result 220. The prompt presentation strategy 212 can control which content in the prompt in the speech recognition result 106 is presented to the user 102, and in what manner the prompt is provided in the processed speech recognition result 220. For example, the prompt in the speech recognition result 106 may include multiple predicted subsequent utterance text sequences, and the prompt presentation strategy 212 can determine that only the highest-ranked predicted subsequent utterance text sequence is included in the processed speech recognition result 220. For example, the prompt presentation strategy 212 may define the desired presentation of the current utterance text and prompts in the processed speech recognition results 220 .
[0036] The processed speech recognition result 220 may be returned to the speech recognition API 112. Thus, the target application 110 may present the processed speech recognition result 220 in the UI.
[0037] According to the architecture 200, by providing the speech recognition result output by the speech recognition service to the application provider platform and forming the speech recognition result to be presented in the application provider platform, the provider of the target application can autonomously set a strategy on how to provide prompts in the speech recognition result. Thus, the application provider platform can obtain customization through the predefined prompt presentation strategy. Thus, in fact, the application provider platform is a "platform user" that directly obtains the improved speech recognition service, and the user 102 is an "end user" that indirectly obtains the improved speech recognition service.
[0038] It should be understood that, although the prompt presentation strategy 212 is shown in the architecture 200 as being included in the application provider platform 210, the prompt presentation strategy 212 may optionally be deployed in the target application 110. Thus, the target application 110 may receive the speech recognition result 106, and after processing the speech recognition result 106 according to the prompt presentation strategy 212, present the processed speech recognition result in the UI. In addition, it should be understood that, although the processed speech recognition result is finally presented in the UI of the target application 110 in the architecture 200, the speech recognition function may also be called by the target application 110 in the background, and thus, the target application platform 210 may not return the processed speech recognition result to the target application 110. In this case, the target application platform 210 may serve as the target receiving endpoint of the speech recognition result 106, and the speech recognition result 106 may be further utilized.
[0039] It should be understood that Figure 1 and Figure 2 Only an exemplary architecture for the deployment of the speech recognition service is shown. Depending on the actual application scenario, the embodiments of the present disclosure may also cover any other deployment architecture.
[0040] Figure 3 An exemplary scenario 300 of providing prompts in real time in speech recognition results according to an embodiment is shown. Assume that an end user 302 is using a slideshow application, and the slideshow application has an exemplary UI 310.
[0041] The UI 310 may include a display area 312 in which the slide content is displayed. The UI 310 may include a speech recognition button 314 for invoking the speech recognition function. The UI 310 may also include a speech recognition result area 316 for displaying the speech recognition result of the speech utterance in real time.
[0042] Assume that user 302 is speaking in voice while presenting a slide, thereby generating a corresponding audio stream 304. As an example, user 302 speaks in English "The next topic is how we leverage Hidden", where "Hidden" is the current voice input. Through the improved speech recognition service according to the embodiment of the present disclosure, the speech recognition result area 316 can display multiple previous speech texts "The next topic is how we leverage", the recognized current speech text "Hidden", and the prompt "Markov Model to evolve". The prompt includes a predicted subsequent speech text sequence, which includes 4 predicted subsequent speech texts "Markov", "Model", "to", and "evolve". In other words, although user 302 has just spoken the word "Hidden", the speech recognition result area 316 has already displayed a prediction of the speech that user 302 may speak next.
[0043] like Figure 3 As shown, the prompt is provided in a different presentation mode from the previous utterance text and the current utterance text. For example, the text "Markov Model to evolve" in the prompt is presented in a different font, bold, and italic.
[0044] Figure 4 An exemplary scenario 400 of providing prompts in real time in speech recognition results according to an embodiment is shown. Assume that the end user Clair is using a remote conference system, and the remote conference system has an exemplary UI 410 .
[0045] The UI 410 may include an information area 412 in which the participants of the meeting are listed, for example, David, Brown, Clair, etc. The UI 410 may include a control area 414 including a plurality of control buttons for implementing different control functions. The control area 414 may include a speech recognition button 416 for invoking the speech recognition function. The UI 410 may also include a speech recognition result area 418 for displaying the speech recognition results of the speech utterances of different participants in the meeting in real time.
[0046] Assume that user Clair is speaking a speech in voice, thereby generating a corresponding audio stream 402. As an example, user Clair speaks "It's a challenge to improve our productivity" in English, where "productivity" is the current voice input. Through the improved speech recognition service according to the disclosed embodiment, the speech recognition result area 418 can display a plurality of previous speech texts "It's a challenge to improve our", the recognized current speech text "productivity", and the prompt "tools inprocessing" that user Clair has spoken. The prompt includes a predicted subsequent speech text sequence, which includes 3 predicted subsequent speech texts "tools", "in" and "processing". In other words, although user Clair has just spoken the word "productivity", the speech recognition result area 418 has already displayed the prediction of the speech that user Clair may speak next.
[0047] like Figure 4 As shown, the prompt is provided in a different presentation mode from the previous utterance text and the current utterance text. For example, the text "tools in processing" in the prompt is presented in a different font and underlined mode.
[0048] It should be understood that Figure 3 and Figure 4 All elements and their layouts in are exemplary, and are only used to intuitively show exemplary scenarios according to embodiments of the present disclosure. In addition, embodiments of the present disclosure are not limited to the exemplary presentation methods shown for providing prompts, but various other presentation methods may be used. In addition, although voice utterances and utterance texts in English are shown, embodiments of the present disclosure are not limited to any specific language.
[0049] Figure 5 An exemplary process 500 for establishing a prompt generator is shown according to an embodiment.
[0050] According to an embodiment of the present disclosure, in order to accurately predict subsequent utterance text and generate prompts, a corpus may be constructed by at least utilizing documents associated with an event and / or a speaker, and then the corpus may be used to establish a prompt generator.
[0051] In one implementation, at least one document associated with the target event T may be obtained, for example, document T_1, document T_2, etc. The target event may refer to an event for which speech recognition will be implemented in a target application, for example, a meeting involving a certain discussion topic held through a remote conference system, a speech on a certain topic made using a slide application, and the like. Taking a speech S involving topic L made through a slide application as an example, the speech S may have one or more associated documents, such as a slide file that the speaker will present during the speech, and these documents include discussion content related to topic L. These documents may provide information that helps predict subsequent utterance texts. For example, assuming that multiple sentences in these documents include the term "Hidden Markov Model", which indicates that the three words "Hidden", "Markov" and "Model" have a high correlation at least in the speech S, then this information will help predict the subsequent utterance text sequence "Markov Model" with a higher probability when the current utterance text is, for example, "Hidden". Optionally, event information associated with the target event T may also be obtained. Event information may include, for example, the event ID of the target event T, the time information of the target event T, the agenda of the target event T, the participant ID of the participant of the target event T, the speaker ID of the speaker associated with the document involved in the target event T, and the like. The participant of the target event T may refer to the personnel who participated in the target event, including the speaker who spoke in the target event and the personnel who did not speak. The time information of the target event T may include, for example, the time point when the target event T starts, the time period of the target event T, and the like. The agenda of the target event T may include, for example, topics, speakers, and the like arranged in chronological order. The establishment of the prompt generator may be further based on the obtained event information. By establishing the prompt generator using the document associated with the target event T and the optional event information before the target event T occurs, the prompt generator may be made to generate prompts adapted to the target event more accurately during the occurrence of the target event T, thereby realizing the customization of the prompt generator for the target event T. The document associated with the target event T and the event information may be provided by the participant, speaker, or other person of the target event T. The obtained documents associated with the target event T and the event information may be stored in the document corpus 510 .
[0052] In one implementation, multiple documents associated with multiple different events can be collected to be used together to establish a prompt generator. For example, documents A_1 and A_2 associated with event A, document B_1 associated with event B, and so on can be obtained. In addition, optionally, event information associated with these events can also be obtained. By using documents associated with multiple events and optional event information to establish a prompt generator, the versatility of the prompt generator can be enhanced. The obtained documents associated with multiple events and event information can be stored in the document corpus 510.
[0053] In one implementation, multiple documents respectively associated with multiple different speakers may be collected for use in establishing a prompt generator. For example, documents M_1 and M_2 associated with speaker M, documents N_1, N_2, and N_3 associated with speaker N, and so on may be obtained. In addition, optionally, speaker IDs respectively associated with these speakers may also be obtained. Although multiple documents for each speaker may be associated with different events, these documents contribute to jointly constructing a corpus specific to the speaker, thereby enabling customization of the prompt generator for the speaker. The documents and speaker IDs associated with a specific speaker may be provided by the speaker or other persons. The obtained documents and speaker IDs associated with multiple speakers may be stored in a document corpus 510.
[0054] Through the above process, the document corpus 510 will store a plurality of documents, each of which may have corresponding tag information, such as event information, speaker ID, etc. At 520, text analysis may be performed on the documents in the document corpus 510 to obtain data for establishing a prompt generator. For example, a word sequence or word set may be extracted from each document by word segmentation, and each word will have the same tag information as the document. In addition, depending on the specific implementation of the prompt generator, the text analysis at 520 may also include possible further processing.
[0055] According to an embodiment of the present disclosure, the prompt generator 530 may be based on a neural network model, for example, a recurrent neural network (RNN) model for text prediction. The neural network model may be trained to predict the next word based on one or more input words using a pre-prepared word sequence. For example, the neural network model may be an n-gram model, where n≥1, so that the next word may be predicted based on n words. In one implementation, the neural network model may be trained using a plurality of word sequences extracted from a plurality of documents associated with a plurality of events through text analysis at 520 to obtain a general prompt generator. In one implementation, the neural network model may be trained using a word sequence extracted from a document associated with a target event through text analysis at 520 to obtain a prompt generator customized for the target event. In one implementation, the neural network model may be trained using a word sequence extracted from a document associated with a specific speaker through text analysis at 520 to obtain a prompt generator customized for the specific speaker. It should be understood that the above multiple implementations may also be combined in any manner. For example, after obtaining a general prompt generator, the general prompt generator can be retrained and optimized using word sequences extracted from documents associated with a target event to obtain a prompt generator customized for the target event. For example, after obtaining a general prompt generator, the general prompt generator can be retrained and optimized using word sequences extracted from documents associated with a specific speaker to obtain a prompt generator customized for the specific speaker. For example, a neural network model can be trained using both word sequences extracted from documents associated with a target event and word sequences extracted from documents associated with a specific speaker to obtain a prompt generator customized for both the target event and the specific speaker.
[0056] In addition, optionally, when training the prompt generator 530 based on the neural network model, the tag information of the words can be further considered. For example, when the neural network model predicts the next word based on one or more input words, a higher weight can be given to candidate words with the same event ID, speaker ID, etc. as the input word, so that these candidate words have a higher probability of being predicted as the next word. This can further improve the prediction accuracy because words with the same event ID and / or speaker ID usually have a higher correlation.
[0057] In addition, optionally, in each prediction step, the neural network model-based prompt generator 530 may also output a plurality of next word candidates with the highest ranking.
[0058] In addition, the hint generator 530 based on the neural network model can be iteratively run to predict a subsequent word sequence in response to an input word. For example, after the hint generator 530 based on the neural network model predicts a first subsequent word at least based on the input word, it can further predict a second subsequent word at least based on the first subsequent word, and so on, and finally obtain a subsequent word sequence including multiple subsequent words. Further, assuming that the hint generator 530 based on the neural network model predicts multiple first subsequent word candidates with the highest ranking based on at least the input word, it can further iteratively predict subsequent words for each first subsequent word candidate, thereby finally obtaining multiple predicted subsequent word sequences. In the case of outputting multiple predicted subsequent word sequences in response to the same input word, the total confidence of each sequence can be calculated respectively, and the sequences can be sorted based on the total confidence. Optionally, at least one subsequent word sequence with the highest confidence ranking can be output.
[0059] According to an embodiment of the present disclosure, prompt generator 530 can be based on a correlation matrix. The correlation matrix can include multiple text items and correlation values between the multiple text items, wherein each text item can be one or more words. For example, the correlation matrix can be an n-ary matrix, wherein n≥1, so that each text item can include n words. If n>1, each text item includes n continuous words in the document. In this case, the text analysis at 520 can further include determining the correlation values between the text items in the extracted multiple word sequences. The correlation value between two text items can be determined by any known means, for example, at least one of the following factors can be considered: the semantic correlation between the two text items, whether the two text items are adjacent, whether the two text items belong to the same document, whether the two text items have the same event ID and / or speaker ID, etc. An exemplary correlation matrix is shown in Table 1 below.
[0060] Text Item topic leverage Hidden Markov Model to evolve scenario … topic -- 1 2 3 4 5 6 7 leverage -1 -- 1 2 3 4 5 6 Hidden -2 -1 -- 1 2 3 4 5 Markov -3 -2 -1 -- 1 2 3 4 Model -4 -3 -2 -1 -- 1 2 3 to -5 -4 -3 -2 -1 -- 1 2 evolve -6 -5 -4 -3 -2 -1 -- 1 scenario -7 -6 -5 -4 -3 -2 -1 -- …
[0061] Table 1
[0062] The correlation matrix in Table 1 is a 1-element matrix, which shows the correlation values between exemplary text items "topic", "leverage", "Hidden", "Markov", "Model", "to", "scenario", etc. The numerical values in Table 1 indicate the correlation values of the text items in the column relative to the text items in the row. Assume that these text items are extracted from a sentence in the document "Then next topic is how we leverage Hidden Markov Model to evolve scenario...". Through the text analysis at 520, the correlation values between these text items can be obtained, for example, "leverage" has a correlation value of "1" relative to "topic", "Model" has a correlation value of "2" relative to "Hidden", "Markov" has a correlation value of "-1" relative to "Model", and so on. The smaller the absolute value of the correlation value between two text items, the higher the correlation between the two text items. A positive correlation value indicates a correlation in language order, and a negative correlation value indicates a correlation in reverse language order. For example, for the text items "Hidden", "Markov" and "Model" in the above exemplary sentences, the language order is "Hidden→Markov→Model", so that "Markov" has a positive correlation value relative to "Hidden" and "Model" has a positive correlation value relative to "Markov", and the inverse language order is "Hidden←Markov←Model", so that "Hidden" has a negative correlation value relative to "Markov" and "Markov" has a negative correlation value relative to "Model". Although the absolute value of the correlation value "-1" of "Hidden" relative to "Markov" is equal to the correlation value "1" of "Model" relative to "Markov", that is, the degree of correlation is the same, in terms of language order, "Model" has a higher correlation relative to "Markov". It should be understood that all elements in Table 1 are exemplary, and in actual applications, the correlation matrix may include more content or have different formats. For example, the text items in the correlation matrix may come from different sentences or documents, the correlation values may be expressed in different ways, and so on. For example, each text item in the correlation matrix may include more than one word.
[0063] For an input word, the hint generator 530 based on the correlation matrix can select or predict the next text item or word according to the correlation value between the text item in the matrix and the input word. For example, assuming that the input word is "Markov", according to the correlation matrix of Table 1, the text item "Model" with the highest correlation with the text item "Markov" can be selected as the predicted next text item or word, wherein the text item "Model" has a positive correlation value "1" relative to the text item "Markov", and the absolute value of the positive correlation value is the smallest compared to other text items. Optionally, for each input word, the hint generator 530 based on the correlation matrix can output multiple next text items or word candidates with the highest ranking. For example, assuming that the input word is "leverage", according to the correlation matrix in Table 1, the two text items "Hidden" and "Model" with the highest correlation with the text item "leverage" can be selected as the two predicted next text items or word candidates, among which the text item "Hidden" with a positive correlation value of "1" is ranked higher than the text item "Model" with a positive correlation value of "2", and is therefore more likely to be the actual next word.
[0064] Although the tag information of the text items or words, such as event ID, speaker ID, etc., is taken into account when determining the correlation value between the text items at 520, the corresponding tag information can also be added to each text item in the correlation matrix. Thus, when the correlation matrix is used to predict the next word based on the input word, a higher weight can be given to the text items with the same event ID, speaker ID, etc. as the input word, so that these text items have a higher probability of being predicted as the next word. This can further improve the prediction accuracy.
[0065] In addition, the prompt maker 530 based on the correlation matrix can be iteratively operated so as to predict or select a subsequent text item sequence or a subsequent word sequence in response to an input word. For example, after predicting the first subsequent text item based on at least the input word, the prompt maker 530 based on the correlation matrix can then predict the second subsequent text item at least based on the first subsequent text item, and so on, and finally obtain a subsequent text item sequence including a plurality of subsequent text items. Further, assuming that the prompt maker 530 based on the correlation matrix predicts a plurality of first subsequent text item candidates with the highest ranking based on at least the input word, then it can then predict the subsequent text item iteratively for each first subsequent text item candidate respectively, thereby finally obtain a plurality of predicted subsequent text item sequences. In the case of outputting a plurality of predicted subsequent text item sequences in response to the same input word, the total correlation of each sequence can be calculated respectively, and these sequences are sorted based on the total correlation. Optionally, at least one text item sequence with the highest correlation ranking can be selected to output.
[0066] It should be understood that the establishment process of the prompt generator 530 discussed above in conjunction with process 500 is merely exemplary, and process 500 may be modified in any manner according to specific application requirements and designs. For example, the document corpus 510 may include one or more of the following: documents associated with the target event T; multiple documents associated with multiple events; and documents associated with a particular speaker. For example, different prompt generators may be established using different corpora in the document corpus 510, and these prompt generators may be used together to predict subsequent utterance text. As an example, a general prompt generator and a customized prompt generator for the target event T may be established in advance, and when predictions are made during the occurrence of the target event T, the best prediction result among the prediction results generated by the two prompt generators may be selected in various ways as the final predicted subsequent utterance text. In addition, it should be understood that the prompt generator 530 established by process 500 can be applied to predict subsequent utterance text or a subsequent utterance text sequence based on at least the current utterance text, wherein the input word provided to the prompt generator 530 can be the current utterance text, and the next word / text item or subsequent word / text item sequence predicted by the prompt generator 530 can be the subsequent utterance text or a subsequent utterance text sequence.
[0067] Figure 6 An exemplary process 600 for providing hints in speech recognition results in real time according to an embodiment is shown.
[0068] Assume that a target event 602 is occurring in a target application, and a speaker 604 is speaking an utterance in speech and thereby generating an audio stream for the target event 602. For a current speech input 610 in the audio stream, for example, a word currently being spoken by the speaker 604, speech recognition may be performed on the speech input 610 at 620 to generate a current utterance text 630 corresponding to the current speech input 610. This may be accomplished by, for example, Figure 1 and Figure 2 The speech recognition at 620 is performed by the speech recognizer 122 in .
[0069] At 640, a prompt 660 including at least one predicted subsequent utterance text sequence may be generated based at least on the current utterance text 630. Figure 5 The prompt generator 530 pre-established in the process 500 is used to perform the prompt generation at 640. For example, the at least one predicted subsequent utterance text sequence can be predicted based on at least the current utterance text 630 by the prompt generator.
[0070] Optionally, in one implementation, the prompt generation at 640 may be further based on at least one previous utterance text 650 recognized before the current utterance text 630, for example, one or more words recognized for one or more voice inputs before the current voice input 610 in the audio stream. Assuming that the prompt generator is n-gram (n>1), the at least one previous utterance text 650 may include, for example, n-1 previous utterance texts.
[0071] Optionally, in one implementation, the event ID 606 of the target event 602 and / or the speaker ID 608 of the speaker 604 may be obtained, and the event ID 606 and / or the speaker ID 608 may be used for prompt generation at 640. The event ID 606 and / or the speaker ID 608 may be provided by a participant, speaker, or other person associated with the target event 602. Optionally, the event ID 606 and / or the speaker ID 608 may be determined based on pre-acquired event information of the target event 602. For example, based on a comparison of the current time with the time information of the target event 602, it may be determined that the target event 602 has occurred and the corresponding event ID may be extracted. For example, the speaker ID corresponding to the current speaker 604 may be determined based on the agenda of the target event 602.
[0072] Depending on the specific implementation of the prompt generator, the event ID 606 and / or the speaker ID 608 may be used in different ways at 640. For example, in one case, the event ID 606 and / or the speaker ID 608 may be used to select a pre-established customized prompt generator corresponding to the event ID 606 and / or the speaker ID 608 to perform the prompt generation at 640. For example, in one case, the event ID 606 and / or the speaker ID 608 may be used to influence the prediction of the subsequent utterance text by the general prompt generator, so that the candidate words with the same event ID 606 and / or the speaker ID 608 may be determined as the subsequent utterance text with a higher probability. For example, in one case, if multiple prompt generators are used to jointly predict the subsequent utterance text, one or more of the prompt generators may predict the subsequent utterance text based on the event ID 606 and / or the speaker ID 608.
[0073] It should be understood that the predicted subsequent utterance text sequence in the prompt 660 may include one or more subsequent utterance texts. For example, after predicting a first subsequent utterance text adjacent to the current utterance text 630 based on at least the current utterance text 630, the second subsequent utterance text may be predicted based on at least the first subsequent utterance text, and so on. These predicted subsequent utterance texts may be included in the prompt 660 in the form of a sequence.
[0074] At 670, the current utterance text 630 and the prompt 660 may be combined to form a speech recognition result 680. Optionally, in the speech recognition result 680, the current utterance text 630 and the prompt 660 may be presented in different ways to visually distinguish the current utterance text 630 and the prompt 660.
[0075] It should be understood that the above-mentioned process 600 is only exemplary, and according to specific application requirements and designs, the process 600 can be modified and modified in any way. For example, as new voice inputs continue to appear in the audio stream, the process 600 can be iteratively executed for the new voice inputs to provide a voice recognition result including a prompt in real time. In addition, the process 600 can also include determining the relevant information of the current utterance text and / or the prompt, and the relevant information is included in the voice recognition result. The relevant information can refer to various information associated with the predicted subsequent utterance text sequence in the current utterance text and / or the prompt. For example, assuming that the current utterance text is "Hidden", and the predicted subsequent utterance text sequence in the prompt is "Markov Model", then the text description, link, etc. associated with "Hidden Markov Model" can be determined as relevant information. The embodiments of the present disclosure can enhance the information richness of the voice recognition result by including relevant information in the voice recognition result, but are not limited to any specific way of determining or obtaining the relevant information. In addition, it should be understood that the voice recognition result 680 can also include one or more previous utterance texts identified.
[0076] Figure 7 An exemplary process 700 of providing prompts in real time in speech recognition results according to an embodiment is shown. Process 700 can be viewed as Figure 6 An exemplary continuation of process 600 in .
[0077] Assume that a first speech recognition result 704 corresponding to the first speech input 702 has been generated through the process 600, wherein the first speech input 702 and the first speech recognition result 704 may correspond to Figure 6 The current voice input 610 and the voice recognition result 680 in the first voice recognition result 704 may include a first speech text 706 and a first prompt 708, which correspond to Figure 6 The current utterance text 630 and prompt 660 in .
[0078] Assume that after obtaining the first voice input 702, a subsequent second voice input 710 is further obtained in the audio stream, and the second voice input 710 becomes the current voice input. At 720, voice recognition can be performed on the second voice input 710 to generate a second utterance text 730 corresponding to the second voice input 710, and the second utterance text 730 becomes the current utterance text. The voice recognition at 720 can be similar to Figure 6 Speech recognition at 620 in.
[0079] At 740, a second prompt 760 may be generated based at least on the second utterance text 730. In one implementation, process 700 may include comparing the second utterance text 730 with the predicted subsequent utterance text sequence in the first prompt 708 at 750, and using the result of the comparison for prompt generation at 740. Through the comparison at 750, it may be identified whether the speaker's speech utterance includes errors or omissions.
[0080] For example, a speaker wants to speak “The next topic is how we leverage HiddenMarkov Model to evolve…” Assume that the first speech input 702 corresponds to “Hidden”, the first utterance text 706 is recognized as “Hidden”, and the first prompt 708 is “Markov Model to”.
[0081] If the speaker's second speech input 710 corresponds to "Macao", which is misspoken, and the second utterance text 730 is recognized as "Macao", it can be found by comparing the second utterance text 730 with the first prompt 708 that the speaker may have mistakenly spoken "Markov" as "Macao" which sounds similar. In this case, the comparison result can be provided to the prompt generation at 740. Accordingly, at 740, a subsequent utterance text sequence can be predicted based on at least the corrected second utterance text "Markov", or alternatively, a respective subsequent utterance text sequence can be predicted based on the recognized second utterance text "Macao" and the corrected second utterance text "Markov".
[0082] If the speaker's second speech input 710 corresponds to "Model", and the second utterance text 730 is recognized as "Model", it can be found that the speaker may have omitted the word "Markov" between the word "Hidden" and the word "Model" by comparing the second utterance text 730 with the first prompt 708. In this case, the comparison result can be provided to the prompt generation at 740. Accordingly, at 740, the subsequent utterance text sequence can be predicted by taking the omitted word "Markov" as the previous utterance text of the second utterance text 730.
[0083] In addition to the above processing for errors or omissions, the prompt generation process at 740 can be similar to Figure 6 It should be understood that the correction of the wrongly spoken words or the addition of the missing words can be included in the second prompt 760 in any manner according to predetermined rules.
[0084] At 770 , the second utterance text 730 and the second prompt 760 may be combined into a second speech recognition result 780 .
[0085] It should be understood that the above process 700 is merely exemplary, and the process 700 may be modified and altered in any manner according to specific application requirements and designs.
[0086] Figure 8 An example 800 of providing prompts in real time in speech recognition results according to an embodiment is shown. Example 800 is intended to illustrate speech recognition results provided in different ways.
[0087] Assume that the audio stream 810 from the speaker includes speech corresponding to "The next topic is how we leverage Hidden", where "Hidden" is the current speech input 812. At 820, the current utterance text "Hidden" corresponding to the current speech input 812 can be identified. In addition, at 830, a prompt can be generated based at least on the current utterance text. The prompt generation at 830 can be similar to Figure 6 The prompt generation at 640 in FIG. The current utterance text and the prompt may be combined into a speech recognition result. Figure 8 Three exemplary speech recognition results provided in different ways are shown.
[0088] In the exemplary speech recognition result 840a, the prompt only includes a predicted subsequent utterance text sequence. As shown, the speech recognition result 840a can include the current utterance text 840a-2 "Hidden" and the prompt 840a-4. The prompt 840a-4 only includes a predicted subsequent utterance text sequence "Markov Model to evolve", wherein the predicted subsequent utterance text sequence includes a plurality of subsequent utterance texts "Markov", "Model", "to", "evolve" etc. predicted in sequence. The word "Markov" is predicted at least based on the current utterance text "Hidden", for example, the prompt generator can predict "Markov" based on "Hidden", or the prompt generator can predict "Markov" based on "leverage" and "Hidden", etc. Similarly, the word "Model" is predicted at least based on the word "Markov", the word "to" is predicted at least based on the word "Model", and the word "evolve" is predicted at least based on the word "to". In addition, the speech recognition result 840a also includes a plurality of previous utterance texts recognized before the current utterance text 840a-2, for example, "leverage", "we", "how", etc. The current utterance text 840a-2 and the prompt 840a-4 are provided in different presentation modes.
[0089] In the exemplary speech recognition result 840b, the prompt includes multiple predicted subsequent utterance text sequences. As shown, the speech recognition result 840b can include the current utterance text 840b-2 "Hidden" and the prompt 840b-4. The prompt 840b-4 includes three predicted subsequent utterance text sequences, for example, "Markov Model to evolve", "LayerNeural Network to", "state information to get", etc., wherein each predicted subsequent utterance text sequence includes multiple subsequent utterance texts predicted in sequence. The multiple predicted subsequent utterance text sequences can be sorted according to total confidence or total relevance. For example, "Markov Model to evolve" is ranked in the first position, which indicates that the speaker then tells the possibility of the predicted subsequent utterance text sequence with the highest. In addition, the speech recognition result 840b also includes multiple previous utterance texts identified before the current utterance text 840b-2. The current utterance text 840b-2 and the prompt 840b-4 are provided in different presentation modes.
[0090] In exemplary speech recognition result 840c, prompt only comprises a predicted subsequent utterance text sequence, but this speech recognition result 840c further comprises the relevant information of current utterance text and / or this prompt.As shown in the figure, speech recognition result 840c can comprise current utterance text 840c-2 "Hidden", prompt 840c-4 and relevant information 840c-6. Prompt 840c-4 comprises a predicted subsequent utterance text sequence "Markov Model to evolve".Related information 840c-6 is the description of "Hidden Markov Model" included in current utterance text 840c-2 and prompt 840c-4.In addition, speech recognition result 840c also comprises a plurality of previous utterance texts identified before current utterance text 840c-2.Current utterance text 840c-2, prompt 840c-4 and relevant information 840c-6 are provided in different presentation modes.
[0091] Fig. 9 An example 900 of providing prompts in real time in speech recognition results according to an embodiment is shown. Example 900 is intended to illustrate an exemplary process of providing prompts in real time as an audio stream proceeds.
[0092] Assume that the speech currently appearing in the audio stream 910 from the speaker corresponds to "Hidden", that is, the speech corresponding to "Hidden" is the current speech input 912. At 920, the current utterance text "Hidden" corresponding to the current speech input 912 can be recognized. In addition, at 930, a prompt can be generated based on at least the current utterance text, and the prompt includes a predicted subsequent utterance text sequence "Markov Model to evolve". A speech recognition result 940 corresponding to the current speech input 912 can be provided. The speech recognition result 940 can include the current utterance text 942 and the prompt 944.
[0093] Assume that in the audio stream 910 from the speaker, the speech corresponding to "Hidden" is followed by the speech corresponding to "Markov", that is, the speech corresponding to "Markov" becomes the current speech input 914. At 920, the current utterance text "Markov" corresponding to the current speech input 914 can be recognized. In addition, at 930, a prompt can be generated based on at least the current utterance text, and the prompt includes a predicted subsequent utterance text sequence "Model to evolvescenario". A speech recognition result 950 corresponding to the current speech input 914 can be provided. The speech recognition result 950 can include the current utterance text 952 and the prompt 954.
[0094] Referring to example 900 , as the audio stream progresses, corresponding prompts may be provided in real time in the speech recognition results for subsequent speech input.
[0095] Fig.10 An example 1000 of providing prompts in real time in speech recognition results according to an embodiment is shown. Example 1000 is intended to illustrate providing prompts in different ways in the case of mis-narration by a speaker.
[0096] Assume that the speech currently appearing in the audio stream 1010 from the speaker corresponds to "Hidden", that is, the speech corresponding to "Hidden" is the current speech input 1012. At 1020, the current utterance text "Hidden" corresponding to the current speech input 1012 can be recognized. In addition, at 1030, a prompt can be generated based on at least the current utterance text, and the prompt includes a predicted subsequent utterance text sequence "Markov Model to evolve". A speech recognition result 1040 corresponding to the current speech input 1012 can be provided. The speech recognition result 1040 can include the current utterance text 1042 and the prompt 1044.
[0097] Assume that in the audio stream 1010 from the speaker, the speech corresponding to "Macao" appears after the speech corresponding to "Hidden", that is, the speech corresponding to "Macao" becomes the current speech input 1014. In fact, the speaker mistakenly speaks "Macao" instead of the intended word "Markov". At 1020, the current utterance text "Macao" corresponding to the current speech input 1014 can be identified.
[0098] By comparing the current utterance text "Macao" with the predicted subsequent utterance text sequence "Markov Model to evolve" included in the prompt 1044 in the speech recognition result 1040, it can be determined that the speaker may have mistakenly spoken the word "Macao" from the predicted utterance text "Markov" in the prompt 1044. Prompts may be provided in the speech recognition result corresponding to the current speech input 1014 according to different strategies.
[0099] According to one strategy, the subsequent utterance text sequence can be predicted based on the incorrectly spoken word "Macao" to form a prompt. For example, at 1030, a prompt can be generated based at least on the current utterance text "Macao", and the prompt includes a predicted subsequent utterance text sequence "City to travel this year". A speech recognition result 1050a corresponding to the current speech input 1014 can be provided, which includes the current utterance text 1050a-2 and the prompt 1050a-4. The prompt 1050a-4 includes the predicted subsequent utterance text sequence "City to travel this year".
[0100] According to another strategy, the subsequent utterance text sequence can be predicted based on the word "Markov" determined to be correct through the above comparison to form a prompt. Figure 7 . For example, at 1030, a prompt may be generated based at least on the word "Markov", the prompt including a predicted subsequent utterance text sequence "Model to evolve". A speech recognition result 1050b corresponding to the current speech input 1014 may be provided, which includes the current utterance text 1050b-2 and the prompt 1050b-4. The prompt 1050b-4 includes an indication of the correct word "(Markov)" and the predicted subsequent utterance text sequence "Model to evolve".
[0101] According to another strategy, the above two strategies can be combined. For example, on the one hand, a subsequent utterance text sequence is predicted based on the word "Macao" told incorrectly to form a first prompt part, and on the other hand, another subsequent utterance text sequence is predicted based on the correct word "Markov" to form a second prompt part. Accordingly, a speech recognition result 1050c corresponding to the current voice input 1014 can be provided, which includes the current utterance text 1050c-2, the first prompt part 1050c-4 and the second prompt part 1050c-6. The first prompt part 1050c-4 includes the predicted subsequent utterance text sequence "City to travel this year", and the second prompt part 1050c-6 includes the indication "(Markov)" to the correct word and the predicted subsequent utterance text sequence "Model to evolve".
[0102] Fig.11 An example 1100 of providing hints in real time in speech recognition results according to an embodiment is shown. Example 1100 is intended to illustrate a way to provide hints when a speaker's speech omits words.
[0103] Assume that the speech currently appearing in the audio stream 1110 from the speaker corresponds to "Hidden", that is, the speech corresponding to "Hidden" is the current speech input 1112. At 1120, the current utterance text "Hidden" corresponding to the current speech input 1112 can be recognized. In addition, at 1130, a prompt can be generated based on at least the current utterance text, and the prompt includes a predicted subsequent utterance text sequence "Markov Model to evolve". A speech recognition result 1140 corresponding to the current speech input 1112 can be provided. The speech recognition result 1140 can include the current utterance text 1142 and the prompt 1144.
[0104] Assume that in the audio stream 1110 from the speaker, the speech corresponding to "Hidden" is followed by the speech corresponding to "Model", that is, the speech corresponding to "Model" becomes the current speech input 1114. In fact, the speaker omitted the word "Markov" in the phrase "Hidden Markov Model" that was intended to be spoken, that is, the word "Model" was spoken directly after the word "Hidden". At 1120, the current utterance text "Model" corresponding to the current speech input 1114 can be identified.
[0105] By comparing the current utterance text "Model" with the predicted subsequent utterance text sequence "Markov Model to evolve" included in the prompt 1144 in the speech recognition result 1140, it can be determined that the speaker may have omitted the predicted utterance text "Markov" in the prompt 1044. According to one strategy, the omitted word "Markov" can be used as the previous utterance text before the current utterance text "Model" so that it can be used to predict the subsequent utterance text sequence. Accordingly, it can be adopted Figure 7 . For example, at 1130, a subsequent utterance text sequence "to evolve scenario" can be predicted based on at least the current utterance text "Model" and the missing word "Markov" as the previous utterance text. A speech recognition result 1150 corresponding to the current speech input 1114 can be provided, which includes the current utterance text 1152, the first prompt part 1154 and the second prompt part 1156. The first prompt part 1154 includes an indication of the missing word "(Markov)", and the second prompt part 1156 includes the predicted subsequent utterance text sequence "to evolve scenario".
[0106] It should be understood that according to specific application requirements and designs, the above combination Figures 8 to 11 The described examples may be modified or combined in any manner. In addition, depending on the specific application scenario, the speech recognition results in these examples may be provided to, for example, end users, platform users, etc.
[0107] Fig.12 The flowchart of an exemplary method 1200 for providing prompts in real time in speech recognition results according to an embodiment is shown.
[0108] At 1210, current speech input in an audio stream for a target event may be obtained.
[0109] At 1220, current utterance text corresponding to the current voice input may be identified.
[0110] At 1230 , a prompt may be generated based at least on the current utterance text, the prompt comprising at least one predicted subsequent utterance text sequence.
[0111] At 1240 , a speech recognition result for the current speech input may be provided, the speech recognition result including the current utterance text and the prompt.
[0112] In one implementation, the current utterance text may include one or more words, each predicted subsequent utterance text sequence may include one or more predicted subsequent utterance texts, and each predicted subsequent utterance text may include one or more words.
[0113] In one implementation, the prompt may be generated further based on at least one previously spoken text recognized before the currently spoken text.
[0114] In one implementation, method 1200 may further include: obtaining an event ID of the target event and / or a speaker ID of a speaker of the current voice input. The prompt may be further generated based on the event ID and / or the speaker ID.
[0115] In one implementation, the generating the prompt may include: predicting the at least one predicted subsequent utterance text sequence based on at least the current utterance text by a pre-established prompt generator.
[0116] The prompt generator may be based on a neural network model.
[0117] The prompt generator may be based on a relevance matrix. The relevance matrix may include a plurality of text items and relevance values between the plurality of text items. The predicting of the at least one predicted subsequent utterance text sequence may include: selecting at least one text item sequence with the highest relevance ranking from the plurality of text items based at least on the current utterance text.
[0118] Method 1200 may also include: obtaining at least one document associated with the target event; and establishing the prompt generator based at least on the at least one document.
[0119] The method 1200 may also include: obtaining event information associated with the target event, the event information including at least one of the following: an event ID of the target event, time information of the target event, an agenda of the target event, a participant ID of a participant of the target event, and a speaker ID of a speaker associated with the at least one document. The prompt generator may be further established based on the event information.
[0120] Method 1200 may also include: obtaining a plurality of documents associated with a plurality of events and / or a plurality of speakers; and establishing the prompt generator based at least on the plurality of documents. Method 1200 may also include: obtaining event information respectively associated with the plurality of events and / or a plurality of speaker IDs respectively corresponding to the plurality of speakers. The prompt generator may be further established based on the event information and / or the plurality of speaker IDs.
[0121] In one implementation, the current utterance text and the prompt may be provided in different presentation modes.
[0122] In one implementation, method 1200 may also include: obtaining a second voice input in the audio stream that follows the current voice input; identifying a second utterance text corresponding to the second voice input; comparing the second utterance text with at least one predicted subsequent utterance text sequence; generating a second prompt based at least on the second utterance text and / or a result of the comparison; and providing a speech recognition result for the second voice input, the speech recognition result including the second utterance text and the second prompt.
[0123] In one implementation, the method 1200 may further include: determining relevant information of the current utterance text and / or the prompt. The speech recognition result may further include the relevant information.
[0124] It should be understood that method 1200 may also include any steps / processes for providing prompts in real time in speech recognition results according to the above-mentioned embodiments of the present disclosure.
[0125] Fig.13 An exemplary apparatus 1300 for providing hints in speech recognition results in real time according to an embodiment is shown.
[0126] The device 1300 may include: a voice input obtaining module 1310, used to obtain a current voice input in an audio stream for a target event; a speech text recognition module 1320, used to identify a current speech text corresponding to the current voice input; a prompt generation module 1330, used to generate a prompt based at least on the current speech text, the prompt including at least one predicted subsequent speech text sequence; and a voice recognition result providing module 1340, used to provide a voice recognition result for the current voice input, the voice recognition result including the current speech text and the prompt.
[0127] In one implementation, the apparatus 1300 may further include: an identification obtaining module, configured to obtain an event ID of the target event and / or a speaker ID of a speaker of the current voice input. The prompt may be further generated based on the event ID and / or the speaker ID.
[0128] In one implementation, the prompt generation module may be configured to: predict the at least one predicted subsequent utterance text sequence based at least on the current utterance text by using a pre-established prompt generator.
[0129] The apparatus 1300 may further include: a document obtaining module, configured to obtain at least one document associated with the target event; and a prompt generator establishing module, configured to establish the prompt generator based at least on the at least one document.
[0130] In addition, the device 1300 may also include any other modules that execute the steps of the method for providing prompts in real time in the speech recognition results according to the above-mentioned embodiment of the present disclosure.
[0131] Fig.14 An exemplary apparatus 1400 for providing hints in speech recognition results in real time according to an embodiment is shown.
[0132] The device 1400 may include: at least one processor 1410; and a memory 1420 storing computer executable instructions. When the computer executable instructions are executed, the at least one processor 1410 may: obtain a current voice input in an audio stream for a target event; identify a current utterance text corresponding to the current voice input; generate a prompt based at least on the current utterance text, the prompt including at least one predicted subsequent utterance text sequence; and provide a voice recognition result for the current voice input, the voice recognition result including the current utterance text and the prompt. In addition, the processor 1410 may also perform any other steps / processes of the method for providing prompts in real time in a voice recognition result according to the above-mentioned embodiment of the present disclosure.
[0133] The embodiments of the present disclosure may be implemented in a non-transitory computer-readable medium. The non-transitory computer-readable medium may include instructions that, when executed, cause one or more processors to perform any operation of the method for providing prompts in real time in speech recognition results according to the above-mentioned embodiments of the present disclosure.
[0134] It should be understood that all operations in the method described above are merely exemplary, and the present disclosure is not limited to any operation in the method or the order of these operations, but should cover all other equivalent transformations under the same or similar concept.
[0135] It should also be understood that all modules in the above described device can be implemented in various ways. These modules can be implemented as hardware, software, or a combination thereof. In addition, any module in these modules can be further divided into submodules or combined together functionally.
[0136] Processors have been described in conjunction with various devices and methods. These processors can be implemented using electronic hardware, computer software or any combination thereof. Whether these processors are implemented as hardware or software will depend on specific application and the overall design constraints imposed on the system. As an example, the processor provided in the present disclosure, any part of the processor or any combination of processors can be implemented as a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), a state machine, a gate logic, a discrete hardware circuit, and other suitable processing components configured to perform the various functions described in the present disclosure. The function of the processor provided in the present disclosure, any part of the processor or any combination of processors can be implemented as software performed by a microprocessor, a microcontroller, a DSP or other suitable platforms.
[0137] Software should be broadly considered to represent instructions, instruction sets, codes, code segments, program codes, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, running threads, processes, functions, etc. Software can reside in a computer-readable medium. A computer-readable medium can include, for example, a memory, which can be, for example, a magnetic storage device (e.g., a hard disk, a floppy disk, a magnetic stripe), an optical disk, a smart card, a flash memory device, a random access memory (RAM), a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a register, or a removable disk. Although the memory is shown as being separated from the processor in the various aspects provided in the present disclosure, the memory can also be located inside the processor (e.g., a cache or a register).
[0138] The above description is provided to enable any person skilled in the art to implement the various aspects described herein. Various modifications to these aspects are apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents of the elements of the various aspects described in this disclosure that are known or to be known to those skilled in the art will be covered by the claims.
Claims
1. A method for providing prompts in real time in speech recognition results, include: Get the current speech input in the audio stream for the target event; Identifying a current utterance text corresponding to the current voice input; generating a prompt based at least on the current utterance, the prompt comprising at least one predicted sequence of subsequent utterances; and providing a speech recognition result for the current speech input, the speech recognition result including the current utterance text and the prompt, Wherein, the generation prompt includes: Obtaining at least one document associated with the target event; establishing a prompt generator based at least on the at least one document to predict the at least one predicted subsequent utterance text sequence; and obtaining event information associated with the target event, the event information comprising at least one of the following: an event identifier (ID) of the target event, time information of the target event, an agenda of the target event, participant IDs of participants of the target event, and a speaker ID of a speaker associated with the at least one document, Wherein, the prompt generator is further established based on the event information.
2. The method according to claim 1, in, The current utterance text includes one or more words, and Each predicted subsequent utterance text sequence includes one or more predicted subsequent utterance texts, and each predicted subsequent utterance text includes one or more words.
3. The method according to claim 1, in, The prompt is further generated based on at least one previously uttered text recognized before the currently uttered text.
4. The method according to claim 1, further comprising: include: Get the speaker ID of the speaker of the current voice input, Wherein, the prompt is further generated based on the speaker ID.
5. The method according to claim 1, in, The at least one predicted subsequent utterance text sequence is predicted based on at least the current utterance text by the pre-established prompt generator.
6. The method according to claim 5, in, The prompt generator is based on a neural network model.
7. The method according to claim 5, in, The hint generator is based on a relevance matrix.
8. The method according to claim 7, in, The correlation matrix includes a plurality of text items and correlation values between the plurality of text items, and The predicting the at least one predicted subsequent utterance text sequence comprises: selecting at least one text item sequence with the highest relevance ranking from the plurality of text items based at least on the current utterance text.
9. The method according to claim 5, further comprising: include: obtaining a plurality of documents associated with a plurality of events and / or a plurality of speakers; as well as The prompt generator is established based at least on the plurality of documents.
10. The method according to claim 9, further comprising: include: obtaining event information respectively associated with the plurality of events and / or a plurality of speaker identifications (IDs) respectively corresponding to the plurality of speakers, Wherein, the prompt generator is further established based on the event information and / or the multiple speaker IDs.
11. The method according to claim 1, in, The current utterance text and the prompt are provided in different presentation modes.
12. The method of claim 1, further comprising: include: Obtaining a second voice input in the audio stream that follows the current voice input; recognizing a second utterance text corresponding to the second voice input; comparing the second utterance to the at least one predicted sequence of subsequent utterances; generating a second prompt based at least on the second utterance text and / or a result of the comparison; as well as A speech recognition result is provided for the second voice input, the speech recognition result including the second utterance text and the second prompt.
13. The method of claim 1, further comprising: include: determining information related to the current utterance text and / or the prompt, and Wherein, the speech recognition result also includes the relevant information.
14. A device for providing prompts in real time in speech recognition results, include: A voice input acquisition module, used to obtain current voice input in an audio stream for a target event; A speech text recognition module, used to recognize the current speech text corresponding to the current voice input; a prompt generation module, configured to generate a prompt based at least on the current utterance text, the prompt comprising at least one predicted subsequent utterance text sequence; A speech recognition result providing module, used for providing a speech recognition result for the current speech input, wherein the speech recognition result includes the current speech text and the prompt; A document obtaining module, used to obtain at least one document associated with the target event; a prompt generator building module for building a prompt generator based at least on the at least one document to predict the at least one predicted subsequent utterance text sequence; as well as an identification obtaining module, configured to obtain event information associated with the target event, the event information comprising at least one of the following: an event identification (ID) of the target event, time information of the target event, an agenda of the target event, a participant ID of a participant of the target event, and a speaker ID of a speaker associated with the at least one document, Wherein, the prompt generator is further established based on the event information.
15. The device according to claim 14, in: The identification acquisition module is used to obtain the speaker ID of the speaker of the current voice input, Wherein, the prompt is further generated based on the speaker ID.
16. The device according to claim 14, in, The prompt generation module is used for: The at least one predicted subsequent utterance text sequence is predicted based on at least the current utterance text by the pre-established prompt generator.
17. A device for providing prompts in real time in speech recognition results, include: at least one processor; as well as a memory storing computer executable instructions that, when executed, cause the at least one processor to: Get the current speech input in the audio stream for the target event, identifying a current utterance text corresponding to the current voice input, generating a prompt based at least on the current utterance, the prompt comprising at least one predicted sequence of subsequent utterances, and providing a speech recognition result for the current speech input, the speech recognition result including the current utterance text and the prompt, Wherein, the generation prompt includes: Obtaining at least one document associated with the target event; establishing a prompt generator based at least on the at least one document to predict the at least one predicted subsequent utterance text sequence; and obtaining event information associated with the target event, the event information comprising at least one of the following: an event identifier (ID) of the target event, time information of the target event, an agenda of the target event, participant IDs of participants of the target event, and a speaker ID of a speaker associated with the at least one document, Wherein, the prompt generator is further established based on the event information.
Citation Information
Patent Citations
Speech recognition with contextual hypothesis probabilities
EP1160767A2
Correction of previous words and other user text input errors
US20160275070A1
Method and apparatus for facilitating persona-based agent interactions with online visitors
US20200098366A1