Long speech recognition method based on large language model post-processing and electronic equipment
By cascading the long speech recognition model to a large language model, the problem of complex model training in the prior art is solved and the problem of inability to inherit the existing model knowledge is realized, and efficient long speech recognition and context information introduction are achieved.
Patent Information
- Application Number
- CN202510199065.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-13
AI Technical Summary
The existing technology requires training of brand new models in long speech recognition, resulting in high data requirements, complex model structure and inability to inherit the knowledge of existing models.
The post-processing method based on the large language model is adopted to cascade the long speech recognition model to the large language model, and the existing speech recognition model is connected to the large language model through a cascading architecture to avoid retraining the model.
It realizes that the existing model capabilities are fully utilized without training a new model, simplifies the model structure, improves the accuracy of long speech recognition, and effectively introduces context information.
Smart Images

Figure CN119993136A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent speech, and in particular to a long speech recognition method, system, electronic device and storage medium based on large language model post-processing. Background Art
[0002] In the long speech recognition task, rich contextual information is crucial to improve the recognition accuracy. Effective use of this contextual information can significantly improve the effect of long speech recognition. For streaming speech recognition tasks, due to the need for real-time processing, when recognizing the current audio segment, the system can only rely on the shorter previous text or historical information that has been received to recognize the current audio segment. Therefore, how to make full use of the existing previous audio or recognition text to enhance the recognition effect of the current audio segment is a research direction for long speech recognition.
[0003] For long speech recognition, existing technologies usually use the following methods: Prompt ASR for contextualized ASR with controllable style (PROMPTASR for contextualized ASR with controllable style), this method encodes the recognition text of the previous text and then uses the cross-attention mechanism to fuse the encoded information into the audio encoder to improve the accuracy of long speech recognition in subsequent speech.
[0004] LLM (Large Language Model) combines speech recognition with LLM, which has excellent contextual reasoning capabilities, to more effectively utilize contextual information, thereby improving the accuracy of speech recognition.
[0005] In the process of implementing the present invention, the inventors found that there are at least the following problems in the related art: The above methods all require training a completely new model. However, this will result in: 1. High-quality training data is required: Retraining the model requires a large amount of carefully designed audio-text pairs as training data. This not only increases the workload of data collection and preprocessing, but may also affect model performance due to insufficient or unbalanced data; 2. Complex model structure and training strategy: After the introduction of context information, additional model components need to be introduced, such as text encoders, cross-attention modules, and audio encoders. Model structure design and training strategies become more complex, which not only increases the difficulty of model design, but also means that more computing resources are needed to support the training and reasoning process. And due to the change in model structure, the existing speech recognition model cannot be used directly and needs to be retrained.
[0006] 3. Unable to inherit knowledge from existing models: For existing speech recognition models, adopting a completely new model architecture often means that the knowledge learned by the old model cannot be directly inherited, resulting in a waste of resources. Summary of the invention
[0007] In order to at least solve the problems in the prior art that existing models are difficult to inherit, have high model data requirements and complex training.
[0008] In a first aspect, an embodiment of the present invention provides a long speech recognition method based on large language model post-processing, comprising: Continuously inputting the long speech into a streaming speech recognition model cascaded with a large language model, and performing speech recognition on the continuously input long speech as ordered i short audio segments in the streaming speech recognition model; Determine N candidate recognition texts and corresponding speech recognition scores for the j-th short audio segment, where 1≤j≤i; The N candidate recognition texts are respectively concatenated with the previous text to obtain the N concatenated texts, and the concatenated texts are input into the large language model through cascading to obtain context understanding scores corresponding to the N candidate recognition texts; Determine the final recognized text of the j-th short audio segment from the N candidate recognized texts based on the speech recognition score and the context understanding score, and use the final recognized text of the j-th short audio segment as the preceding text of the j+1-th short audio segment; When it is detected that the long speech input is completed, the recognition result of the long speech is generated in order using the final recognition texts of the short audio segments.
[0009] In a second aspect, an embodiment of the present invention provides a long speech recognition system based on large language model post-processing, comprising: A segment recognition module, used for continuously inputting a long speech into a streaming speech recognition model cascaded with a large language model, and treating the continuously input long speech as ordered i short audio segments for speech recognition in the streaming speech recognition model; A score determination module, used to determine N candidate recognition texts and corresponding speech recognition scores for the j-th short audio segment, where 1≤j≤i; A concatenation module, used for concatenating the N candidate recognition texts with the previous texts respectively to obtain the N concatenated texts, and inputting the concatenated texts into the large language model through cascading to obtain context understanding scores corresponding to the N candidate recognition texts; a final text determination module, configured to determine a final recognized text of the j-th short audio segment from the N candidate recognized texts based on the speech recognition score and the context understanding score, and use the final recognized text of the j-th short audio segment as a preceding text of the j+1-th short audio segment; The recognition module is used to generate the recognition result of the long speech in an orderly manner by using the final recognition texts of each short audio segment when it is detected that the long speech input is completed.
[0010] According to a third aspect, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the steps of the long speech recognition method based on large language model post-processing according to any embodiment of the present invention.
[0011] In a fourth aspect, an embodiment of the present invention provides a storage medium having a computer program stored thereon, characterized in that when the program is executed by a processor, the steps of the long speech recognition method based on large language model post-processing of any embodiment of the present invention are implemented.
[0012] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program / instructions, characterized in that when the computer program / instructions are executed by a processor, the steps of the long speech recognition method based on large language model post-processing of any embodiment of the present invention are implemented.
[0013] The beneficial effects of the embodiments of the present invention are: the speech recognition model is cascaded with a large language model, the deployment is flexible and elastic, no training is required, the capabilities of the existing model can be fully utilized, and no additional model structure is required. The large language model can be used as post-processing to introduce contextual information and improve the accuracy of long speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0015] Figure 1 is a flow chart of a long speech recognition method based on large language model post-processing provided by one embodiment of the present invention; Figure 2 is a specific schematic flow chart of a long speech recognition method based on large language model post-processing provided by one embodiment of the present invention; Figure 3 It is a schematic diagram of the structure of a long speech recognition system based on large language model post-processing provided by one embodiment of the present invention; Figure 4 A schematic diagram of the structure of an embodiment of an electronic device for long speech recognition based on large language model post-processing provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0017] like Figure 1 The flowchart of a long speech recognition method based on large language model post-processing provided by an embodiment of the present invention includes the following steps: S11: continuously inputting the long speech into a streaming speech recognition model cascaded with a large language model, and performing speech recognition on the continuously input long speech as ordered i short audio segments in the streaming speech recognition model; S12: Determine N candidate recognition texts and corresponding speech recognition scores for the j-th short audio segment, where 1≤j≤i; S13: splicing the N candidate recognition texts with the previous texts respectively to obtain the N spliced texts, and inputting the spliced texts into the large language model through cascading to obtain the context understanding scores corresponding to the N candidate recognition texts; S14: determining a final recognized text of the j-th short audio segment from the N candidate recognized texts based on the speech recognition score and the context understanding score, and using the final recognized text of the j-th short audio segment for a preceding text of the j+1-th short audio segment; S15: When it is detected that the long speech input is completed, the recognition result of the long speech is generated in order using the final recognition texts of the short audio segments.
[0018] In this embodiment, in order to realize long speech recognition in the prior art, solutions are usually considered from the aspects of data and model parameters, such as increasing the data scale, adjusting the model structure and parameters, etc. to improve the model performance. However, this will encounter the problem that the existing model is difficult to inherit, the model data requirements are high, and the training is complex.
[0019] In order to address the shortcomings of existing technologies, this method mainly considers the following two core points: 1. How to maximize the use of existing model capabilities and avoid the need for retraining without training a new model; 2. How to model contextual information in the simplest way. For these two points, this method uses a cascade architecture and a pre-trained large language model to solve them.
[0020] For step S11, the streaming speech recognition model used in this method can be a speech recognition model in the prior art, but it needs to be cascaded with the LLM (Large Language Model). In this way, the cascade architecture can be used to connect the existing speech recognition model with the large language model without the need for training from scratch, saving a lot of time and computing resources.
[0021] Furthermore, objective analysis is performed on long speech. Although the speech is long, considering the user's speaking habits, intonation, rhythm, etc., the long speech is not entirely uninterrupted speech. It is usually composed of a series of ordered short audio clips. The long speech is divided into multiple clips according to speaking habits, intonation, rhythm, and these clips can be regarded as short sentences. In the streaming speech recognition system, these short sentences will be recognized one by one in the order they appear in the audio and the corresponding text will be output.
[0022] Based on the above considerations, the long audio stream is sent to the speech recognition model in the form of chunks. Among them, since the present application is a scenario of long audio recognition, the conversation is usually not completed with a few short words such as "wake up". For example, in the AI call summary scenario, the user continuously inputs long voice in the call, and no one can be sure how long the call will be, and how long the audio is. It may be 1 minute, 10 minutes, or even 1 hour of long voice. Therefore, considering that the server cannot know the content size of the long audio in advance, and considering that if it is a full-duplex conversation, it is necessary to continuously identify the user's long audio, and feedback may be sent to the user at the same time, therefore, the transmission is in the form of chunks. Furthermore, these short sentences (that is, short audio clips) in the long voice are continuously sent to the streaming speech recognition system in the prior art for speech recognition.
[0023] For step S12, for each short audio clip, N possible recognition texts and their respective speech recognition scores are obtained, which are called N-Best and N-Best ASR scores. For the first short sentence, N possible recognition texts are identified and the ASR score results are evaluated for each of the N possible sentences.
[0024] For steps S13 and S14, after obtaining N candidate recognition texts, they are concatenated with the previous text to obtain the N concatenated texts.
[0025] The preceding text used for concatenating the N candidate recognition texts of the j-th short audio segment includes the final recognition text of the j-1-th short audio segment; When j is 1, the final recognition text is determined directly using the speech recognition scores corresponding to the N candidate recognition texts.
[0026] In this embodiment, since the recognition is performed sequentially in a streaming manner, for the first short audio segment input, since there is no preceding text, the ASR scores are directly sorted, and the one with the highest score is selected as the final recognized text of the current first short audio segment. At this time, the final recognized text of the first short audio segment is used as the preceding text of the second short audio segment.
[0027] In the same way, determine the N candidate recognition texts of the second short audio clip and the ASR score result. And concatenate the N candidate recognition texts of the second short audio clip with the previous text and input them into the large language model for context understanding, where the large language model includes: general dialogue basic model, LLaMA model, Qwen model. Obtain the corresponding context understanding score (LLM score). Furthermore, the defects of the streaming speech recognition model in the prior art in understanding the long speech context are avoided. The context information for modeling the speech recognition model is simplified through the large language model.
[0028] As an implementation manner, determining the final recognized text of the j-th short audio segment from the N candidate recognized texts based on the speech recognition score and the context understanding score includes: Determine the sum of the speech recognition scores and the context understanding scores of each of the N candidate recognition texts; The candidate recognition text with the highest total score is used as the final recognition text of the j-th short audio segment.
[0029] In this embodiment, the N texts are sorted according to the total result of the ASR score and the LLM score, and the text with the highest total score is selected as the final recognized text of the current short sentence. At this time, the final recognized text of the second short audio segment is used for the preceding text of the third short audio segment. For example, the preceding text of the fifth short audio segment includes the final recognized text of the first four segments.
[0030] Specifically, for example, the speech recognition system recognizes N results: text_1_1-text_1_N (where text_1 refers to the first short sentence, text_1_n refers to the nth possible recognition result in the first short sentence, 1≤n≤N). Suppose text_1_6 has the highest ASR score.
[0031] As an implementation mode, when it is only the first short sentence, the result with the highest ASR score can be directly used as the current final recognition text, for example, the current final recognition text is text_1_6.
[0032] Alternatively, text_1_1-text_1_N may be input into LLM to obtain the corresponding LLM score result. The current final recognition text is determined based on the highest comprehensive score of the ASR score and the LLM score.
[0033] At this time, text_1_6 is used as the previous text of the second short sentence. Similarly, for the second short sentence, N possible recognition texts and the N possible sentences text_2_1- text_2_N are identified, and the ASR score results are evaluated respectively.
[0034] Then, text_1_6 is combined with text_2_1- text_2_N respectively, and the concatenated N results are input into LLM to obtain the LLM score result.
[0035] If the highest combined score of ASR score and LLM score is text_1_6- text_2_3, then text_2_3 is used as the preceding text of the third short sentence, which may include text_1_6 and text_2_3. Based on the same steps above, each short sentence in the long speech is recognized in turn, and the final recognition text is determined by continuous splicing until the long speech recognition is completed.
[0036] In step S15, when it is detected that the long speech input is completed, the recognition result of the long speech is generated in order using the final recognition texts of the short audio segments. For example, text_1_6, text_2_3, text_3_8, etc. are spelled out in order to generate the recognition result of the long speech.
[0037] As an implementation manner, after obtaining the context understanding scores corresponding to the N candidate recognition texts, the method further includes: Determine the final recognized text of the j-th short audio segment from the concatenated texts corresponding to the N candidate recognized texts based on the speech recognition score and the context understanding score, and use the final recognized text of the j-th short audio segment as the preceding text of the j+1-th short audio segment; When it is detected that the long speech input is completed, the final recognition text of the i-th short audio segment is used as the recognition result of the long speech.
[0038] This method also provides another recognition method, which tests the ability of large language models more. Large language models have more input information, which can improve the recognition effect of long speech.
[0039] Taking the above example as an example, if the spliced text_1_6- text_2_3 is used as the preceding text of the third short sentence, the preceding text is no longer the separate words "text_1_6" and "text_2_3", but a continuously expanding spliced segment "text_1_6- text_2_3", which will be continuously expanded in the subsequent short sentence recognition.
[0040] Based on the same steps as above, each short sentence in the long speech is recognized in turn, and the final recognized text is continuously spliced and expanded until the long speech recognition is completed, and the final recognized text of the last short audio segment is obtained.
[0041] In summary, the overall process of this method is as follows Figure 2 As shown, the following steps are included: Step 1: Feed the long audio stream into the speech recognition model in fixed chunks. Step 2: The speech recognition model decodes the input chunk and obtains multiple candidate texts, namely the N-best recognized texts and the corresponding N-best scores; Step 3: Splice the obtained multiple candidate texts with the previous texts respectively to obtain N newly spliced text contents; for the first chunk, since there is no historical information, the previous text is empty; Step 4: Send the N newly concatenated texts to the large language model for secondary re-scoring to obtain the LLM scores of the N concatenated texts; Step 5: Calculate the total score based on the LLM score result and the N-best score, and then re-rank the N-best recognized texts; Step 6: Based on the sorting results, select the result with the highest score as the final text; Step 7: If the audio stream has not ended, repeat steps 2 to 6 until the audio stream ends; It can be seen from this implementation that the speech recognition model is cascaded with a large language model, which is flexible and elastic to deploy, does not require training, can fully utilize the capabilities of the existing model, and can apply the large language model as post-processing without the need for an additional model structure to introduce contextual information and improve the accuracy of long speech recognition.
[0042] like Figure 3 The figure shows a schematic diagram of the structure of a long speech recognition system based on large language model post-processing provided by an embodiment of the present invention. The system can execute the long speech recognition method based on large language model post-processing described in any of the above embodiments and be configured in a terminal.
[0043] The present embodiment provides a long speech recognition system 10 based on large language model post-processing, comprising: a segment recognition module 11 , a score determination module 12 , a splicing module 13 , a final text determination module 14 and a recognition module 15 .
[0044] Among them, the segment recognition module 11 is used to continuously input the long speech into the streaming speech recognition model cascaded with the large language model, and the continuously input long speech is used as ordered i short audio segments for speech recognition in the streaming speech recognition model; the score determination module 12 is used to determine N candidate recognition texts and corresponding speech recognition scores of the j-th short audio segment, wherein 1≤j≤i; the splicing module 13 is used to splice the N candidate recognition texts with the previous texts respectively to obtain the N spliced texts, and input the spliced texts into the large language model through cascading to obtain the context understanding scores corresponding to the N candidate recognition texts; the final text determination module 14 is used to determine the final recognition text of the j-th short audio segment from the N candidate recognition texts based on the speech recognition score and the context understanding score, and use the final recognition text of the j-th short audio segment for the previous text of the j+1-th short audio segment; the recognition module 15 is used to generate the recognition result of the long speech in order using the final recognition texts of each short audio segment when it is detected that the long speech input is completed.
[0045] The embodiment of the present invention further provides a non-volatile computer storage medium, the computer storage medium stores computer executable instructions, and the computer executable instructions can execute the long speech recognition method based on large language model post-processing in any of the above method embodiments; As an implementation mode, the non-volatile computer storage medium of the present invention stores computer executable instructions, and the computer executable instructions are configured as follows: Continuously inputting the long speech into a streaming speech recognition model cascaded with a large language model, and performing speech recognition on the continuously input long speech as ordered i short audio segments in the streaming speech recognition model; Determine N candidate recognition texts and corresponding speech recognition scores for the j-th short audio segment, where 1≤j≤i; The N candidate recognition texts are respectively concatenated with the previous text to obtain the N concatenated texts, and the concatenated texts are input into the large language model through cascading to obtain context understanding scores corresponding to the N candidate recognition texts; Determine the final recognized text of the j-th short audio segment from the N candidate recognized texts based on the speech recognition score and the context understanding score, and use the final recognized text of the j-th short audio segment as the preceding text of the j+1-th short audio segment; When it is detected that the long speech input is completed, the recognition result of the long speech is generated in order using the final recognition texts of the short audio segments.
[0046] As a non-volatile computer-readable storage medium, it can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the method in the embodiment of the present invention. One or more program instructions are stored in the non-volatile computer-readable storage medium, and when executed by the processor, the long speech recognition method based on large language model post-processing in any of the above method embodiments is executed.
[0047] Figure 4 is a schematic diagram of the hardware structure of an electronic device of a long speech recognition method based on large language model post-processing provided by another embodiment of the present application, such as Figure 4 As shown, the device includes: One or more processors 410 and memory 420, Figure 4 A processor 410 is taken as an example. The device of the long speech recognition method based on large language model post-processing may further include: an input device 430 and an output device 440.
[0048] The processor 410, the memory 420, the input device 430 and the output device 440 may be connected via a bus or other means. Figure 4 The example of connecting through bus is taken in the following.
[0049] The memory 420, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the long speech recognition method based on large language model post-processing in the embodiment of the present application. The processor 410 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory 420, that is, the long speech recognition method based on large language model post-processing in the above method embodiment is implemented.
[0050] The memory 420 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data, etc. In addition, the memory 420 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 420 may optionally include a memory remotely arranged relative to the processor 410, and these remote memories may be connected to the mobile device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0051] The input device 430 can receive input digital or character information. The output device 440 can include a display device such as a display screen.
[0052] The one or more modules are stored in the memory 420, and when executed by the one or more processors 410, the long speech recognition method based on large language model post-processing in any of the above method embodiments is executed.
[0053] The above-mentioned product can execute the method provided in the embodiment of the present application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of the present application.
[0054] The non-volatile computer-readable storage medium may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the device, etc. In addition, the non-volatile computer-readable storage medium may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0055] An embodiment of the present invention also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the steps of the long speech recognition method based on large language model post-processing of any embodiment of the present invention.
[0056] The electronic device of the embodiment of the present application exists in various forms, including but not limited to: (1) Mobile communication equipment: This type of equipment is characterized by having mobile communication functions and its main purpose is to provide voice and data communications. This type of terminal includes: smart phones, multimedia phones, functional phones, and low-end phones.
[0057] (2) Ultra-mobile personal computer devices: These devices fall into the category of personal computers, have computing and processing capabilities, and generally also have mobile Internet access features. These terminals include: PDAs, MIDs, and UMPC devices, such as tablet computers.
[0058] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players, handheld game consoles, e-books, smart toys, and portable car navigation devices.
[0059] (4) Other electronic devices with data processing functions.
[0060] In this article, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include" and "comprise" include not only those elements, but also other elements not explicitly listed, or also include elements inherent to such processes, methods, articles or equipment. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the existence of other identical elements in the process, method, article or equipment that includes the elements.
[0061] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0062] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A long speech recognition method based on large language model post-processing, comprising: Continuously inputting the long speech into a streaming speech recognition model cascaded with a large language model, and performing speech recognition on the continuously input long speech as ordered i short audio segments in the streaming speech recognition model; Determine N candidate recognition texts and corresponding speech recognition scores for the j-th short audio segment, where 1≤j≤i; The N candidate recognition texts are respectively concatenated with the previous text to obtain the N concatenated texts, and the concatenated texts are input into the large language model through cascading to obtain context understanding scores corresponding to the N candidate recognition texts; Determine the final recognized text of the j-th short audio segment from the N candidate recognized texts based on the speech recognition score and the context understanding score, and use the final recognized text of the j-th short audio segment as the preceding text of the j+1-th short audio segment; When it is detected that the long speech input is completed, the recognition result of the long speech is generated in order using the final recognition texts of the short audio segments.
2. The method according to claim 1, wherein: The step of determining the final recognized text of the j-th short audio segment from the N candidate recognized texts based on the speech recognition score and the context understanding score comprises: Determine the sum of the speech recognition scores and the context understanding scores of each of the N candidate recognition texts; The candidate recognition text with the highest total score is used as the final recognition text of the j-th short audio segment.
3. The method according to claim 1, wherein: The previous text used for concatenating the N candidate recognition texts of the j-th short audio segment includes the final recognition text of the j-1-th short audio segment; When j is 1, the final recognized text is determined directly using the speech recognition scores corresponding to the N candidate recognized texts.
4. The method according to claim 1, wherein: After obtaining the context understanding scores corresponding to the N candidate recognition texts, the method further includes: Determine the final recognized text of the j-th short audio segment from the concatenated texts corresponding to the N candidate recognized texts based on the speech recognition score and the context understanding score, and use the final recognized text of the j-th short audio segment as the preceding text of the j+1-th short audio segment; When it is detected that the long speech input is completed, the final recognition text of the i-th short audio segment is used as the recognition result of the long speech.
5. The method according to claim 1, wherein: The large language model cascaded with the streaming speech recognition model includes: a general dialogue basic model, an LLaMA model, and a Qwen model.
6. A long speech recognition system based on large language model post-processing, comprising: A segment recognition module, used for continuously inputting a long speech into a streaming speech recognition model cascaded with a large language model, and treating the continuously input long speech as ordered i short audio segments for speech recognition in the streaming speech recognition model; A score determination module, used to determine N candidate recognition texts and corresponding speech recognition scores for the j-th short audio segment, where 1≤j≤i; A concatenation module, used for concatenating the N candidate recognition texts with the previous texts respectively to obtain the N concatenated texts, and inputting the concatenated texts into the large language model through cascading to obtain context understanding scores corresponding to the N candidate recognition texts; a final text determination module, configured to determine a final recognized text of the j-th short audio segment from the N candidate recognized texts based on the speech recognition score and the context understanding score, and use the final recognized text of the j-th short audio segment as a preceding text of the j+1-th short audio segment; The recognition module is used to generate the recognition result of the long speech in an orderly manner by using the final recognition texts of each short audio segment when it is detected that the long speech input is completed.
7. The method according to claim 6, wherein: The final text determination module is also used for: Determine the final recognized text of the j-th short audio segment from the concatenated texts corresponding to the N candidate recognized texts based on the speech recognition score and the context understanding score, and use the final recognized text of the j-th short audio segment as the preceding text of the j+1-th short audio segment; The recognition module is also used for: when it is detected that the long speech input is completed, taking the final recognition text of the i-th short audio segment as the recognition result of the long speech.
8. A storage medium having a computer program product stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 4 are implemented.
9. A computer program product having instructions embedded on a storage medium, wherein the instructions implement the steps of the method according to any one of claims 1 to 4.
10. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method described in any one of claims 1 to 4.
Citation Information
Patent Citations
Streaming speech recognition method, terminal equipment and medium
CN113838468A
Speech recognition method and device, equipment and storage medium
CN115240685A
Speech recognition method, device and system, electronic equipment and readable storage medium
CN116434771A
Speech recognition method and device, equipment and medium
CN118942462A