Speech Recognition Method, System, Electronic Device and Storage Medium

By searching the audio stream for the first and second decoding paths, and combining CPU and GPU thread processing, the problem of low tail accuracy of real-time voice recognition is solved, and the feedback efficiency of call center customer service is improved.

CN114743540BActive Publication Date: 2025-07-29CTRIP TRAVEL INFORMATION TECH (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210392131.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-14
Publication Date
2025-07-29
Estimated Expiration
2042-04-14

AI Technical Summary

Technical Problem

In the prior art, the tail accuracy of real-time voice recognition results is low, which affects the timely feedback from the call center customer service.

Method used

The first decoding path search for the audio stream containing multi-frame voice frames is performed in real-time speech recognition, the audio stream is split at the pause, and the second decoding path search is performed based on the target audio band to obtain real-time speech recognition results, and the processing is combined with CPU and GPU threads, and the audio data is stored and separated using a streaming media module.

Benefits of technology

It improves the accuracy of real-time voice recognition, especially the recognition accuracy of audio tail, and supports timely and efficient feedback from call center customer service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114743540B_ABST
    Figure CN114743540B_ABST
Patent Text Reader

Abstract

The present invention discloses a voice recognition method, system, electronic device and storage medium. The voice recognition method includes: performing a first decoding path search for real-time voice recognition on an audio stream containing multiple frames of voice frames, and using the voice frame currently undergoing voice recognition as the target frame; if a pause is recognized after the target frame, splitting the audio stream at the pause to obtain a target audio segment; and performing a second decoding path search based on the target audio segment to obtain the real-time voice recognition result of the target frame. In the process of performing real-time voice recognition, the voice recognition method provided by the present invention uses the audio frame currently being recognized as the target frame, and does not output the recognition result of the target frame for the time being. Instead, after reaching the end point of the audio segment, a secondary decoding path search is performed on the target frame based on the entire audio segment, and then the best recognition result is output, effectively overcoming the defect of low recognition accuracy at the tail of the audio during real-time voice recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly relates to a voice recognition method, system, electronic device, and storage medium. Background Art

[0002] As a connection hub between an OTA (Online Travel Agency) company and customers, a call center is an important part of the entire service chain. For OTA customer service, if "speech-to-text conversion while speaking" can be supported during the voice call with a hotel or a customer, that is, the dialogue audio can be real-time transcribed into text information through a speech recognition algorithm, it will be convenient for the customer service to effectively respond to the customer's consultation in a timely and accurate manner.

[0003] Based on this, in the prior art, streaming speech recognition is usually used in combination with streaming media to achieve the above "speech-to-text conversion while speaking". Specifically, first, the audio data is compressed through streaming media and sent and stored in segments in the network in the form of a stream, so as to obtain the audio of the ongoing call; then, the audio stream in the streaming media is recognized and decoded through streaming speech recognition, and the corresponding text information is output, so as to continuously return the text transcribed by the speech recognition during the call duration to support the customer service staff to make timely feedback.

[0004] However, as Figure 1 shown, the process of streaming speech recognition is essentially a process of online inference. Since online inference cannot know the audio information generated after the current moment, compared with offline inference, the context information that can be used during inference is much less than that of offline inference. Based on this, in the streaming speech recognition algorithm, although the intermediate recognition results of real-time decoding are continuously output as the call progresses, since these intermediate recognition results lack sufficient context information for assistance in recognition, their accuracy at the end is usually relatively low, which is not conducive to supporting the customer service to accurately and timely obtain information and make feedback. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the defect that the accuracy of the recognition result at the end of real-time speech recognition in the prior art is relatively low, and to provide a voice recognition method, system, electronic device, and storage medium.

[0006] The present invention solves the above technical problem through the following technical solutions:

[0007] In a first aspect, the present invention provides a voice recognition method, and the voice recognition method includes:

[0008] Performing a first decoding path search for real-time speech recognition on an audio stream containing multiple speech frames, and using the speech frame currently undergoing speech recognition as the target frame;

[0009] If a pause is recognized after the target frame, the audio stream is split at the pause to obtain a target audio segment;

[0010] Based on the target audio segment, a second decoding path search is performed to obtain a real-time speech recognition result for the target frame.

[0011] Preferably, before the step of performing the first decoding path search for real-time speech recognition on an audio stream containing multiple speech frames, the method further includes:

[0012] Obtain an initial audio stream;

[0013] If the initial audio stream includes at least two channels, perform channel separation on the initial audio stream to obtain the target audio stream.

[0014] Preferably, before the step of performing the second decoding path search based on the target audio segment, the method further includes:

[0015] If the number of audio frames after the target frame in the target audio segment is less than a preset value, the last frame of the target audio segment is copied multiple times until the number of audio frames after the target frame in the target audio segment is greater than or equal to the preset value.

[0016] Preferably, the target audio segment includes ID information;

[0017] The speech recognition method further includes:

[0018] Perform speech recognition on several target audio segments with the same ID information to update the real-time speech recognition result.

[0019] In a second aspect, the present invention provides a speech recognition system, which includes:

[0020] A decoder module, configured to perform real-time speech recognition on a target audio stream and use the speech frame currently undergoing speech recognition as the target frame, and the real-time speech recognition includes a first decoding path search;

[0021] An audio splitting module, configured to split the audio stream at the pause when a pause after the target frame is recognized to obtain a target audio segment;

[0022] The decoder module is further configured to perform a second decoding path search based on the target audio segment to obtain a real-time speech recognition result for the target frame.

[0023] Preferably, the system further includes:

[0024] A streaming media module, configured to obtain an initial audio stream;

[0025] A channel separation module, configured to separate channels of the initial audio stream to obtain the target audio stream when the initial audio stream includes at least two channels;

[0026] Optionally, the real-time speech recognition further includes feature extraction, acoustic score calculation, and obtaining a decoding result. The decoder module includes a CPU thread and a GPU thread. The CPU thread is configured to perform the feature extraction and obtain the decoding result, and the GPU thread is configured to perform the acoustic score calculation, the first decoding path search, and the second decoding path search. Data is transmitted between the GPU thread and the CPU thread in a shared queue manner.

[0027] Preferably, the decoder module is further configured to:

[0028] When the number of audio frames after the target frame in the target audio segment is less than a preset value, copy the last frame of the target audio segment multiple times until the number of audio frames after the target frame in the target audio segment is greater than or equal to the preset value.

[0029] Preferably, the target audio segment includes ID information;

[0030] The decoder module is further configured to: perform speech recognition on a plurality of target audio segments with the same ID information to update the real-time speech recognition result.

[0031] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the speech recognition method as described above when executing the computer program.

[0032] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, wherein the computer program implements the speech recognition method as described above when executed by a processor.

[0033] The positive and progressive effects of the present invention are as follows: During the real-time speech recognition process of the speech recognition method provided by the present invention, the currently recognized audio frame is used as the target frame, and the recognition result of the target frame is not output temporarily. Instead, after reaching the end point of the audio segment, a second decoding path search is performed on the target frame based on the entire audio segment, and then the best recognition result is output, thereby realizing more accurate speech recognition based on the context information corresponding to the target frame, effectively overcoming the defect of low recognition accuracy at the end of the audio during real-time speech recognition, and improving the user experience. Description of the Drawings

[0034] Figure 1Schematic diagram for comparing the quantity of context information available for online inference and offline inference in the prior art.

[0035] Figure 2 First process schematic diagram of the speech recognition method according to Embodiment 1 of the present invention.

[0036] Figure 3 Schematic diagram of the workflow of streaming real-time speech recognition in the prior art.

[0037] Figure 4 Schematic diagram of the sub-steps of Step S1 of the speech recognition method according to Embodiment 1 of the present invention.

[0038] Figure 5 Schematic diagram of the specific implementation scenario of the speech recognition method according to Embodiment 1 of the present invention.

[0039] Figure 6 Schematic diagram of data transmission between the streaming media and the decoder module of the speech recognition method according to Embodiment 1 of the present invention.

[0040] Figure 7 Module schematic diagram of the speech recognition system according to Embodiment 2 of the present invention.

[0041] Figure 8 Dual-thread schematic diagram inside the decoder module of the speech recognition system according to Embodiment 2 of the present invention.

[0042] Figure 9 Structural schematic diagram of the electronic device for implementing the speech recognition method according to Embodiment 3 of the present invention. Specific implementation manners

[0043] The present invention will be further described below by way of examples, but the present invention is not limited to the scope of the described examples.

[0044] Embodiment 1

[0045] This embodiment discloses a speech recognition method. As Figure 2 shown, the speech recognition method includes:

[0046] S1. Perform the first decoding path search for real-time speech recognition on the audio stream containing multiple speech frames, and use the speech frame currently undergoing speech recognition as the target frame;

[0047] S2. If a pause is recognized after the target frame, split the audio stream at the pause to obtain the target audio segment;

[0048] S3. Perform the second decoding path search based on the target audio segment to obtain the real-time speech recognition result of the target frame.

[0049] In this embodiment, the above speech recognition method will be described in detail by taking the streaming speech recognition of the telephone content of the OTA call center as an example.

[0050] As Figure 3 shown, in the prior art, the workflow of streaming real-time speech recognition mainly includes feature extraction, acoustic score calculation, decoding path search, and obtaining the decoding result.

[0051] For step S1, in the streaming speech recognition algorithm, as the call progresses, the speech recognition algorithm will continuously generate intermediate decoding results. However, the speech recognition method in this embodiment does not directly output the decoding result corresponding to the target frame being recognized at the current moment. Instead, while performing the first decoding path search on the target frame, it continuously performs streaming speech recognition on the audio stream, so that after obtaining sufficient context information, it can perform a second decoding path search on the target frame again, thereby improving the accuracy of the finally output recognition result.

[0052] For step S2, when a pause is detected during the recognition process, it can be considered that the speaker has completed a complete expression of meaning (for example, the speaker has finished speaking a sentence). Therefore, at this time, the target audio segment obtained by splitting with this pause contains complete context information, and the recognition result output based on this target audio segment often has a high confidence level.

[0053] In this embodiment, in order to support the timely return of highly reliable transcription text during a long call, a streaming noise detection algorithm is used to detect the target audio stream, and splitting is performed when the end point of the pause is detected, so that the obtained target audio segment contains at least one complete expression of meaning. In the prior art, common streaming noise detection algorithms include, but are not limited to, the judgment method based on energy threshold and the detection method based on neural network.

[0054] For step S3, using the complete context information in the target audio segment can reduce the loss of recognition information, thereby improving the accuracy of the tail of the recognition result, so that the transcription of the entire call audio can be more natural and fluent.

[0055] In addition, since the pause time during the user's speech is usually limited, the time interval between two decoding path searches in this embodiment can be basically ignored, and the recognition result can still be output in real time to support the timely response of the customer service.

[0056] Since the telephone speech recognition in the OTA scenario basically belongs to conversational speech recognition, there are two independent channels for the OTA customer service and the customer. Their conversations alternate and last for a long time. Therefore, as a preferred embodiment, as Figure 4 shown, before step S1, it further includes:

[0057] S101. Obtain an initial audio stream;

[0058] S102. If the initial audio stream includes at least two channels, perform channel separation on the initial audio stream to obtain a target audio stream.

[0059] Since in this embodiment, the object for which speech recognition is required is the telephone content of the OTA call center, it has the characteristics of a large number of simultaneous calls and long call durations within the same time period. Therefore, the amount of real-time audio data generated is extremely large.

[0060] Based on this, as Figure 5 shown, in this embodiment, a streaming media is used to store the call audio stream in real time. After the call starts, as the call progresses, the call center generates a series of audio data. Then, the streaming media compresses the above audio data and sends the data in the form of network packets, so that the audio data can be sent as a target audio stream like running water or stored in the streaming media module for downstream services to subscribe to.

[0061] Specifically, when receiving a call start signal, the ASR algorithm engine starts to circularly pull the initial audio stream from the streaming media side. After obtaining the initial audio stream, it separates the two-channel audio of the customer service and the customer and uses two threads to process them separately to reduce the mutual interference of the two-channel audio in streaming speech recognition, improve the coherence between the contents of each channel, and thus improve the efficiency of streaming speech recognition and the accuracy of the final result.

[0062] Moreover, the use of streaming media can also decouple the ASR (automatic speech recognition) algorithm at the back end from other front-end modules, so as to more flexibly support different ASR algorithm engines to work in parallel and further improve the efficiency of speech recognition.

[0063] As Figure 6 shown, in order to increase the throughput during processing, in this embodiment, the decoder module of the ASR algorithm engine obtains the target audio stream through the mode of a shared queue.

[0064] Among them, the call audio stored in the streaming media is written into the shared queue in the form of an audio stream, and the decoder module, as a resident thread, continuously obtains the target audio stream from the shared queue through circulation for decoding, and broadcasts the decoding result with the highest score through the message queue for downstream services to subscribe to and use.

[0065] Preferably, to support the scenario requirements of high concurrency and low latency in a call, in this embodiment, a combination of CPU threads and GPU threads is used for further optimization. After the CPU thread extracts the features of the target audio stream, the task is encapsulated into a task queue, and the GPU thread suitable for intensive computing is used to calculate the acoustic score and search for the decoding path, and then the decoding result is passed back to the CPU thread for combination and broadcasting.

[0066] Preferably, the above-mentioned target audio segment includes ID information;

[0067] The speech recognition method further includes:

[0068] Performing speech recognition on several target audio segments with the same ID information to update the real-time speech recognition result.

[0069] During the streaming decoding process, since it is necessary to consider the context information within the target audio segment, in this embodiment, a unique ID is set for each target audio stream, and the target audio segments belonging to the same ID can be used as context information for each other to provide to the decoder to further improve the decoding accuracy.

[0070] In some cases, there may be too few audio frames after the target frame, resulting in insufficient subsequent context information. Therefore, in a preferred embodiment, before step S3, the above method further includes:

[0071] If the number of audio frames after the target frame in the target audio segment is less than the preset value, the last frame of the target audio segment is copied multiple times until the number of audio frames after the target frame in the target audio segment is greater than or equal to the preset value.

[0072] For example, the above preset value can be 1. If the number of audio frames after the target frame in the target audio segment is 0, at this time, in order for the speech recognition method in this embodiment to execute smoothly, the last frame of the target audio segment can be copied, and several pause frames are used as subsequent context information to indicate that there is no corresponding subsequent context information for the target frame.

[0073] It should be noted that the preset value of 1 here is only for illustrative purposes and is not limited thereto. The preset can also be set to any positive integer according to actual needs.

[0074] In one embodiment, since the customer service conversations in the OTA industry have strong domain relevance, therefore, in this embodiment, different ASR algorithm engines are adapted for customer service in different domains, and the ASR algorithm engine includes customized training of vocabulary and corpus specific to the domain scenario to improve the recognition accuracy of call audio in different domains.

[0075] In one embodiment, it further includes transcoding the audio stream into the pcm mode that can be recognized by the streaming speech recognition model, and then transmitting the transcoded audio stream to the decoder module of the speech recognition algorithm for streaming decoding, so that the audio stream can be adapted to the speech recognition algorithm model.

[0076] In the process of real-time speech recognition of the speech recognition method in this embodiment, the currently recognized audio frame is used as the target frame, and the recognition result of the target frame is not output temporarily. Instead, after reaching the end point of the audio segment, a secondary decoding path search is performed on the target frame based on the entire audio segment, and then the best recognition result is output, thereby realizing more accurate speech recognition based on the context information corresponding to the target frame, effectively overcoming the defect of low recognition accuracy at the end of the audio during real-time speech recognition, and improving the user experience.

[0077] Embodiment 2

[0078] This embodiment discloses a speech recognition system, as Figure 7 shown. The speech recognition system includes:

[0079] Decoder module 1, which is used for performing real-time speech recognition on the target audio stream, and using the speech frame that is currently being recognized as the target frame. The above real-time speech recognition includes the first decoding path search;

[0080] Audio segmentation module 2, which is used for segmenting the audio stream at the pause after the target frame is recognized to obtain the target audio segment;

[0081] Decoder module 1 is also used for performing a second decoding path search based on the target audio segment to obtain the real-time speech recognition result of the target frame.

[0082] This embodiment takes the streaming speech recognition of the phone content of the OTA call center as an example to elaborate on the above speech recognition system in detail

[0083] As Figure 3 shown, the working process of streaming real-time speech recognition in the prior art mainly includes feature extraction, acoustic score calculation, decoding path search, and obtaining the decoding result.

[0084] For decoder module 1, in the streaming speech recognition algorithm, as the call progresses, the speech recognition algorithm will continuously generate intermediate decoding results. However, decoder module 1 in this embodiment does not directly output the decoding result corresponding to the target frame that is currently being recognized, but will continuously perform streaming speech recognition on the audio stream while performing the first decoding path search on the target frame, so that the target frame can be decoded again after obtaining sufficient context information, thereby improving the accuracy of the finally output recognition result.

[0085] For the audio segmentation module 2, when a pause is detected during the recognition process, it can be considered that the speaker has completed a complete expression of meaning (for example, the speaker has finished a sentence). Therefore, the target audio segment obtained by segmenting at this pause contains complete context information, and the recognition result output based on this target audio segment often has a high confidence level.

[0086] In this embodiment, in order to support the timely return of highly reliable transcription text during a long call, the audio segmentation module 2 uses a streaming noise detection algorithm to detect the target audio stream and performs segmentation when the end point of the pause is detected, so that the obtained target audio segment contains at least one complete expression of meaning. In the prior art, common streaming noise detection algorithms include, but are not limited to, the judgment method based on energy threshold and the detection method based on neural network.

[0087] For the decoder module 1, using the complete context information in the target audio segment can reduce the loss of recognition information, thereby improving the accuracy of the tail of the recognition result, so that the transcription of the entire call audio can be more natural and fluent.

[0088] In addition, since the pause time during the user's speech is usually limited, the time interval between two decoding path searches in this embodiment can be basically ignored, and the recognition result can still be output in real time to support the timely response of the customer service.

[0089] Since the telephone speech recognition in the OTA scenario is basically a conversational speech recognition, there are two independent channels for the OTA customer service and the customer. Their conversations alternate and last for a long time. Therefore, as a preferred embodiment, the system in this embodiment further includes:

[0090] A streaming media module 3 for obtaining an initial audio stream;

[0091] A channel separation module 4 for separating channels of the initial audio stream to obtain a target audio stream when the initial audio stream includes at least two channels;

[0092] And / or, the above real-time speech recognition further includes feature extraction, acoustic score calculation, and obtaining a decoding result. As Figure 8 shown, the decoder module 1 includes a CPU thread and a GPU thread. The CPU thread is used for feature extraction and obtaining a decoding result, and the GPU thread is used for acoustic score calculation, the first decoding path search, and the second decoding path search. Data is transmitted between the GPU thread and the CPU thread in the form of a shared queue.

[0093] Since in this embodiment, the object for which speech recognition is required is the telephone content of the OTA call center, it has the characteristics of a large number of simultaneous calls and long call durations within the same time period. Therefore, the amount of real-time audio data generated therein is extremely large.

[0094] Based on this, in this embodiment, the streaming media module 2 is used to store the call audio stream in real time. After the call starts, as the call progresses, the call center generates a series of audio data. Then, the streaming media module 2 compresses the above audio data and sends the data in the form of network packets, so that the audio data can be sent or stored in the streaming media module 2 as a target audio stream like running water for downstream services to subscribe to.

[0095] Specifically, when receiving the signal that the call starts, the ASR algorithm engine starts to circularly pull the initial audio stream from the streaming media end. After obtaining the initial audio stream, the channel separation module 4 separates the two-channel audio of the customer service and the customer, and two threads are used in the decoder module 1 to process them separately to reduce the mutual interference of the two-channel audio in streaming speech recognition, improve the coherence between the contents of each channel, and thus improve the efficiency of streaming speech recognition and the accuracy of the final result.

[0096] Moreover, the use of the streaming media module 3 can also decouple the ASR (automatic speech recognition) algorithm at the backend from other modules at the frontend, so as to more flexibly support different ASR algorithm engines to work in parallel, further improving the efficiency of speech recognition.

[0097] As Figure 6 shown, in order to increase the throughput during processing, in this embodiment, the decoder module 1 of the ASR algorithm engine obtains the target audio stream through the mode of a shared queue.

[0098] Among them, the call audio stored in the streaming media module 3 is written into the shared queue in the form of an audio stream, and the decoder module 1, as a resident thread, continuously obtains the target audio stream from the shared queue through circulation for decoding, and broadcasts the decoding result with the highest score through the message queue for downstream services to subscribe to and use.

[0099] Preferably, in order to support the scenario requirements of high concurrency and low latency during calls, the decoder module 1 in this embodiment is further optimized by combining CPU threads and GPU threads. After the CPU thread extracts the features of the target audio stream, it encapsulates the task into a task queue, and the GPU thread suitable for intensive computing is used to perform acoustic score calculation and decoding path search, and then the decoding result is passed back to the CPU thread for combination and broadcast.

[0100] Preferably, the above target audio segment includes ID information;

[0101] The decoder module 1 is further configured to: perform speech recognition on a plurality of target audio segments with the same ID information to update the real-time speech recognition result.

[0102] During the streaming decoding process, since the context information within the target audio segment needs to be considered, in this embodiment, a unique ID is set for each target audio stream, and the target audio segments belonging to the same ID can be used as context information for each other and provided to the decoder module 1 to further improve the decoding accuracy.

[0103] In some cases, there may be too few audio frames after the target frame, so that insufficient subsequent information cannot be obtained. Therefore, in a preferred embodiment, the decoder module 1 is further configured to:

[0104] When the number of audio frames after the target frame in the target audio segment is less than a preset value, the last frame of the target audio segment is copied multiple times until the number of audio frames after the target frame in the target audio segment is greater than or equal to the preset value.

[0105] For example, the above preset value can be 1. If the number of audio frames after the target frame in the target audio segment is 0, at this time, in order for the speech recognition method in this embodiment to be successfully executed, the last frame of the target audio segment can be copied, and several pause frames are used as subsequent information to indicate that there is no corresponding subsequent information for the target frame.

[0106] It should be noted that the preset value of 1 here is only for illustrative purposes and is not limited thereto. The preset can also be set to any integer greater than zero according to actual needs.

[0107] In one embodiment, since the customer service conversations in the OTA industry have strong domain relevance, in this embodiment, different ASR algorithm engines are adapted for customer service in different domains. The ASR algorithm engine includes customized training of vocabulary and corpus specific to the domain scenario to improve the recognition accuracy of call audio in different domains.

[0108] In one embodiment, it further includes transcoding the audio stream into the pcm mode that can be recognized by the streaming speech recognition model, and then transmitting the transcoded audio stream to the decoder module of the speech recognition algorithm for streaming decoding, so that the audio stream can be adapted to the speech recognition algorithm model.

[0109] During the process of real-time speech recognition, the speech recognition method in this embodiment takes the currently recognized audio frame as the target frame and does not output the recognition result of this target frame temporarily. Instead, after reaching the end point of the audio segment, it performs a secondary decoding path search for this target frame based on the entire audio segment, and then outputs the best recognition result, thereby achieving more accurate speech recognition based on the context information corresponding to this target frame, effectively overcoming the defect that the recognition accuracy of the tail of the audio during real-time speech recognition is relatively low, and improving the user experience.

[0110] Embodiment 3

[0111] Figure 9 FIG. 7 is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the speech recognition method in Embodiment 1. Figure 9 The displayed electronic device 30 is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0112] As Figure 9 shown, the electronic device 30 may be presented in the form of a general-purpose computing device. For example, it may be a smart watch. The components of the electronic device 30 may include, but are not limited to: the at least one processor 31 described above, the at least one memory 32 described above, and a bus 33 connecting different system components (including the memory 32 and the processor 31).

[0113] The bus 33 includes a data bus, an address bus, and a control bus.

[0114] The memory 32 may include volatile memory, such as random access memory (RAM) 321 and / or cache memory 322, and may further include read-only memory (ROM) 323.

[0115] The memory 32 may further include a program / utilities 325 having a set (at least one) of program modules 324. Such program modules 324 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0116] The processor 31 executes various functional applications and data processing by running the computer program stored in the memory 32, such as the speech recognition method in Embodiment 1 of the present invention.

[0117] The electronic device 30 can also communicate with one or more external devices 34 (such as a mobile phone). Such communication can be carried out through the input / output (I / O) interface 35. Moreover, the model generation device 30 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 36. As Figure 9 shown, the network adapter 36 communicates with other modules of the model generation device 30 through the bus 33. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the model generation device 30, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (redundant array of independent disks) systems, tape drives, and data backup storage systems, etc.

[0118] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present invention, the features and functions of the two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0119] Embodiment 4

[0120] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the voice recognition method in Embodiment 1.

[0121] Among them, the more specific computer-readable storage medium that can be adopted can include but not limited to: portable disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories, optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0122] In a possible implementation manner, the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to enable the terminal device to execute the voice recognition method in Embodiment 1.

[0123] Among them, the program code for executing the present invention can be written in any combination of one or more programming languages. The program code can be completely executed on the user device, partially executed on the user device, executed as an independent software package, partially executed on the user device and partially executed on a remote device, or completely executed on a remote device.

[0124] Although the specific embodiments of the present invention have been described above, those skilled in the art should understand that this is only an example, and the protection scope of the present invention is defined by the appended claims. Without departing from the principles and essence of the present invention, those skilled in the art can make various changes or modifications to these embodiments, but these changes and modifications all fall within the protection scope of the present invention.

Claims

1. A voice recognition method, characterized in that, The speech recognition method includes: Performing a first decoding path search for real-time speech recognition on an audio stream containing multiple frames of speech frames, and using the speech frame currently undergoing speech recognition as the target frame; If a pause is recognized after the target frame, splitting the audio stream at the pause to obtain a target audio segment; Performing a second decoding path search based on the target audio segment to obtain the real-time speech recognition result for the target frame, where the target audio segment includes complete context information.

2. The speech recognition method according to claim 1, characterized in that, Before the step of performing the first decoding path search for real-time speech recognition on an audio stream containing multiple frames of speech frames, it further includes: Obtaining an initial audio stream; If the initial audio stream includes at least two channels, performing channel separation on the initial audio stream to obtain the target audio stream.

3. The voice recognition method according to claim 1, characterized in that Before the step of performing the second decoding path search based on the target audio segment, it further includes: If the number of audio frames after the target frame in the target audio segment is less than a preset value, copying the last frame of the target audio segment multiple times until the number of audio frames after the target frame in the target audio segment is greater than or equal to the preset value.

4. The speech recognition method according to claim 1, wherein The target audio segment includes ID information; The speech recognition method further includes: Performing speech recognition on several target audio segments with the same ID information to update the real-time speech recognition result.

5. A voice recognition system, characterized in that, The speech recognition system includes: A decoder module for performing real-time speech recognition on a target audio stream and using the speech frame currently undergoing speech recognition as the target frame, where the real-time speech recognition includes a first decoding path search; An audio splitting module for splitting the audio stream at the pause when a pause after the target frame is recognized to obtain a target audio segment; The decoder module is further used to perform a second decoding path search based on the target audio segment to obtain the real-time speech recognition result for the target frame, where the target audio segment includes complete context information.

6. The voice recognition system according to claim 5, wherein The system further includes: A streaming media module for obtaining an initial audio stream; A channel separation module for performing channel separation on the initial audio stream to obtain the target audio stream when the initial audio stream includes at least two channels; And / or, the real-time speech recognition further includes feature extraction, acoustic score calculation, and obtaining a decoding result. The decoder module includes a CPU thread and a GPU thread. The CPU thread is used for performing the feature extraction and obtaining the decoding result, and the GPU thread is used for performing the acoustic score calculation, the first decoding path search, and the second decoding path search. Data is transmitted between the GPU thread and the CPU thread in a shared queue manner.

7. The voice recognition system according to claim 5, characterized in that, The decoder module is further used for: When the number of audio frames after the target frame in the target audio segment is less than a preset value, copying the last frame of the target audio segment multiple times until the number of audio frames after the target frame in the target audio segment is greater than or equal to the preset value.

8. The voice recognition system according to claim 5, wherein, The target audio segment includes ID information; The decoder module is further used for: performing speech recognition on several target audio segments with the same ID information to update the real-time speech recognition result.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the speech recognition method described in any one of claims 1-4.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the speech recognition method described in any one of claims 1-4.

Citation Information

Patent Citations

  • Method and device for decoding voice data

    CN102436816A

  • Voice recognition method and device, computer equipment and computer readable storage medium

    CN112750425A