Voice processing method and device, electronic equipment and program product

By using historical decoding results and the longest common prefix import verification mechanism in the speech buffer, the streaming speech recognition method of the Whisper model is improved, which solves the problems of speech recognition delay and low accuracy, and achieves more efficient and accurate speech recognition.

CN120690207APending Publication Date: 2025-09-23LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511074600.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The Whisper model has problems with large speech recognition delay and recognition rate loss in streaming speech recognition, especially speech truncation at the boundaries of speech segments, which leads to low recognition efficiency and accuracy.

Method used

By utilizing historical decoding results in the speech buffer as audio and text context data, the decoding method of the Whisper model is improved, and the longest common prefix import verification mechanism and sentence segmentation algorithm within the sentence are adopted to improve decoding accuracy and consistency.

Benefits of technology

It improves the efficiency and accuracy of streaming speech recognition, reduces computational load, enhances the model's ability to understand and robustness of speech data, and reduces repeated decoding and delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690207A_ABST
    Figure CN120690207A_ABST
Patent Text Reader

Abstract

The invention discloses a voice processing method and device, electronic equipment and a program product, and the method comprises the steps: reading first voice data and second voice data from a voice buffer area; wherein the first voice data represents at least part of voice data in the current recognized sentence; the second voice data represents decoded voice data before the current recognized sentence; decoding the coded data of the first voice data according to first prompt information and the coded data of the second voice data to obtain a first decoding result corresponding to the first voice data; the first prompt information at least comprises a second decoding result corresponding to the second voice data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to, but is not limited to, the field of computer technology, and in particular to a speech processing method, device, electronic device, and program product. Background Art

[0002] Speech recognition converts human speech into digital information that machines can understand and process, enabling functions such as voice input, voice control, and speech transcription. Common speech recognition models include the Hidden Markov Model (HMM), Deep Neural Network (DNN), Recurrent Neural Network (RNN), and the Whisper model. The Whisper model boasts leading recognition accuracy and can process speech in 97 languages.

[0003] Streaming speech recognition, which involves recording a speech stream while simultaneously recognizing and outputting partial or final results, is a key application scenario in the field of speech recognition. However, the Whisper model recognizes speech in segments—for example, a 30-second segment. This results in significant speech recognition latency, and speech truncation at segment boundaries often results in a loss in recognition rate. Therefore, improving the efficiency and accuracy of streaming speech recognition has become a pressing issue. Summary of the Invention

[0004] In view of this, the present application at least provides a speech processing method, device, electronic device and program product.

[0005] The technical solution of this application is achieved as follows:

[0006] In one aspect, the present application provides a speech processing method, the method comprising:

[0007] Reading first voice data and second voice data from the voice buffer; wherein the first voice data represents at least part of the voice data in the current recognition sentence; and the second voice data represents the decoded voice data before the current recognition sentence;

[0008] According to the first prompt information and the encoded data of the second voice data, the encoded data of the first voice data is decoded to obtain a first decoding result corresponding to the first voice data; the first prompt information at least includes a second decoding result corresponding to the second voice data.

[0009] In some implementations, after obtaining the first decoding result corresponding to the first voice data, the method further includes:

[0010] Using the first decoding result to update the first prompt information, to obtain the second prompt information;

[0011] Reading third voice data from the voice buffer; the third voice data represents the voice data buffered after the first voice data in the currently recognized sentence;

[0012] The encoded data of the third voice data is decoded according to the second prompt information and the encoded data of the second voice data to obtain a third decoding result.

[0013] In some embodiments, using the first decoding result to update the first prompt information to obtain the second prompt information includes: using a first partial result of the first decoding result to update the first prompt information to obtain the second prompt information; the first decoding result includes the first partial result decoded first and the second partial result decoded later;

[0014] According to the second prompt information and the encoded data of the second voice data, the third voice data is decoded to obtain a third decoding result, including: according to the second prompt information and the encoded data of the second voice data, the encoded data corresponding to the second part of the result and the encoded data corresponding to the third voice data are decoded to obtain a fourth decoding result and a fifth decoding result respectively; if the fourth decoding result matches the second part of the result, the fifth decoding result is used as the third decoding result.

[0015] In some embodiments, the method further comprises:

[0016] In response to the fourth decoding result not matching the second partial result, deleting the decoding result of the current recognized sentence from the first prompt information to obtain third prompt information;

[0017] Re-decode from the starting position of the current recognition sentence according to the third prompt information and the second voice data.

[0018] In some embodiments, the first decoding result further includes a third partial result; the third partial result represents a decoding result generated after the second partial result;

[0019] The method further includes: taking the text data corresponding to the first part of the result and the second part of the result as output results.

[0020] In some embodiments, the method further comprises:

[0021] Using the first decoding result to update the first prompt information, to obtain the second prompt information;

[0022] In response to the voice decoding result in the second prompt information being the same as the voice decoding result in the first prompt information, removing the voice data of the first duration at the head of the voice buffer, and adding the voice data to be recognized of the first duration to the tail of the voice buffer;

[0023] The first prompt information is used to decode the voice data in the voice buffer.

[0024] In some embodiments, the method further comprises:

[0025] Perform sentence segmentation processing on the speech data to be recognized to obtain at least one sentence;

[0026] Store at least one sentence in a speech buffer.

[0027] In another aspect, the present application provides a speech processing device, comprising:

[0028] A speech reading module is configured to read first speech data and second speech data from a speech buffer; wherein the first speech data represents at least part of speech data in a currently recognized sentence; and the second speech data represents decoded speech data before the currently recognized sentence;

[0029] The decoding module is used to decode the first voice data according to the first prompt information and the second voice data to obtain a first decoding result corresponding to the first voice data; the first prompt information at least includes a second decoding result corresponding to the second voice data.

[0030] On the other hand, the present application also provides an electronic device, comprising at least one processor; a speech recognition model is running on at least one processor; wherein,

[0031] A speech recognition model is used to read first speech data and second speech data from a speech buffer; wherein the first speech data represents at least part of the speech data in the currently recognized sentence; the second speech data represents the decoded speech data before the currently recognized sentence; according to the first prompt information and the encoded data of the second speech data, the encoded data of the first speech data is decoded to obtain a first decoding result corresponding to the first speech data; the first prompt information at least includes a second decoding result corresponding to the second speech data.

[0032] On the other hand, the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, implements some or all of the steps in the above method.

[0033] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the technical solutions of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.

[0035] Figure 1 A schematic diagram of the implementation flow of a speech processing method provided in this application;

[0036] Figure 2 A schematic diagram of the implementation flow of a speech processing method provided in this application;

[0037] Figure 3 A schematic diagram of the implementation flow of a speech processing method provided in this application;

[0038] Figure 4 A schematic diagram of an implementation flow of an embodiment provided in this application;

[0039] Figure 5 A schematic diagram of an implementation flow of another embodiment provided in this application;

[0040] Figure 6 A schematic diagram of a decoding process according to another embodiment of the present invention;

[0041] Figure 7 A schematic diagram of the structure of a speech processing device provided in this application;

[0042] Figure 8 A schematic diagram of the hardware entity of an electronic device provided in this application. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions of this application are further elaborated in detail below with reference to the accompanying drawings and examples. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0044] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0045] The terms "first / second / third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first / second / third" can be interchanged with a specific order or sequence where permitted so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing this application only and are not intended to limit this application.

[0047] In response to the problem that the Whisper model is not suitable for real-time streaming speech recognition, the relevant technology improves the Whisper model by introducing algorithms such as Local Agreement to achieve real-time streaming speech recognition. However, this method still has the following obvious limitations: first, the decoding result of the previous sentence of the current sentence is used as prompt information to guide the decoding of the current sentence, and before decoding the period, the speech in the speech buffer needs to be decoded from the beginning each time, which leads to repeated decoding and high decoding delay. Second, the sentence segmentation relies on the punctuation marks output by the model, and the output of punctuation marks by the model is a probabilistic event, so it is easy to fail to output the period for a long time, resulting in delayed output or loss of the next sentence result.

[0048] Based on this, the present application provides a speech processing method, which can be performed by an electronic device, which can be various types of terminals such as laptops, tablet computers, desktop computers, set-top boxes, mobile devices (for example, mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), etc., and can also be implemented as a server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0049] Below, the technical solution of this application will be clearly and completely described in conjunction with the drawings in this application.

[0050] Figure 1 A flow chart of the implementation of a speech processing method provided in this application is shown as follows: Figure 1 As shown, the method includes the following steps S11 to S12:

[0051] Step S11, reading first voice data and second voice data from a voice buffer; wherein the first voice data represents at least part of the voice data in the current recognition sentence; and the second voice data represents decoded voice data before the current recognition sentence.

[0052] Here, the speech buffer refers to a temporary storage area for speech data to be recognized. For example, in streaming speech recognition, the speech to be recognized is input segment by segment, and the speech buffer is used to accumulate these speech segments for overall or step-by-step decoding.

[0053] In some embodiments, the voice buffer can be implemented as any type of storage area in the electronic device. For example, the voice buffer can be implemented as a memory area, specifically, the voice buffer can be implemented as a ring buffer, a linear buffer, etc.

[0054] The first voice data is at least part of the voice data in the currently recognized sentence, for example, the first to P-th characters, or the P-th to Q-th characters in the currently recognized sentence; wherein P and Q are both integers greater than 1.

[0055] The second voice data refers to the decoded voice data between the currently recognized sentences, for example, the N sentences of voice data preceding the currently recognized sentence; where N is greater than 0, i.e., the second voice data may include at least one complete sentence or a portion of a sentence. The decoding result of the second voice data is the second decoding result. In some embodiments, the second decoding result may be stored in any data buffer.

[0056] Step S12: Decode the encoded data of the first voice data according to the first prompt information and the encoded data of the second voice data to obtain a first decoding result corresponding to the first voice data; the first prompt information at least includes a second decoding result corresponding to the second voice data.

[0057] Here, the first prompt information is prompt information used to guide the model to decode the first voice data.

[0058] The first prompt information at least includes a second decoding result corresponding to the second voice data, so that the second decoding result is used as text context data for decoding the first voice data.

[0059] After reading the first and second voice data from the voice buffer, the first and second voice data are respectively encoded using an encoder to obtain encoded data. The encoded data of the second voice data used to generate the second decoding result is different from the encoded data of the second voice data. When decoding the first voice data, the second voice data read from the voice buffer is re-encoded to obtain the encoded data of the second voice data.

[0060] In this way, the encoded data of the first voice data is decoded according to the first prompt information and the encoded data of the second voice data, and the second voice data and its historical decoding results can be used together as context data for decoding the first voice data for model reference to obtain a first decoding result that is consistent with the previous decoding result.

[0061] In the speech processing method provided by the present application, the decoded second speech data will not be deleted from the speech buffer immediately after the decoding is completed, but will be re-input into the encoder for encoding, and the encoding result and the historical decoding result of the second speech data (i.e., the second decoding result) will be used as the audio context and text context data of the first speech data decoded later. In this way, compared with the related art that only provides text context data to the model, the solution provided by the present application provides the model with audio context and text context data at the same time when using the model to decode the first speech data, that is, providing the model with longer and different types of context data, thereby enhancing the model's ability to understand the first speech data, thereby improving the accuracy, consistency and stability of the decoding results.

[0062] In some embodiments, as Figure 2 As shown, after obtaining the first decoding result corresponding to the first voice data, that is, after the above step S12, the method further includes the following steps S21 to S23:

[0063] Step S21: Use the first decoding result to update the first prompt information to obtain second prompt information.

[0064] Here, after obtaining the first decoding result, the first decoding result is updated to the first prompt information to obtain the second prompt information, so that the second prompt information is used to continue decoding the voice data in the current recognition sentence. It can be seen that the second prompt information includes at least the first decoding result corresponding to the first voice data and the second decoding result corresponding to the second voice data.

[0065] Step S22 , reading third voice data from the voice buffer; the third voice data represents the voice data in the currently recognized sentence that is buffered after the first voice data.

[0066] Here, the third voice data refers to the voice data cached after the first voice data in the currently recognized sentence. In some embodiments, the third voice data can be voice data of the same or different length as the first voice data.

[0067] In some embodiments, after the first voice data is decoded, the third voice data can be read from the voice buffer; the current recognized sentence can also be read from the voice buffer at one time, that is, the first voice data and the second voice data are read from the voice buffer at the same time, and the current recognized sentence is sent to the encoder for encoding, and then the decoder is used to decode the encoded data of the first voice data and the encoded data of the second voice data in turn.

[0068] Step S23: Decode the encoded data of the third voice data according to the second prompt information and the encoded data of the second voice data to obtain a third decoding result.

[0069] Here, based on the first decoding result of the first voice data, the second decoding result of the second voice data, and the encoded data of the second voice data, the encoded data of the third voice data is decoded to obtain a third decoding result of the third voice data.

[0070] Compared to the problem of repeated decoding caused by using the decoding result of the previous sentence of the current recognition sentence as prompt information to guide the decoding of the current recognition sentence in the related art, the above embodiment provided by the present application uses sentences as units for cyclic decoding. For the current recognition sentence, when decoding the third voice data cached later, the historical decoding result in the current recognition sentence (i.e., the decoding result of the first voice data) is used as the longest common prefix to construct prompt information and guide the decoding of the third voice data. In this way, when decoding the current recognition sentence, the decoding granularity of the model is refined, and the model does not need to start decoding from the first word of the sentence each time, thereby improving the decoding speed of the model, reducing the load on the computing chip, and improving the streaming recognition effect.

[0071] In some embodiments, after determining the first decoding result corresponding to the first voice data, in order to improve the accuracy of the model output result, the first decoding result is verified, and the first decoding result is output after the first decoding result passes the verification, and the second prompt information updated based on the first decoding result is determined as valid prompt information.

[0072] In order to verify the accuracy of the first decoding result, the present application proposes an intra-sentence longest common prefix import verification mechanism, that is, within the scope of the currently recognized sentence, the historical decoding results of the sentence are used as the longest common prefix, and then part of the longest common prefix is ​​used to construct prompt information, and the other part of the longest common prefix is ​​used as verification information.

[0073] Thus, in some implementations, the first prompt information is updated using the first decoding result to obtain the second prompt information, that is, the above step S21, can be implemented as the following step S211:

[0074] Step S211: using a first partial result in the first decoding result to update the first prompt information, to obtain the second prompt information; the first decoding result includes a first partial result decoded first and a second partial result decoded later.

[0075] Here, the first decoding result is split into a first partial result decoded first and a second partial result decoded later, and the first partial result is updated to the first prompt information (ie, the prompt information is constructed using the first partial result) to obtain the second prompt information.

[0076] In some embodiments, the lengths of the first partial result and the second partial result can be determined in any manner. For example, the first decoding result can be split into the first partial result and the second partial result according to a preset length ratio based on the total length of the first decoding result. For another example, the first R characters or word units in the first decoding result can be used as the first partial result, and the remaining characters or word units can be used as the second partial result, where R is an integer greater than 0.

[0077] For example, when the first decoding result is "The weather is good today, everyone is planning", "The weather is not good today" can be used as the first partial result and "It's wrong, everyone is planning" can be used as the second partial result.

[0078] In this way, the third voice data is decoded according to the second prompt information and the encoded data of the second voice data to obtain a third decoding result. That is, the above step S23 can be implemented as the following step S231:

[0079] Step S231: Decode the encoded data corresponding to the second partial result and the encoded data corresponding to the third voice data according to the second prompt information and the encoded data of the second voice data to obtain a fourth decoding result and a fifth decoding result, respectively; if the fourth decoding result matches the second partial result, use the fifth decoding result as the third decoding result.

[0080] Here, according to the second prompt information and the second voice data, the encoded data corresponding to the second partial result is re-decoded to obtain a fourth decoding result; and the encoded data corresponding to the third voice data is decoded to obtain a fifth decoding result.

[0081] Thus, when the fourth decoding result matches the second partial result, it indicates that the decoding result for the first voice data is accurate, and the second prompt information is reliable. Therefore, the first decoding result is used as the output result corresponding to the first voice data, and the fifth decoding result is used as the third decoding result corresponding to the third voice data.

[0082] For example, in the above example, if the second decoding result (i.e., the fourth decoding result) of the encoded data corresponding to the second part result "wrong everyone's plan" is still "wrong everyone's plan", it is confirmed that the first decoding result "nice weather today everyone's plan" is accurate.

[0083] Similarly, the same method as steps S211 and S231 above can be used to verify the third decoding result.

[0084] In the above embodiments provided by the present application, based on the in-sentence longest common prefix import verification mechanism, the historical decoding result in the current recognized sentence is used as the longest common prefix, and a part of the longest common prefix is used to construct a prompt message, and another part of the longest common prefix is used as verification information to implement the accuracy verification of the first decoding result, thereby improving the accuracy of the model decoding result.

[0085] In some embodiments, the decoding result can be verified by importing the longest common prefix within the sentence according to a preset rule. For example, every time the decoder decodes N word tokens, the longest common prefix import verification is performed once; or for another example, every time the decoder performs M decoding steps, the longest common prefix import verification is performed once; and so on.

[0086] In some embodiments, the method further includes the following steps S24 to S25:

[0087] Step S24, in response to the fourth decoding result not matching the second part result, delete the decoding result of the current recognized sentence from the first prompt message to obtain a third prompt message.

[0088] Here, the fourth decoding result not matching the second part result means that the similarity between the fourth decoding result and the second part result is less than the similarity threshold. For example, in the above example, if the fourth decoding result is "wrong adult's plan", which is different from the second part result "wrong everyone's plan", it is considered that the fourth decoding result does not match the second part result.

[0089] When the fourth decoding result does not match the second part result, it means that the first decoding result is inaccurate. Therefore, in order to improve the decoding accuracy of the current recognized sentence, all the decoding results of the current recognized sentence included in the first prompt message are deleted to obtain a third prompt message.

[0090] Step S25, according to the third prompt message and the second voice data, re-decode from the starting position of the current recognized sentence.

[0091] Here, re-decoding the current recognized sentence according to the third prompt message and the second voice data means starting to re-decode from the starting position of the current recognized sentence.

[0092] In the above embodiments provided by the present application, when the first decoding result fails to pass the verification, the current recognized sentence is re-decoded using the third hint information, which can improve the decoding accuracy of the current recognized sentence and avoid the influence of the incorrect first decoding result on the subsequent decoding process.

[0093] In some embodiments, the first decoding result further includes a third partial result; the third partial result represents the decoding result generated after the second partial result;

[0094] The method further includes: using the text data corresponding to the first partial result and the second partial result as the output result.

[0095] Here, according to the decoding sequence, the first decoding result is divided into three partial results, namely the first partial result, the second partial result, and the third partial result. For example, in the above example, the first decoding result "The weather is nice today everyone plan" is divided into three partial results, namely "The weather is not", "wrong everyone", and "plan" in sequence.

[0096] Since the third partial result is the tangent point of the first speech data and the third speech data and has a relatively low decoding accuracy, when verifying the first decoding result, the third partial result is excluded, and only the second partial result is used as the verification information. For example, in the above example, "wrong everyone" is used to verify the first decoding result.

[0097] In this way, when it is determined that the first decoding result passes the verification, the text data corresponding to the first partial result and the second partial result is used as the output result to improve the accuracy of the output result.

[0098] In some embodiments, when the length of the current recognized sentence is greater than the maximum length of the speech segment (chunk) that the model can process, the speech at the head of the sentence (i.e., the decoded speech at the head of the sentence in the speech buffer) is removed to meet the input requirements of the speech segment. The length of the speech segment that the model can process can be any preset length, for example, 10 seconds, 20 seconds, 30 seconds, etc.

[0099] Here, when removing the speech at the head of the sentence, when calculating the longest common prefix of the current recognized sentence, the already determined longest common prefix and the latest decoded decoding result are still used as the longest common prefix of the current recognized sentence, and the hint information is constructed using this longest common prefix.

[0100] In some embodiments, as Figure 3 shown, the method further includes the following steps S31 to step S33:

[0101] Step S31, updating the first hint information using the first decoding result to obtain a second hint information.

[0102] Here, after obtaining the first decoding result, the first decoding result is updated to the first prompt information to obtain the second prompt information, so as to continue to perform decoding on the speech data in the current recognition sentence using the second prompt information.

[0103] Step S32: In response to the voice decoding result in the second prompt information being the same as the voice decoding result in the first prompt information, remove the voice data of the first duration at the head of the voice buffer, and add the voice data to be recognized of the first duration to the tail of the voice buffer.

[0104] Here, the speech decoding result in the second prompt information is compared with the speech decoding result in the first prompt information to determine whether the two are the same, that is, to determine whether the longest common prefix in the second prompt information is the same as that in the first prompt information; if they are not the same, it is considered that the first decoding result is not empty, or the first speech data is speech containing semantic content; if they are the same, it is considered that the first decoding result is empty, and a speech content jump occurs at the first speech data.

[0105] When it is determined that a voice jump occurs, the content of the voice buffer is updated, that is, the voice data of the first duration at the head of the voice buffer is removed, and the voice data to be recognized of the first duration is added to the tail of the voice buffer.

[0106] In some implementations, the first duration may be any preset duration, for example, 1 second, 2 seconds, or other durations.

[0107] Step S33: Decoding the voice data in the voice buffer using the first prompt information.

[0108] Here, after the content of the voice buffer is updated, the voice buffered in the voice buffer is decoded using the first prompt information.

[0109] During implementation, if no decoding result is generated after decoding the updated voice buffer content using the first prompt information, or if the voice decoding result in the second prompt information obtained after updating the first prompt information using the generated decoding result is still the same as the voice decoding result in the first prompt information, then the above steps S32 and S33 are executed in a loop until a decoding result is generated and the voice decoding results in the first prompt information and the second prompt information are different, that is, the longest common prefix changes.

[0110] In the above embodiments provided in the present application, by detecting whether voice jumps occur in the currently decoded voice data, the robustness of the model in the voice processing process is improved.

[0111] In some embodiments, the method further comprises the following steps S13 to S14:

[0112] Step S13, performing sentence segmentation processing on the speech data to be recognized to obtain at least one sentence;

[0113] Step S14: storing the at least one sentence in the voice buffer.

[0114] Here, before the speech data to be recognized is cached in the speech buffer, sentence segmentation is performed on the speech data to be recognized to determine the starting point and the end point of the sentence in the speech data to be recognized, that is, to obtain at least one sentence.

[0115] In some embodiments, a specified speech detection algorithm may be used to segment the speech data to be recognized and generate a Begin of Sentence (BOS) and an End of Sentence (EOS) marker in the speech data to indicate the start and end of a sentence, respectively. For example, a Voice Activity Detection (VAD) and / or Speech Activity Detection (SAD) algorithm may be used to segment the speech data to be recognized.

[0116] In some implementations, while segmenting the speech data to be recognized, speech and non-speech detection is performed on the speech data to be recognized, and the non-speech portion is removed to reduce the encoding and decoding burden of the model. The non-speech portion refers to non-human speech or speech that does not contain semantic meaning.

[0117] In this way, after determining at least one sentence, the at least one sentence is stored in a speech buffer so as to process the speech data using a model.

[0118] Compared with the solution in the related art that relies on models to generate punctuation marks, in the above embodiments of the present application, before decoding the voice data, the voice data to be recognized is segmented using a specified segmentation algorithm, which can improve the accuracy of segmentation and thus improve the accuracy of the voice decoding results.

[0119] Next, combine Figure 4 , the speech processing process in an embodiment provided by the present application is described, wherein the embodiment is implemented by a speech processing system for implementing the speech processing method provided by the present application; the speech processing system is constructed based on the Wishper model and includes a sentence segmentation module 410, a feature extraction module 420 and a recognition task controller 430; wherein,

[0120] First, the speech data to be recognized 401 is input into the segmentation module 410, so as to generate speech pulse code modulation (PCM) data 402 by the segmentation module 410;

[0121] Here, the sentence segmentation module 410 can use the VAD algorithm and the SAD algorithm to detect the speech part and the non-speech part in the speech data to be recognized, thereby cutting off the non-speech part and generating a start mark and an end mark for the speech part in units of sentences to obtain output data in PCM format.

[0122] Then, the sentence segmentation module 410 sends the speech PCM data 402 to the feature extraction module 420, so that the feature extraction module 420 extracts audio features from the speech PCM data 402 to generate a log-Mel spectrogram 403.

[0123] Here, the feature extraction module 420 uses an open-source Fast Fourier Transform (FF) acceleration optimization algorithm to perform feature extraction on the speech PCM data 402 to obtain a logarithmic Mel spectrum 403 .

[0124] Afterwards, the feature extraction module 420 sends the logarithmic mel spectrum 403 to the inference module 431 in the recognition task controller 430 , so as to generate the speech recognition result 404 using the inference module 431 ;

[0125] Here, the inference module 431 includes a prompt information generation unit and an inference unit; wherein the prompt information generation unit is used to generate prompt information input to the inference unit; the inference unit includes an encoder and a decoder to encode and decode the speech data.

[0126] Afterwards, the inference module 431 sends the speech recognition result 404 to the result combination module 432 in the recognition task controller 430 , so that the result combination module 432 combines the multiple speech recognition results 404 to obtain the output text 405 .

[0127] Next, combine Figure 5 , the implementation process of an embodiment of the present application is described. Figure 5 As shown, this embodiment includes the following steps S51 to S57:

[0128] Step S51, reading the fifth voice data decoded before the current recognition sentence from the voice buffer; then executing step S52;

[0129] Step S52, encoding the fifth voice data using an encoder to obtain fifth encoded data; then, executing step S53;

[0130] Step S53, inputting the fourth prompt information and the fifth encoded data into a decoder, so as to decode the encoded data of the sixth voice data using the decoder to obtain a sixth decoding result;

[0131] Here, the fourth prompt information includes the fifth decoding result and the seventh decoding result corresponding to the fifth voice data; wherein the seventh decoding result represents the first part of the result, that is, the first part of the result in the longest common prefix determined according to the historical decoding results of the currently recognized sentence; the longest common prefix also includes the second part of the result.

[0132] Step S54, determining whether the sixth decoding result includes the second part of the longest common prefix; if so, executing step S55; if not, executing step S56;

[0133] Step S55, read new voice data from the voice buffer and continue decoding;

[0134] Step S56, deleting the speech decoding result of the current recognition sentence from the fourth prompt information to obtain the fifth prompt information; then executing step S57;

[0135] Step S57: re-decode the currently recognized sentence based on the fifth prompt information.

[0136] Next, combine Figure 6 , the process of decoding speech in one embodiment of the present application is described.

[0137] The data inside the block 610 represents the voice data buffered in the voice buffer;

[0138] The data marked in italics in box 610 represents the decoded speech data before the current recognition sentence;

[0139] The data marked with a horizontal line in the box 610 represents the data currently being fed into the encoder for encoding;

[0140] The data marked with a horizontal line outside the box 610 represents prompt information;

[0141] The data outside the box 610 and not marked with a horizontal line represents the speech recognition data that has been confirmed and output.

[0142] like Figure 6 As shown, in the first decoding 620:

[0143] The first line 621 shows that the voice data corresponding to the decoded content "Thank you, Mr. President" is retained in the voice buffer as voice context data, and the corresponding text data is used as the text context in the prompt information. At the same time, the partial decoding result "today," which has been decoded in the current recognized sentence, is used as the longest common prefix to construct the prompt information. The current recognized sentence is "Today, I am very grateful to Mr. Zhang for his report."

[0144] According to the prompt information shown in the first line 621, the current recognized sentence is decoded to determine a decoding result "I", and the decoding result is used as a new common prefix to update the prompt information. Thus, the updated prompt information is shown in the second line 622, including "Thank you, Mr. President. Today, I";

[0145] According to the prompt information shown in the second line 622, decoding is continued to obtain a new decoding result "Thank you very much, Mr. Zhang". In this way, the output obtained by the first decoding 620 is "Thank you, Mr. President. Today, I am very grateful to Mr. Zhang", as shown in the third line 623.

[0146] In the second decoding 630:

[0147] For the new decoding result "Thank you very much, Mr. Zhang" obtained from the first decoding 620, a portion of the result "Thank you very much" is used to construct a prompt message, and the other portion "Mr. Zhang" is used as verification information. In this way, the newly constructed prompt message is shown in the first line 631 as "Thank you, Mr. President. I am very grateful today."

[0148] Thus, based on the prompt information shown in the first line 631, the voice data is further decoded to obtain the output "Thank you, Mr. President. I am very grateful for Mr. Zhang's report today." Here, the information "Mr. Zhang" decoded in the second decoding 630 is consistent with the verification information "Mr. Zhang", so it can be determined that the decoding result of the first decoding 620 is accurate.

[0149] Afterwards, the prompt information can be reconstructed using the newly decoded data “Mr. Zhang did this” in the second decoding 630 , and the voice data in the voice buffer can be decoded using the reconstructed prompt information.

[0150] In the third decoding 640:

[0151] After decoding the current recognition sentence "Today, I am very grateful to Mr. Zhang for his report.", the voice data "Thank you, Mr. President." cached in the voice buffer is removed, and the voice data corresponding to "Today, I am very grateful to Mr. Zhang for his report." is retained in the voice buffer to serve as voice context data for subsequent sentences;

[0152] At the same time, "Thank you, Mr. President. Today, I am very grateful to Mr. Zhang for his report" is used as the text context data for subsequent decoding;

[0153] In this way, the next sentence is decoded based on the new speech context and text context data.

[0154] Based on the foregoing embodiments, the present application provides a speech processing device, which includes the various units included and the various modules included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.

[0155] Figure 7 A schematic diagram of the structure of a speech processing device provided in this application is shown in FIG. Figure 7 As shown, the speech processing device 700 includes: a speech reading module 710 and a decoding module 720, wherein:

[0156] The speech reading module 710 is configured to read first speech data and second speech data from the speech buffer; wherein the first speech data represents at least part of the speech data in the currently recognized sentence; and the second speech data represents decoded speech data before the currently recognized sentence;

[0157] The decoding module 720 is used to decode the first voice data according to the first prompt information and the second voice data to obtain a first decoding result corresponding to the first voice data; the first prompt information at least includes a second decoding result corresponding to the second voice data.

[0158] In some embodiments, the speech processing device 700 further includes:

[0159] a first updating module, configured to update the first prompt information using the first decoding result to obtain second prompt information;

[0160] The speech reading module 710 is further configured to read third speech data from the speech buffer; the third speech data represents speech data in the currently recognized sentence that is buffered after the first speech data;

[0161] The decoding module 720 is further configured to decode the encoded data of the third voice data according to the second prompt information and the encoded data of the second voice data to obtain a third decoding result.

[0162] In some embodiments, the first updating module is configured to update the first prompt information using a first partial result of the first decoding result to obtain the second prompt information; the first decoding result includes a first partial result decoded earlier and a second partial result decoded later;

[0163] The decoding module 720 is also used to perform decoding on the encoded data corresponding to the second partial result and the encoded data corresponding to the third voice data according to the second prompt information and the encoded data of the second voice data, to obtain a fourth decoding result and a fifth decoding result, respectively; if the fourth decoding result matches the second partial result, the fifth decoding result is used as the third decoding result.

[0164] In some embodiments, the first updating module is further configured to, in response to the fourth decoding result not matching the second partial result, delete the decoding result of the current recognized sentence from the first prompt information to obtain third prompt information;

[0165] The decoding module 720 is further configured to re-decode from the starting position of the currently recognized sentence according to the third prompt information and the second voice data.

[0166] In some embodiments, the first decoding result further includes a third partial result; the third partial result represents a decoding result generated after the second partial result;

[0167] The speech processing device 700 further includes an output determination module; the output determination module is configured to use the text data corresponding to the first partial result and the second partial result as output results.

[0168] In some implementations, the speech processing device 700 further includes:

[0169] a first updating module, configured to update the first prompt information using the first decoding result to obtain second prompt information;

[0170] a second updating module, configured to, in response to a voice decoding result in the second prompt information being the same as a voice decoding result in the first prompt information, remove the voice data of the first duration from the head of the voice buffer, and add the voice data to be recognized of the first duration to the tail of the voice buffer;

[0171] The decoding module 720 is configured to decode the voice data in the voice buffer using the first prompt information.

[0172] In some implementations, the speech processing device 700 further includes:

[0173] A sentence segmentation module, configured to perform sentence segmentation processing on the speech data to be recognized to obtain at least one sentence;

[0174] The second updating module is configured to store the at least one sentence in the speech buffer.

[0175] In addition, the present application also provides an electronic device. Figure 8 As shown, the electronic device 800 includes at least one processor 810; a speech recognition model is running on the at least one processor 810; wherein,

[0176] The speech recognition model is used to read first speech data and second speech data from a speech buffer; wherein the first speech data represents at least part of the speech data in a currently recognized sentence; the second speech data represents decoded speech data before the currently recognized sentence; based on first prompt information and the encoded data of the second speech data, the encoded data of the first speech data is decoded to obtain a first decoding result corresponding to the first speech data; the first prompt information at least includes a second decoding result corresponding to the second speech data.

[0177] It should be noted that the description of the above device and equipment embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. In some embodiments, the functions, units, components, or modules included in the devices and equipment provided in the embodiments of the present application can be used to perform the methods described in the above method embodiments. For technical details not disclosed in the device and equipment embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0178] In the embodiment of the present application, if the above-mentioned speech processing method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific hardware, software or firmware, or any combination of hardware, software and firmware.

[0179] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above method. The computer-readable storage medium may be transient or non-transient.

[0180] An embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code is run in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.

[0181] An embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, implements some or all of the steps in the above method. The computer program product can be implemented specifically by hardware, software, or a combination thereof. In some embodiments, the computer program product is embodied as a computer storage medium. In other embodiments, the computer program product is embodied as a software product, such as a software development kit (SDK), etc.

[0182] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between the various embodiments, and their similarities or similarities can be referenced to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the description of the method embodiments of this application for understanding.

[0183] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.

[0184] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0185] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0186] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0187] In addition, all functional units in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.

[0188] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.

[0189] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0190] The above is only an implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.

Claims

1. A speech processing method, comprising: Reading first voice data and second voice data from a voice buffer; wherein the first voice data represents at least part of the voice data in the current recognition sentence; and the second voice data represents decoded voice data before the current recognition sentence; According to the first prompt information and the encoded data of the second voice data, the encoded data of the first voice data is decoded to obtain a first decoding result corresponding to the first voice data; the first prompt information at least includes a second decoding result corresponding to the second voice data.

2. The method according to claim 1, after obtaining the first decoding result corresponding to the first speech data, further comprising: Using the first decoding result to update the first prompt information to obtain second prompt information; Reading third voice data from the voice buffer; The third voice data represents the voice data in the currently recognized sentence that is buffered after the first voice data; The encoded data of the third voice data is decoded according to the second prompt information and the encoded data of the second voice data to obtain a third decoding result.

3. The method according to claim 2, wherein the updating of the first prompt information using the first decoding result to obtain the second prompt information comprises: Using a first part of the first decoding result to update the first prompt information, to obtain the second prompt information; The first decoding result includes a first partial result decoded first and a second partial result decoded later; The method of decoding the third voice data according to the second prompt information and the encoded data of the second voice data to obtain a third decoding result includes: decoding the encoded data corresponding to the second partial result and the encoded data of the third voice data according to the second prompt information and the encoded data of the second voice data to obtain a fourth decoding result and a fifth decoding result, respectively; if the fourth decoding result matches the second partial result, using the fifth decoding result as the third decoding result.

4. The method according to claim 3, further comprising: In response to the fourth decoding result not matching the second partial result, deleting the decoding result of the current recognized sentence from the first prompt information to obtain third prompt information; Re-decoding is performed from the starting position of the currently recognized sentence according to the third prompt information and the second voice data.

5. The method according to claim 3, wherein the first decoding result further includes a third partial result; the third partial result represents a decoding result generated after the second partial result; The method further comprises: The text data corresponding to the first partial result and the second partial result are used as output results.

6. The method according to any one of claims 1 to 5, further comprising: Using the first decoding result to update the first prompt information to obtain second prompt information; In response to the voice decoding result in the second prompt information being the same as the voice decoding result in the first prompt information, removing the voice data of the first duration at the head of the voice buffer, and adding the voice data to be recognized of the first duration to the tail of the voice buffer; The voice data in the voice buffer is decoded using the first prompt information.

7. The method according to any one of claims 1 to 6, further comprising: Perform sentence segmentation processing on the speech data to be recognized to obtain at least one sentence; The at least one sentence is stored in the speech buffer.

8. A speech processing device, comprising: A speech reading module is configured to read first speech data and second speech data from a speech buffer; wherein the first speech data represents at least part of speech data in a currently recognized sentence; and the second speech data represents decoded speech data before the currently recognized sentence; A decoding module is used to decode the first voice data according to the first prompt information and the second voice data to obtain a first decoding result corresponding to the first voice data; the first prompt information at least includes a second decoding result corresponding to the second voice data.

9. An electronic device comprising at least one processor; a speech recognition model is running on the at least one processor; wherein: The speech recognition model is used to read first speech data and second speech data from a speech buffer; wherein the first speech data represents at least part of the speech data in a currently recognized sentence; the second speech data represents decoded speech data before the currently recognized sentence; based on first prompt information and the encoded data of the second speech data, the encoded data of the first speech data is decoded to obtain a first decoding result corresponding to the first speech data; the first prompt information at least includes a second decoding result corresponding to the second speech data.

10. A computer program product, comprising a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is read and executed by a computer, the steps of the method according to any one of claims 1 to 7 are implemented.