Punctuation and Smoothness Integrated Speech Recognition Method, System and Electronic Device

By adopting a punctuation and smooth integrated speech recognition model on the localized terminal, the problem of slow real-time long speech transcription recognition speed caused by the limitation of computing resources is solved, efficient speech recognition and punctuation and smooth prediction are achieved, and the readability and application scope of the recognition software are improved.

CN116229968BActive Publication Date: 2025-06-17AISPEECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310145773.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-21
Publication Date
2025-06-17
Estimated Expiration
2043-02-21

AI Technical Summary

Technical Problem

When real-time long voice transliteration function is implemented on localized terminals, computing resources are limited, resulting in slow recognition speed. If the model is reduced to improve the speed, it will affect the recognition accuracy.

Method used

A punctuation and smooth integrated speech recognition model is adopted. The model includes an encoder and a decoder for text prediction, punctuation prediction, and smooth prediction. It can identify, punctuation and smooth prediction through hidden layer features to reduce resource usage.

Benefits of technology

In the localization scenario where computing resources are limited, voice recognition performance comparable to cloud is achieved, the subjective readability of recognition software is improved, and the application scope of long speech recognition transcription localization is expanded.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116229968B_ABST
    Figure CN116229968B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a speech recognition method, system and electronic device for punctuation and smoothness integration. The method includes: inputting an audio into a recognition-punctuation-smoothness integration model; determining hidden layer features of the audio through an encoder; a decoder sequentially performs recognition prediction on m characters in the audio according to the hidden layer features, performs punctuation prediction and smoothness prediction after the nth character after the recognition prediction of the nth character, obtains an intermediate recognition result, and performs recognition prediction, punctuation prediction and smoothness prediction of the (n + 1)th character according to the intermediate recognition result and the hidden layer features until the prediction of the mth character is completed, and obtains a final recognition result. The embodiment of the present invention reduces the occupation of computing resources, can be applied to intelligent devices with relatively weak computing capabilities, and expands the application range of long speech recognition and transcription localization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent speech, and in particular to a speech recognition method, system and electronic device integrating punctuation and smoothness. Background Art

[0002] With the development of intelligent voice, products with long voice real-time transcription function provide voice services for users. Among them, long voice refers to the scene of continuous voice recognition, such as conference transcription, audio and video subtitles, etc. Real-time transcription refers to the streaming recognition of voice results, punctuation, and smoothing of spoken phrases (such as: um, ah) while speaking.

[0003] The speech recognition model of the traditional transcription function is based on cloud-based reasoning, and usually adopts a separate model approach. That is, one model is used for speech recognition, one model is used for punctuation, and one model is used for smooth spoken language.

[0004] In the process of implementing the present invention, the inventors found that there are at least the following problems in the related art:

[0005] As users' demand for privacy protection continues to grow, more and more voice services refuse to transmit voice data to the cloud for recognition. Traditional local end-side recognition only supports simple command control, and the performance of long speech real-time recognition is far from the cloud level. Compared with cloud devices, the computing resources on the local end are very limited.

[0006] The traditional method uses three sets of models with large parameters, which takes up a lot of computing resources. The cloud has powerful processing performance to process, but the computing capacity of the user's local smart device is relatively weak and the computing resources are limited. The recognition speed is relatively slow in local scenarios with limited computing resources. Although the model can be reduced to improve the speed of local recognition, this will seriously affect the accuracy of recognition. Summary of the invention

[0007] In order to at least solve the problem of localized real-time transcription of long speech in the existing technology, due to the relatively weak computing power and limited computing resources of the user's local smart device, the recognition speed is relatively slow. If the model is reduced to ensure the recognition speed, it will affect the recognition accuracy.

[0008] In a first aspect, an embodiment of the present invention provides a punctuation and smooth integrated speech recognition method, comprising:

[0009] Input the audio to a recognition-punctuation-smoothness integrated model, wherein the recognition-punctuation-smoothness integrated model includes an encoder and a decoder for text prediction, punctuation prediction, and smoothness prediction;

[0010] Determining, by the encoder, latent features of the audio;

[0011] The decoder sequentially performs recognition and prediction on m characters in the audio according to the hidden layer features, and performs punctuation prediction and smoothing prediction after the nth character after the recognition and prediction of the nth character, to obtain an intermediate recognition result, and performs recognition prediction, punctuation prediction and smoothing prediction of the (n + 1)th character according to the intermediate recognition result and the hidden layer features, until the prediction of the mth character is completed, to obtain a final recognition result, where 1 ≤ n ≤ m.

[0012] In a second aspect, an embodiment of the present invention provides a speech recognition system for integrated punctuation and smoothing, including:

[0013] A data receiving program module, configured to input an audio into a recognition-punctuation-smoothing integrated model, where the recognition-punctuation-smoothing integrated model includes an encoder and a decoder for text prediction, punctuation prediction, and smoothing prediction;

[0014] An encoding program module, configured to determine the hidden layer features of the audio through the encoder;

[0015] An integrated recognition program module, configured to enable the decoder to sequentially perform recognition and prediction on m characters in the audio according to the hidden layer features, perform punctuation prediction and smoothing prediction after the nth character after the recognition and prediction of the nth character, to obtain an intermediate recognition result, and perform recognition prediction, punctuation prediction and smoothing prediction of the (n + 1)th character according to the intermediate recognition result and the hidden layer features, until the prediction of the mth character is completed, to obtain a final recognition result, where 1 ≤ n ≤ m.

[0016] In a third aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the steps of the method for integrated punctuation and smoothing of speech recognition according to any embodiment of the present invention.

[0017] In a fourth aspect, an embodiment of the present invention provides a storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the method for integrated punctuation and smoothing of speech recognition according to any embodiment of the present invention are implemented.

[0018] The beneficial effects of the embodiments of the present invention are as follows: For the real-time long speech recognition application for local terminals, the recognition-punctuation-smoothing integrated model of this method can greatly reduce resource occupancy. Only a very small amount of additional computational effort is required to solve the punctuation and smoothing problems, which has no impact on the performance of speech recognition while greatly improving the subjective readability of the recognition software. The local speech recognition for user privacy can be further developed and can be applied to intelligent devices with relatively weak computing power, expanding the application scope of long speech recognition transcription localization. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0020] Figure 1 is a flowchart of a speech recognition method with integrated punctuation and smoothing provided by an embodiment of the present invention;

[0021] Figure 2 is a structural schematic diagram of a speech recognition method with integrated punctuation and smoothing provided by an embodiment of the present invention;

[0022] Figure 3 is an integrated flowchart schematic diagram of a speech recognition method with integrated punctuation and smoothing provided by an embodiment of the present invention;

[0023] Figure 4 is a schematic diagram of memory delay test data of a speech recognition method with integrated punctuation and smoothing provided by an embodiment of the present invention;

[0024] Figure 5 is a schematic diagram of recognition rate and F1 test data of a speech recognition method with integrated punctuation and smoothing provided by an embodiment of the present invention;

[0025] Figure 6 is a structural schematic diagram of a speech recognition system with integrated punctuation and smoothing provided by an embodiment of the present invention;

[0026] Figure 7 is a structural schematic diagram of an embodiment of an electronic device for speech recognition with integrated punctuation and smoothing provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0028] As Figure 1 shown in the flowchart of a punctuation and smooth integrated speech recognition method provided by an embodiment of the present invention, the method includes the following steps:

[0029] S11: Input the audio into an identification-punctuation-smooth integration model, where the identification-punctuation-smooth integration model includes an encoder and a decoder for text prediction, punctuation prediction, and smooth prediction;

[0030] S12: Determine the hidden layer features of the audio through the encoder;

[0031] S13: The decoder sequentially performs identification prediction on m characters in the audio according to the hidden layer features, performs punctuation prediction and smooth prediction after the nth character after the identification prediction of the nth character, obtains an intermediate identification result, and performs identification prediction, punctuation prediction, and smooth prediction of the (n + 1)th character according to the intermediate identification result and the hidden layer features until the prediction of the mth character is completed, obtaining a final identification result, where 1 ≤ n ≤ m.

[0032] In this embodiment, it is found that in the prior art, to implement the long speech real-time transcription function, three sets of models are required: one set of models for speech recognition, one set of models for punctuation, and one set of models for oral smoothness. The parameter quantities of the three sets of models are relatively large, which consume a lot of computing resources and are difficult to apply to the user's local scenario. At the same time, during the training process of the three sets of models in the prior art, the input during the training and inference of punctuation is mismatched. During training, most of the corpora are obtained by crawling network resources and are mainly text corpora. The input during inference is the text of the speech recognition result, which results in a mismatch in the training between the models and may cause recognition errors, thereby affecting the accuracy of long speech real-time transcription. Considering the above defects, this method no longer regards punctuation generation and oral smoothness as independent post-processing modules, but combines them with speech recognition. In actual use, it can achieve performance comparable to that of the cloud with only a small amount of additional parameters and computing resources.

[0033] For step S11, this method can be adapted to intelligent devices, such as meeting recording devices, so that users can use the meeting recording devices equipped with this method for real-time transcription of long speeches. Users pay more attention to the meeting content and do not want to use the cloud for real-time transcription. As an implementation, the recognition-punctuation-smoothing integrated model of this method is applied to local speech recognition. In this implementation, users can use the meeting recording device for local real-time transcription, ensuring the security and privacy of data.

[0034] The meeting recording device continuously maintains VAD (Voice Activity Detection) and continuously collects the audio input by the user during the meeting in real time. The real-time audio is input into the recognition-punctuation-smoothing integrated model. The recognition-punctuation-smoothing integrated model includes an encoder and decoders for text prediction, punctuation prediction, and smoothing prediction. The structure of the recognition-punctuation-smoothing integrated model of this method is as Figure 2 shown.

[0035] For step S12, the audio features of the meeting audio collected in real time are extracted, the hidden layer features are determined through the encoder, and further accurate prediction is made through the hidden layer features.

[0036] For step S13, the real-time transcription of the long speech of the meeting record does not input a whole segment of audio, but inputs it continuously over time. Therefore, the decoder needs to identify and predict the text at the current time based on the continuously determined hidden layer features.

[0037] The decoder can determine whether the audio in the current period is the user's speech based on the hidden layer features. If it is the user's speech, direct speech recognition is performed to obtain the recognition result of the current word. Considering that this method integrates punctuation prediction and smoothing prediction, the audio of a fixed number of frames following the current word is detected. For example, if it is detected that there is no user speech and a pause occurs, the punctuation after the current word can be determined. When determining whether it is the user's speech, the encoder can directly mark the smoothing state of the determined text or punctuation. For example, if it is speech, it is marked as 0, and if there is no user speech and a pause occurs, it is marked as 1 (the above markings are only examples, and the marking method can be adjusted according to different requirements and is not limited here). As Figure 3 shown, it is the recognition process of the recognition-punctuation-smoothing integration. When the first word "bye" is recognized, continue to judge the punctuation prediction and smoothing prediction after "bye": "no punctuation" and "0" to obtain the intermediate recognition result. Based on the intermediate recognition result, the prediction of the next word is made, and the subsequent recognized words are corrected and adjusted based on the historically determined intermediate recognition results. Through the above prediction method until all the text is predicted, that is, until the meeting ends and the user no longer inputs speech.

[0038] For the training of the recognition-punctuation-smoothing integrated model, which is obtained by training with training audio and the reference recognition result with punctuation of the training audio, it includes:

[0039] Determine the hidden layer features of the training audio through the encoder of the recognition-punctuation-smoothing integrated model;

[0040] The decoder of the recognition-punctuation-smoothing integrated model sequentially performs recognition predictions on m characters in the training audio according to the hidden layer features. After the nth character recognition prediction, perform punctuation prediction and smoothing prediction after the nth character to obtain an intermediate recognition result, and perform the (n + 1)th character recognition prediction, punctuation prediction, and smoothing prediction according to the intermediate recognition result and the hidden layer features until the final predicted recognition result is obtained after predicting the mth character;

[0041] Train the recognition-punctuation-smoothing integrated model based on the loss function determined by the reference recognition result and the final predicted recognition result until the text and punctuation in the final predicted recognition result approach the reference recognition result.

[0042] In this embodiment, the training process is the error function for comparing the predicted recognition result of the training audio with the pre-prepared reference recognition result. The process of predicted recognition has been described in steps S11 - S13 and will not be elaborated here. Through training, the integrated model can achieve operations of speech recognition, punctuation marking, and sentence smoothing with only a small amount of additional parameters and computing resources.

[0043] Train the recognition-punctuation-smoothing integrated model trained by this method, as Figure 4 shown, which shows the memory occupied and latency used by the model, and it can be seen that it is superior to the existing methods.

[0044] As Figure 5 shown, it is a display of the recognition rate and F1 value, and it can be seen that the long speech transcription performance of this method is basically equivalent to that of the existing technology.

[0045] It can be seen from this embodiment that for the real-time long speech recognition application facing the local terminal, the recognition-punctuation-smoothing integrated model of this method can greatly reduce the resource occupation. Only a very small amount of additional computation is required to solve the punctuation and smoothing problems, which has no impact on the performance of speech recognition while greatly improving the subjective readability of the recognition software. The local speech recognition for user privacy can be further developed and can be applied to intelligent devices with relatively weak computing power, expanding the application scope of long speech recognition and transcription localization.

[0046] As Figure 6The following is a schematic structural diagram of a punctuation and smoothness integrated speech recognition system provided by an embodiment of the present invention. The system can execute the punctuation and smoothness integrated speech recognition method described in any of the above embodiments and is configured in a terminal.

[0047] A punctuation and smoothness integrated speech recognition system 10 provided in this embodiment includes: a data receiving program module 11, an encoding program module 12, and an integrated recognition program module 13.

[0048] Among them, the data receiving program module 11 is used to input audio into an identification-punctuation-smoothness integrated model, where the identification-punctuation-smoothness integrated model includes an encoder and decoders for text prediction, punctuation prediction, and smoothness prediction; the encoding program module 12 is used to determine the hidden layer features of the audio through the encoder; the integrated recognition program module 13 is used for the decoder to sequentially perform recognition predictions on m characters in the audio, perform punctuation prediction and smoothness prediction after the nth character after the nth character recognition prediction to obtain an intermediate recognition result, and perform the n+1th character recognition prediction, punctuation prediction, and smoothness prediction based on the intermediate recognition result and the hidden layer features until the mth character is predicted to obtain a final recognition result, where 1≤n≤m.

[0049] The embodiment of the present invention also provides a non-volatile computer storage medium. The computer storage medium stores computer-executable instructions, and the computer-executable instructions can execute the punctuation and smoothness integrated speech recognition method in any of the above method embodiments;

[0050] As an implementation manner, the non-volatile computer storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as:

[0051] Input audio into an identification-punctuation-smoothness integrated model, where the identification-punctuation-smoothness integrated model includes an encoder and decoders for text prediction, punctuation prediction, and smoothness prediction;

[0052] Determine the hidden layer features of the audio through the encoder;

[0053] The decoder sequentially performs recognition predictions on m characters in the audio according to the hidden layer features, performs punctuation prediction and smoothness prediction after the nth character after the nth character recognition prediction to obtain an intermediate recognition result, and performs the n+1th character recognition prediction, punctuation prediction, and smoothness prediction based on the intermediate recognition result and the hidden layer features until the mth character is predicted to obtain a final recognition result, where 1≤n≤m.

[0054] As a non-volatile computer-readable storage medium, it can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the method in the embodiments of the present invention. One or more program instructions are stored in the non-volatile computer-readable storage medium and, when executed by a processor, perform the punctuation and smooth integrated speech recognition method in any of the above method embodiments.

[0055] Figure 7 It is a schematic diagram of the hardware structure of an electronic device for the punctuation and smooth integrated speech recognition method provided in another embodiment of the present application, as Figure 7 shown. The device includes:

[0056] One or more processors 710 and a memory 720, Figure 7 Taking one processor 710 as an example. The device for the punctuation and smooth integrated speech recognition method may further include: an input device 730 and an output device 740.

[0057] The processor 710, the memory 720, the input device 730, and the output device 740 can be connected through a bus or other means, Figure 7 Taking connection through a bus as an example.

[0058] The memory 720, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the punctuation and smooth integrated speech recognition method in the embodiments of the present application. The processor 710 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 720, that is, implements the punctuation and smooth integrated speech recognition method in the above method embodiments.

[0059] The memory 720 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data, etc. In addition, the memory 720 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 720 may optionally include a memory remotely set relative to the processor 710, and these remote memories can be connected to the mobile device through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0060] The input device 730 can receive input digital or character information. The output device 740 may include a display device such as a display screen.

[0061] The one or more modules are stored in the memory 720 and, when executed by the one or more processors 710, perform the punctuation and smooth integrated speech recognition method in any of the above method embodiments.

[0062] The above product can execute the method provided in the embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference may be made to the method provided in the embodiments of the present application.

[0063] The non-volatile computer-readable storage medium may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the device, etc. In addition, the non-volatile computer-readable storage medium may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely provided with respect to the processor, and these remote memories may be connected to the device through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0064] An embodiment of the present invention further provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the punctuation and smooth integrated speech recognition method in any embodiment of the present invention.

[0065] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to:

[0066] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0067] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDA, MID, and UMPC devices, such as tablet computers.

[0068] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and smart toys and portable in-vehicle navigation devices.

[0069] (4) Other electronic devices with data processing functions.

[0070] In this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising" and "including" not only include those elements, but also other elements not expressly listed, or also include elements inherent to such a process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the said element.

[0071] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed over multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.

[0072] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0073] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A punctuation and smoothness integrated speech recognition method, comprising: Input the audio into the integrated recognition-punctuation-smoothing model, where the integrated recognition-punctuation-smoothing model includes an encoder and a decoder for text prediction, punctuation prediction, and smoothing prediction; Determine the hidden layer features of the audio through the encoder; The decoder sequentially performs recognition prediction on m characters in the audio according to the hidden layer features, performs punctuation prediction and smoothing prediction after the nth character after the recognition prediction of the nth character, obtains an intermediate recognition result, and performs recognition prediction, punctuation prediction, and smoothing prediction on the (n + 1)th character according to the intermediate recognition result and the hidden layer features until the prediction of the mth character is completed, obtaining a final recognition result, where 1 ≤ n ≤ m.

2. The method according to claim 1, wherein, The integrated recognition-punctuation-smoothing model is trained by training audio and the benchmark recognition result with punctuation of the training audio, including: Determine the hidden layer features of the training audio through the encoder of the integrated recognition-punctuation-smoothing model; The decoder of the integrated recognition-punctuation-smoothing model sequentially performs recognition prediction on m characters in the training audio according to the hidden layer features, performs punctuation prediction and smoothing prediction after the nth character after the recognition prediction of the nth character, obtains an intermediate recognition result, and performs recognition prediction, punctuation prediction, and smoothing prediction on the (n + 1)th character according to the intermediate recognition result and the hidden layer features until the prediction of the mth character is completed to obtain a final predicted recognition result; Train the integrated recognition-punctuation-smoothing model based on the loss function determined by the benchmark recognition result and the final predicted recognition result until the text and punctuation in the final predicted recognition result approach the benchmark recognition result.

3. The method according to claim 1, wherein, The integrated recognition-punctuation-smoothing model is applied to local speech recognition.

4. The method according to claim 1, wherein, When n = 1, the decoder directly determines the character recognition result of the first character according to the hidden layer features.

5. A punctuation and smoothness integrated speech recognition system, comprising: A data receiving program module for inputting the audio into the integrated recognition-punctuation-smoothing model, where the integrated recognition-punctuation-smoothing model includes an encoder and a decoder for text prediction, punctuation prediction, and smoothing prediction; An encoding program module for determining the hidden layer features of the audio through the encoder; An integrated recognition program module for the decoder to sequentially perform recognition prediction on m characters in the audio according to the hidden layer features, perform punctuation prediction and smoothing prediction after the nth character after the recognition prediction of the nth character, obtain an intermediate recognition result, and perform recognition prediction, punctuation prediction, and smoothing prediction on the (n + 1)th character according to the intermediate recognition result and the hidden layer features until the prediction of the mth character is completed, obtaining a final recognition result, where 1 ≤ n ≤ m.

6. The system according to claim 5, wherein, The integrated recognition-punctuation-smoothing model is trained by training audio and the benchmark recognition result with punctuation of the training audio, including: Determine the hidden layer features of the training audio through the encoder of the integrated recognition-punctuation-smoothing model; The decoder of the recognition-punctuation-smoothing integrated model sequentially performs recognition predictions on m characters in the training audio according to the hidden layer features. After the recognition prediction of the nth character, punctuation prediction and smoothing prediction after the nth character are performed to obtain an intermediate recognition result, and recognition prediction, punctuation prediction and smoothing prediction of the (n + 1)th character are performed according to the intermediate recognition result and the hidden layer features until the final prediction recognition result is obtained after predicting the mth character; The recognition-punctuation-smoothing integrated model is trained based on the loss function determined by the reference recognition result and the final prediction recognition result until the text and punctuation in the final prediction recognition result approach the reference recognition result.

7. The system according to claim 5, wherein, The recognition-punctuation-smoothing integrated model is applied to local speech recognition.

8. The system according to claim 5, wherein, When n = 1, the decoder directly determines the character recognition result of the first character according to the hidden layer features.

9. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method according to any one of claims 1-4.

10. A storage medium, on which a computer program is stored, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Voice interactive method based on VoiceXML movable termination and movable termination

    CN101527755A

  • Three-dimensional vehicular access collaborative simulation system

    CN103021026A