Punctuation prediction method and apparatus, and speech recognition device

By combining a punctuation prediction model with original audio information to correct the probability of non-punctuation tags in transcribed text, the problem of inaccurate punctuation prediction in speech recognition is solved, improving the readability of transcribed text and the efficiency of voice interaction.

CN115662432BActive Publication Date: 2025-11-04HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211184502.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2025-11-04
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

In existing technologies, punctuation prediction in speech recognition is not very accurate when the transcribed text is too short, and it is easy to mis-punctuate or omit punctuation.

Method used

The probability of punctuation tags and non-punctuation tags appearing after each character in the transcribed text is obtained by using a punctuation prediction model. The probability of non-punctuation tags is then corrected by combining the first and second text information of the original audio, and the corrected prediction probability information of non-punctuation tags is obtained.

Benefits of technology

It improves the accuracy and rationality of punctuation prediction, enhances the readability of transcribed text, and improves the efficiency of voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115662432B_ABST
    Figure CN115662432B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a punctuation prediction method, device and speech recognition equipment. The method comprises: obtaining, based on a punctuation prediction model, a probability of a punctuation label and a probability of a non-punctuation label appearing after each character in a transcription text, obtaining first text information and second text information corresponding to the original audio, correcting the probability of the non-punctuation label appearing after each character in the transcription text according to the first text information and the second text information, and obtaining the prediction probability information of the corrected non-punctuation label. The transcription text is a text sequence obtained by performing speech recognition processing on the original audio; the first text information is text information obtained by performing audio truncation processing on the original audio, including speech feature information and non-speech feature information corresponding to the original audio; and the second text information is text information obtained by performing decoding processing on the original audio, including transcription text characters and transcription space characters corresponding to the original audio. The present method can improve the accuracy of punctuation prediction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of natural language processing, and particularly relates to a punctuation prediction method and device and a speech recognition device. BACKGROUND

[0002] In the scene of human-computer interaction, speech recognition plays a crucial role in natural language understanding and natural language generation. Punctuation prediction on the text transcribed by speech recognition is an important work for semantic understanding and interaction. Correct punctuation annotation plays a great auxiliary role in semantic understanding. Ambiguous or mismatched punctuation will have the effect of words not conveying the intended meaning and misleading, and thus affect the overall process of voice interaction.

[0003] In related technologies, punctuation prediction in speech recognition is mostly based on the context of the transcribed text to output the punctuation of the whole sentence. However, when the transcribed text is too short, there is not enough information available, and the above-mentioned method has a high probability of mislabeling or missing labeling. Therefore, how to improve the accuracy of punctuation prediction in speech recognition is a current problem to be solved. SUMMARY

[0004] To solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a display device and a punctuation prediction method, which can determine the semantic understanding content corresponding to the word segmentation of various food materials when the user speaks the storage of various food materials, effectively improve the semantic understanding ability of the display device, and further improve the user experience.

[0005] When a user watches multimedia content, target barrage content corresponding to the multimedia content is provided for the user, so that the user can send the barrage content that the user wants to send, and the user's watching experience and satisfaction are improved.

[0006] In a first aspect, the present disclosure provides a punctuation prediction method, comprising:

[0007] obtaining, based on a punctuation prediction model, a probability of a punctuation label and a probability of a non-punctuation label after each character in a transcribed text; the transcribed text is a text sequence obtained by performing speech recognition on original audio, and the punctuation prediction model is used to output the probability of the punctuation label and the probability of the non-punctuation label after each character in the transcribed text;

[0008] obtaining first text information corresponding to the original audio and second text information corresponding to the original audio; the first text information is text information obtained by performing audio truncation processing on the original audio, the first text information includes speech feature information and non-speech feature information corresponding to the original audio, and the second text information is text information obtained by performing decoding processing on the original audio, the second text information includes transcribed text characters and transcribed space characters corresponding to the original audio;

[0009] correcting a probability of a non-punctuation label appearing after each character in the transcribed text according to the first text information and the second text information, and obtaining predicted probability information of the corrected non-punctuation label.

[0010] As an optional implementation of an embodiment of the present disclosure, the correcting a probability of a non-punctuation label appearing after each character in the transcribed text according to the first text information and the second text information, and obtaining predicted probability information of the corrected non-punctuation label includes:

[0011] obtaining a first normalized position weight, a second normalized position weight, and a third normalized position weight;

[0012] The first normalized position weight is a normalized weight of the number of space characters after each character in the first text information, the second normalized position weight is a normalized weight of the number of space characters after each character in the second text information, and the third normalized position weight is a sum of the first normalized position weight and the second normalized position weight.

[0013] obtaining a corrected probability of a non-punctuation label appearing after each character according to the probability of a non-punctuation label appearing after each character in the transcribed text, the third normalized position weight, and an audio intervention adjustment parameter;

[0014] obtaining predicted probability information of the corrected non-punctuation label according to the corrected probability of a non-punctuation label appearing after each character.

[0015] As an optional implementation of an embodiment of the present disclosure, the method further includes:

[0016] The probability of a punctuation label appearing after each character in the transcribed text and the probability of a non-punctuation label appearing after each character in the transcribed text are negatively correlated.

[0017] As an optional implementation of an embodiment of the present disclosure, the obtaining a corrected probability of a non-punctuation label appearing after each character according to the probability of a non-punctuation label appearing after each character in the transcribed text, the third normalized position weight, and an audio intervention adjustment parameter includes:

[0018] According to the preset correction manner, the probability of the non-punctuation label appearing after each character in the transcribed text, the third normalized position weight, and the audio intervention adjustment parameter are calculated to obtain the probability of the non-punctuation label appearing after each character after correction.

[0019] As an optional implementation of the embodiment of the present disclosure, when the preset correction manner is the first preset correction manner, the calculation of the probability of the non-punctuation label appearing after each character in the transcribed text, the third normalized position weight, and the audio intervention adjustment parameter according to the preset correction manner to obtain the probability of the non-punctuation label appearing after each character after correction includes:

[0020] The probability of the non-punctuation label appearing after each character after correction is determined according to the sum of the probability of the non-punctuation label appearing after each character in the transcribed text and the product of the audio intervention adjustment parameter and the third normalized position weight.

[0021] As an optional implementation of the embodiment of the present disclosure, when the preset correction manner is the second preset correction manner, the calculation of the probability of the non-punctuation label appearing after each character in the transcribed text, the third normalized position weight, and the audio intervention adjustment parameter according to the preset correction manner to obtain the probability of the non-punctuation label appearing after each character after correction includes:

[0022] The probability of the non-punctuation label appearing after each character after correction is determined according to the product of the probability of the non-punctuation label appearing after each character in the transcribed text, the audio intervention adjustment parameter, and the third normalized position weight.

[0023] As an optional implementation of the embodiment of the present disclosure, when the preset correction manner is the first preset correction manner, the obtaining of the first normalized position weight, the second normalized position weight, and the third normalized position weight includes:

[0024] The first normalized position weight is obtained according to the total number of empty characters in the first text information, the average number of empty characters in the first text information, and the number of empty characters after each character in the first text information.

[0025] The second normalized position weight is obtained according to the total number of empty characters in the second text information, the average number of empty characters in the second text information, and the number of empty characters after each character in the second text information.

[0026] The third normalized position weight is obtained according to the sum of the first normalized position weight and the second normalized position weight.

[0027] As an optional implementation of the embodiment of the present disclosure, when the preset correction manner is the second preset correction manner, the obtaining the first normalized position weight, the second normalized position weight, and the third normalized position weight comprises:

[0028] obtaining the first normalized position weight according to the maximum number of space characters after each character in the N characters of the first text information, the number of space characters after each character in the first text information, and a scaling factor;

[0029] obtaining the second normalized position weight according to the maximum number of space characters after each character in the N characters of the second text information, the number of space characters after each character in the second text information, and a scaling factor;

[0030] obtaining the third normalized position weight according to the sum of the first normalized position weight and the second normalized position weight;

[0031] wherein N is an integer greater than or equal to 1.

[0032] In a second aspect, a semantic understanding apparatus is provided, and the method comprises:

[0033] a punctuation probability obtaining module configured to obtain, based on a punctuation prediction model, a probability of a punctuation label and a probability of a non-punctuation label after each character in a transcription text; the transcription text is a text sequence obtained by performing speech recognition on original audio, and the punctuation prediction model is configured to output the probability of the punctuation label and the probability of the non-punctuation label after each character in the transcription text;

[0034] an audio text obtaining module configured to obtain first text information corresponding to original audio and second text information corresponding to the original audio; the first text information is text information obtained by performing audio truncation on the original audio, and the first text information comprises speech feature information and non-speech feature information corresponding to the original audio; the second text information is text information obtained by performing decoding on the original audio, and the second text information comprises transcription text characters and transcription space characters corresponding to the original audio;

[0035] a punctuation probability correction module configured to correct, according to the first text information and the second text information, the probability of the non-punctuation label after each character in the transcription text, and obtain predicted probability information of the corrected non-punctuation label.

[0036] As an optional implementation of the embodiment of the present disclosure, the punctuation probability correction module comprises:

[0037] a weight obtaining unit, configured to obtain a first normalized position weight, a second normalized position weight, and a third normalized position weight;

[0038] The first normalized position weight is a normalized weight of the number of space characters after each character in the first text information, the second normalized position weight is a normalized weight of the number of space characters after each character in the second text information, and the third normalized position weight is a sum of the first normalized position weight and the second normalized position weight.

[0039] a character probability correction unit, configured to obtain a probability of a non-punctuation label appearing after each character in the corrected text according to a probability of a non-punctuation label appearing after each character in the transcription text, the third normalized position weight, and an audio intervention adjustment parameter;

[0040] a punctuation probability correction unit, configured to obtain corrected prediction probability information of the non-punctuation label according to the probability of the non-punctuation label appearing after each character in the corrected text.

[0041] As an optional implementation of the embodiment of the present disclosure, the probability of the punctuation label appearing after each character in the transcription text and the probability of the non-punctuation label appearing after each character in the transcription text are negatively correlated.

[0042] As an optional implementation of the embodiment of the present disclosure, the character probability correction unit is specifically configured to:

[0043] According to a preset correction mode, the probability of the non-punctuation label appearing after each character in the transcription text, the third normalized position weight, and the audio intervention adjustment parameter are calculated to obtain the probability of the non-punctuation label appearing after each character in the corrected text. The preset correction mode includes a first preset correction mode and a second preset correction mode.

[0044] As an optional implementation of the embodiment of the present disclosure, when the preset correction mode is the first preset correction mode, the calculation of the probability of the non-punctuation label appearing after each character in the transcription text, the third normalized position weight, and the audio intervention adjustment parameter according to the preset correction mode to obtain the probability of the non-punctuation label appearing after each character in the corrected text includes:

[0045] The probability of the non-punctuation label appearing after each character in the corrected text is determined according to a sum of the probability of the non-punctuation label appearing after each character in the transcription text and a product of the audio intervention adjustment parameter and the third normalized position weight.

[0046] As an optional implementation of the embodiment of the present disclosure, when the preset correction manner is the second preset correction manner, the calculating, according to the preset correction manner, of the probability of the non-punctuation label appearing after each character in the transcribed text, the third normalized position weight, and the audio intervention adjustment parameter to obtain the probability of the non-punctuation label appearing after each corrected character includes:

[0047] determining the probability of the non-punctuation label appearing after each corrected character according to the product of the probability of the non-punctuation label appearing after each character in the transcribed text, the audio intervention adjustment parameter, and the third normalized position weight.

[0048] As an optional implementation of the embodiment of the present disclosure, when the preset correction manner is the first preset correction manner, the weight obtaining unit is specifically configured to:

[0049] obtain a first normalized position weight according to the total number of empty characters in the first text information, the average number of empty characters in the first text information, and the number of empty characters after each character in the first text information;

[0050] obtain a second normalized position weight according to the total number of empty characters in the second text information, the average number of empty characters in the second text information, and the number of empty characters after each character in the second text information;

[0051] obtain a third normalized position weight according to the sum of the first normalized position weight and the second normalized position weight.

[0052] As an optional implementation of the embodiment of the present disclosure, when the preset correction manner is the second preset correction manner, the weight obtaining unit is specifically configured to:

[0053] obtain a first normalized position weight according to the maximum number of empty characters after each of the N characters in the first text information, the number of empty characters after each character in the first text information, and a scaling factor;

[0054] obtain a second normalized position weight according to the maximum number of empty characters after each of the N characters in the second text information, the number of empty characters after each character in the second text information, and a scaling factor;

[0055] obtain a third normalized position weight according to the sum of the first normalized position weight and the second normalized position weight.

[0056] wherein N is an integer greater than or equal to 1.

[0057] In a third aspect, a speech recognition device is provided, comprising a memory and a processor, the memory storing a computer program, and the processor implementing the punctuation prediction method according to the first aspect when executing the computer program.

[0058] In a fourth aspect, a computer-readable storage medium is provided, comprising a computer program stored thereon, and the computer program, when executed by a processor, implements the punctuation prediction method according to the first aspect.

[0059] The technical solution provided by the embodiments of the present disclosure has the following advantages compared with the prior art: the punctuation prediction model is used to obtain the probability of a punctuation tag appearing after each character in a transcription text and the probability of a non-punctuation tag appearing after each character in the transcription text, the first text information corresponding to the original audio and the second text information corresponding to the original audio are obtained, the probability of a non-punctuation tag appearing after each character in the transcription text is corrected according to the first text information corresponding to the original audio and the second text information corresponding to the original audio, and the corrected prediction probability information of the non-punctuation tag is obtained. The transcription text is a text sequence obtained by performing speech recognition processing on the original audio, the punctuation prediction model is used to output the probability of a punctuation tag appearing after each character in the transcription text and the probability of a non-punctuation tag, the first text information is text information obtained by performing audio truncation processing on the original audio, the first text information includes speech feature information and non-speech feature information corresponding to the original audio, and the second text information is text information obtained by performing decoding processing on the original audio, the second text information includes transcription character and transcription space character corresponding to the original audio. By using the non-speech feature information corresponding to the original audio and the character and space character corresponding to the original audio, the probability of a non-punctuation tag appearing after each character in the original audio is corrected, not only based on the context meaning of the text after speech recognition transcription, but also combining the two kinds of audio information of the first text information and the second text information to predict the probability of a non-punctuation tag appearing after each character, and then obtain the corrected prediction probability information of the non-punctuation tag. Since the higher the probability of a non-punctuation tag appearing after each character, the lower the probability of a punctuation tag appearing, the accuracy and rationality of punctuation prediction can be improved by obtaining the corrected prediction probability information of the non-punctuation tag, and the readability of the transcription text is enhanced and the efficiency of speech interaction is improved. BRIEF DESCRIPTION OF DRAWINGS

[0060] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, further serve to explain the principles of the present disclosure.

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without any creative effort.

[0062] Figure 1A A schematic diagram of a label prediction process in the prior art;

[0063] Figure 1B A schematic diagram of an application scenario of a label prediction process in the embodiments of the present disclosure;

[0064] Figure 2A A hardware configuration block diagram of a speech recognition device according to one or more embodiments of the present disclosure;

[0065] Figure 2B A software configuration block diagram of a speech recognition device according to one or more embodiments of the present disclosure;

[0066] Figure 2C An icon control interface display schematic diagram of an application program included in a speech recognition device according to one or more embodiments of the present disclosure;

[0067] Figure 3A One of the flow schematic diagrams of a punctuation prediction method provided in the embodiments of the present disclosure;

[0068] Figure 3B An automatic speech recognition transcription result schematic diagram of an audio signal provided in the embodiments of the present disclosure;

[0069] Figure 3C A result schematic diagram of an audio truncation processing on an original audio provided in the embodiments of the present disclosure;

[0070] Figure 3D A result schematic diagram of a decoding processing on an original audio provided in the embodiments of the present disclosure;

[0071] Figure 4 The second flow schematic diagram of a punctuation prediction method provided in the embodiments of the present disclosure;

[0072] Figure 5A The third flow schematic diagram of a punctuation prediction method provided in the embodiments of the present disclosure;

[0073] Figure 5B Another result schematic diagram of an audio truncation processing on an original audio provided in the embodiments of the present disclosure;

[0074] Figure 5CA schematic diagram of audio fusion correction after audio truncation processing and decoding processing of original audio is provided for the embodiments of the present disclosure.

[0075] Figure 6 A fourth schematic diagram of a flow of a punctuation prediction method is provided for the embodiments of the present disclosure.

[0076] Figure 7 A schematic diagram of a structure of a label prediction device is provided for the embodiments of the present disclosure.

[0077] Figure 8 A schematic diagram of a structure of a speech recognition device is provided for the embodiments of the present disclosure. DETAILED DESCRIPTION

[0078] In order to more clearly understand the above-mentioned purposes, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0079] In the following description, many specific details are set forth in order to fully understand the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some of the embodiments of the present disclosure, not all the embodiments.

[0080] The terms "first", "second", "third", and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit the specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchanged under appropriate circumstances.

[0081] The terms "include" and "have" and any variations thereof are intended to cover but not exclusive inclusion, for example, a product or device including a series of components does not have to be limited to all the components clearly listed, but can include other components not clearly listed or inherent to these products or devices.

[0082] With the rapid development of intelligent technologies and the increasing popularity of smart devices, speech recognition technology is playing an increasingly important role in various fields such as home appliances, automotive electronics, and consumer electronics. Speech recognition technology is the technology that enables machines to convert speech signals into corresponding text or commands through recognition and understanding. In human-computer interaction scenarios, speech recognition plays a crucial role in natural language understanding and generation. The accuracy of transcribed text is the foundation and bottleneck of downstream tasks. Therefore, much exploratory work remains to be done on how to improve the accuracy of speech recognition. Furthermore, predicting punctuation based on transcribed text is another key task in semantic understanding and interaction. Correct punctuation greatly assists semantic understanding. Ambiguous or mismatched punctuation can even be misleading, thus affecting the overall process of voice interaction.

[0083] Currently, punctuation prediction in speech recognition is mostly based on the contextual semantics of transcribed text to output the punctuation of the entire sentence. Figure 1A This is a schematic diagram of punctuation prediction methods in existing speech recognition technology. For example... Figure 1A As shown, its main implementation process is as follows: the original audio is processed to obtain transcribed text, the transcribed text is input into the punctuation prediction model, and the probability of outputting punctuation at each position is output. However, in this method, when the transcribed text is too short and the available text information is insufficient, the probability of mis-punctuation or omission is extremely high, thus resulting in low accuracy of punctuation prediction.

[0084] To address the shortcomings of the aforementioned methods, this embodiment first obtains the probability of punctuation tags and non-punctuation tags appearing after each character in the transcribed text based on a punctuation prediction model. It then obtains first text information and second text information corresponding to the original audio. Based on the first and second text information, the probability of non-punctuation tags appearing after each character in the transcribed text is corrected, resulting in the corrected predicted probability information for non-punctuation tags. Here, the transcribed text is a text sequence obtained by performing speech recognition processing on the original audio. The punctuation prediction model outputs the probability of non-punctuation tags appearing after each character in the transcribed text. The first text information is text information obtained by audio truncation processing on the original audio, including speech feature information and non-speech feature information corresponding to the original audio. The second text information is text information obtained by decoding the original audio, including transcribed text characters and transcribed empty characters corresponding to the original audio. By using the non-speech feature information and the corresponding text characters and blank characters in the original audio, the probability of non-punctuation marks appearing after each character in the original audio is corrected. This is based not only on the contextual semantics of the text transcribed from speech recognition, but also on the prediction of the probability of non-punctuation marks appearing after each character by combining the first and second text information of the audio. This results in the obtained predicted probability information of corrected non-punctuation marks. Since the higher the probability of non-punctuation marks appearing after each character, the lower the probability of punctuation marks appearing, obtaining the predicted probability information of corrected non-punctuation marks can improve the accuracy and rationality of punctuation prediction, enhance the readability of the transcribed text, and improve the efficiency of voice interaction.

[0085] For example, such as Figure 1B As shown, Figure 1B This is a schematic diagram illustrating an application scenario of a semantic recognition punctuation prediction process for an intelligent device provided in this embodiment of the disclosure. Figure 1B In this context, the process of speech recognition punctuation prediction can be applied to voice interaction scenarios between users and smart home devices. For example, the smart devices in this scenario could be smart device 100 (…). Figure 1B Examples include smart refrigerators and smart devices 101. Figure 1B Examples include smart washing machines and smart devices 102. Figure 1B For example, in smart devices with voice recognition capabilities, such as smart display devices, users need to issue voice commands to control the smart devices in the scenario. When the smart device receives the voice command, it performs voice recognition on the voice command, and predicts the text punctuation based on the transcribed text. This makes it easier for the smart device to display the text according to the punctuation prediction results, which helps to improve the readability of the transcribed text and further improves the efficiency of voice interaction.

[0086] The punctuation prediction method provided in this disclosure can be implemented based on a computer device, or a functional module or functional entity within a computer device.

[0087] The computer equipment can be a personal computer (PC), server, mobile phone, tablet computer, laptop computer, mainframe computer, etc., and this disclosure does not specifically limit it.

[0088] For example, Figure 2A This is a hardware configuration block diagram of a computer device according to one or more embodiments of the present disclosure. Figure 2A As shown, the computer device includes at least one of the following: a tuner / demodulator 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface 280. The controller 250 includes a central processing unit, a video processor, an audio processor, a graphics processor, RAM, ROM, and a first to nth interface for input / output. The display 260 can be at least one of a liquid crystal display, an OLED display, a touch display, and a projection display, and can also be a projection device and a projection screen. The tuner / demodulator 210 receives broadcast television signals via wired or wireless means, and demodulates audio and video signals, such as EPG audio and video data signals, from multiple wireless or wired broadcast television signals. The communicator 220 is a component used to communicate with external devices or servers according to various communication protocol types. For example, the communicator may include at least one of a Wi-Fi module, a Bluetooth module, a wired Ethernet module, other network communication protocol chips or near-field communication protocol chips, and an infrared receiver. The computer equipment can establish the transmission and reception of control signals and data signals with the server or local control equipment through the communicator 220. The detector 230 is used to collect signals from the external environment or to interact with the outside world. The controller 250 and the tuner / demodulator 210 can be located in different separate devices; that is, the tuner / demodulator 210 can also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0089] In some embodiments, the controller 250 controls the operation of the computer device and responds to user operations through various software control programs stored in memory. The controller 250 controls the overall operation of the computer device. The user can input commands through a graphical user interface (GUI) displayed on the monitor 260, and the user input interface receives the user input commands through the GUI. Alternatively, the user can input commands by entering specific sounds or gestures, and the user input interface receives the user input commands by recognizing the sounds or gestures through sensors.

[0090] Figure 2B As a software configuration diagram of a computer device according to one or more embodiments of the present disclosure, as shown in FIG. 1, the system is divided into four layers from top to bottom, namely, an Applications layer (referred to as an "application layer" for short), an Application Framework layer (referred to as a "framework layer" for short), an Android runtime and system library layer (referred to as a "system runtime library layer" for short), and a kernel layer. Figure 2B

[0091] Figure 2C As an icon control interface display diagram of an application included in a smart device (mainly a smart play device, such as a smart television, a digital theater system, or a video server) according to one or more embodiments of the present disclosure, as shown in FIG. 2, the application layer includes at least one application, which can display a corresponding icon control in a display, such as a live television application icon control, a video on demand (VOD) application icon control, a media center application icon control, an application center icon control, a game application icon control, and the like. The live television application can provide live television through different signal sources. The VOD application can provide videos from different storage sources. Unlike the live television application, the VOD provides video display from certain storage sources. The media center application can provide various multimedia content play applications. The application center can provide storage of various applications. Figure 2C

[0092] To further illustrate the present solution, the following will be described in an exemplary manner in combination with Figure 3A , Figure 4 , Figure 5A , Figure 6 It should be understood that although each step in the flowcharts of 3A, Figure 4 , Figure 5A , Figure 6 is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, 3A, Figure 4 , Figure 5A , Figure 6 ​​At least one of the steps in the method can include a plurality of sub-steps or a plurality of stages, which are not necessarily performed at the same time, but can be performed at different times, and the order of the execution of the sub-steps or stages is not necessarily sequential, but can be performed alternately or alternately with at least one of the other steps or the sub-steps or stages of the other steps. The implementation of the punctuation prediction method provided in the embodiments of the present disclosure is used as a reference.

[0093] As shown in Figure 3A , the method specifically includes the following steps:

[0094] S31, obtaining a probability of a punctuation tag appearing after each character in the speech recognition transcription text based on a punctuation prediction model.

[0095] The speech recognition transcription text is a text sequence obtained by performing speech recognition processing on the original audio. Specifically, the speech recognition transcription text is a text sequence obtained by processing the original audio signal based on ASR (Automatic Speech Recognition) technology. As shown in Figure 3B , the input of speech recognition is generally a time-domain speech signal represented by a series of vectors, and the output is text.

[0096] The punctuation prediction model is used to output the probability of a punctuation tag appearing after each character in the speech recognition transcription text and the probability of a non-punctuation tag. It should be noted that punctuation tags can be divided into four categories, namely: comma, period, question mark, and exclamation mark. A non-punctuation tag can be understood as a literal character.

[0097] Specifically, the network structure of the punctuation prediction model is divided into two parts, which are a pre-trained language model and a bidirectional LSTM (Long Short-Term Memory) model. For text punctuation prediction, the context information of the text is important, so a bidirectional LSTM model is added after the pre-trained language model to obtain the global dependency information of the text. The pre-trained model of the punctuation prediction model uses a transformer-based structure, and the multi-head attention can better encode the upper and lower information of the text.

[0098] The input of the punctuation prediction model is the transcription text of speech recognition, and the output is the probability of a punctuation tag appearing after each character and the probability of a non-punctuation tag. For example, the input of the punctuation prediction model can be the transcription text "today's weather is how", "hello, I am the technical support of A company" and the like.

[0099] The punctuation probability is corrected in combination with the acquired VAD audio information. In combination with audio signals such as pauses and speech rates, weight is added to positions where punctuation should be output in the transcribed text but no punctuation is output in the punctuation prediction model, and punctuation is output.

[0100] S32, acquiring first text information corresponding to the original audio and second text information corresponding to the original audio.

[0101] The first text information is text information obtained by performing audio truncation processing on an audio signal of the original audio, and includes speech feature information and non-speech feature information corresponding to the original audio. The second text information is text information obtained by performing decoding processing on the original audio, and includes transcribed text characters and transcribed null characters corresponding to the original audio.

[0102] Specifically, the audio truncation processing is performed by a VAD (Voice Activity Detection) technology. Voice activity detection, also known as speech endpoint detection, is generally used to identify the presence and absence of speech in an audio signal. Typically, a VAD algorithm divides the audio signal into a speech portion, a non-speech portion, and a silent portion. For example, when speech is detected, 1 is output, otherwise, 0 is output. In this embodiment, the first text information is the speech feature information and the non-speech feature information output by decoding in the original audio signal. For example, a schematic diagram of the information contained in the first text information is shown in Figure 3C .

[0103] The decoding processing is a decoding operation based on automatic speech recognition technology. During the decoding processing, if there is no text character output at the current time, a is output, which may represent silence or other conditions. A schematic diagram of the information contained in the second text information is shown in Figure 3D . The speech decoding uses a prefix-beam-search algorithm to search for N paths and finally obtains the optimal result. For example, the transcribed text is "that's good you off work no", and the corresponding intermediate transcribed text may include but is not limited to the following ways: (1) that 's good you off work no (2) that 's good you off work no (3) That OK You Down Class No. Finally get a best result: for example, That Just OK You Down Class No

[0104] S33、According to the first text information and the second text information, the probability of each character in the transcription text after the non-punctuation label is corrected, and the prediction probability information of the corrected non-punctuation label is obtained.

[0105] Specifically, after obtaining the probability of each character in the transcription text after the punctuation label and the probability of each character in the transcription text after the non-punctuation label, the first text information and the second text information are combined. Two kinds of audio information are used to correct the probability of each character in the transcription text after the non-punctuation label, so as to obtain the corrected punctuation information. One way that can be implemented is to input the transcription text into the punctuation prediction model, output the probability of each position appearing five labels respectively, and combine the obtained VAD audio information and the ASR continuous transcription audio information. The position where the punctuation is not output by the prediction model is increased in weight, and the punctuation is output.

[0106] Exemplarily, the original audio collected in real time is decoded into transcription text by ASR, for example, the transcription text is: “That OK you down class no”. The transcription text is input into the punctuation prediction model, and the output is “0000000 question mark”. The original audio is processed to obtain the first text information: “ That Just OK You Down Class No ”, the second text information is: “11110001110001111000001111000110001111000111110000”, and finally the result after the intervention of the above two kinds of audio information is obtained: “That OK, you down class no?”.

[0107] In some embodiments, the probability of each character in the transcription text after the punctuation label and the probability of each character in the transcription text after the non-punctuation label are negatively correlated.

[0108] Specifically, when the probability of a punctuation tag appearing after a certain character in the transcribed text is relatively high, correspondingly, the probability of a non-punctuation tag appearing after that character is relatively low; when the probability of a punctuation tag appearing after a certain character in the transcribed text is relatively low, correspondingly, the probability of a non-punctuation tag appearing after that character is relatively high. Exemplarily, assuming the transcribed text is "Hello", the probability of a punctuation tag appearing after the character "你" is relatively low, while the probability of a non-punctuation tag appearing after it is relatively high.

[0109] In the embodiments of the present disclosure, by using the non-speech feature information corresponding to the original audio, the text characters and null characters corresponding to the original audio, the probability of a non-punctuation tag appearing after each character in the original audio is corrected. Not only based on the context semantics of the text after speech recognition transcription, but also by combining the two types of audio information, the first text information and the second text information, to predict the probability of a non-punctuation tag appearing after each character, and then obtain the predicted probability information of the corrected non-punctuation tag. Also, since the higher the probability of a non-punctuation tag appearing after each character, the lower the probability of a punctuation tag appearing, so by obtaining the predicted probability information of the corrected non-punctuation tag, the accuracy and rationality of punctuation prediction can be improved, while enhancing the readability of the transcribed text and improving the efficiency of voice interaction.

[0110] Figure 4 It is a flowchart showing another punctuation prediction method provided by the embodiments of the present disclosure. This embodiment is further extended and optimized based on Figure 3A . Optionally, this embodiment mainly describes the process of step S33 (correcting the probability of a non-punctuation tag appearing after each character in the transcribed text according to the first text information and the second text information, and obtaining the predicted probability information of the corrected non-punctuation tag).

[0111] S431. Obtain the first normalized position weight, the second normalized position weight, and the third normalized position weight.

[0112] Among them, the first normalized position weight is the normalized weight of the number of null characters after each character in the first text information, the second normalized position weight is the normalized weight of the number of null characters after each character in the second text information, and the third normalized position weight is the sum of the first normalized position weight and the second normalized position weight.

[0113] S432. According to the probability of a non-punctuation tag appearing after each character in the transcribed text, the third normalized position weight, and the audio intervention adjustment parameter, obtain the probability of a non-punctuation tag appearing after each corrected character.

[0114] The audio intervention adjustment parameter can adjust the intervention degree of the audio information for different scenes to achieve optimization in specific scenes. For example, the audio intervention adjustment parameter can be 0.5, 0.2, 0.3, or other reasonable values, which are not limited here.

[0115] After determining the probability of a non-punctuation label appearing after each character in the transcribed text, the third normalized position weight, and the audio intervention adjustment parameter, the probability of a non-punctuation label appearing after each character in the corrected transcribed text can be obtained in the following ways.

[0116] In some embodiments, when the preset correction method is the first preset correction method, the probability of a non-punctuation label appearing after each character in the transcribed text, the third normalized position weight, and the audio intervention adjustment parameter are calculated according to the preset correction method to obtain the probability of a non-punctuation label appearing after each character after correction, which can be realized by the following way:

[0117] According to the probability of a non-punctuation label appearing after each character in the transcribed text and the sum of the product of the audio intervention adjustment parameter and the third normalized position weight, the probability of a non-punctuation label appearing after each character after correction is determined.

[0118] Specifically, the probability of a non-punctuation label appearing after each character after correction can be calculated by the following formula (1):

[0119] L Oi =l Oi +β*δ i Formula (1)

[0120] wherein L Oi represents the probability of a non-punctuation label appearing after each character after correction, l Oi represents the probability of a non-punctuation label appearing after each character output by the punctuation prediction model, β represents the audio intervention adjustment parameter, and δ i represents the third normalized position weight.

[0121] For example, when the probability of a non-punctuation label appearing after each character output by the punctuation prediction model is 0.2, the audio intervention adjustment parameter is 0.5, and the third normalized position weight is 0.4, the probability of a non-punctuation label appearing after each character after correction L Oi = 0.2 + 0.5 * 0.4 = 0.4.

[0122] In some embodiments, when the preset correction manner is the second preset correction manner, the probability of the non-punctuation label appearing after each character in the corrected transcription text is calculated according to the preset correction manner, the probability of the non-punctuation label appearing after each character in the transcription text, the third normalized position weight, and the audio intervention adjustment parameter. The probability of the non-punctuation label appearing after each character in the corrected transcription text can be obtained by the following method:

[0123] The probability of the non-punctuation label appearing after each character in the corrected transcription text is determined according to the product of the probability of the non-punctuation label appearing after each character in the transcription text, the audio intervention adjustment parameter, and the third normalized position weight.

[0124] Specifically, the probability of the non-punctuation label appearing after each character in the corrected transcription text can be calculated by the following formula (2):

[0125] L Oi =l Oi *β*δ i Formula (2)

[0126] Wherein, L Oi represents the probability of the non-punctuation label appearing after each character in the corrected transcription text, l Oi represents the probability of the non-punctuation label appearing after each character output by the punctuation prediction model, β represents the audio intervention adjustment parameter, and δ i represents the third normalized position weight.

[0127] For example, when the probability of the non-punctuation label appearing after each character output by the punctuation prediction model is 0.2, the audio intervention adjustment parameter is 0.5, and the third normalized position weight is 0.4, the probability of the non-punctuation label appearing after each character in the corrected transcription text L Oi = 0.2 * 0.5 * 0.4 = 0.04.

[0128] S433, the probability of the punctuation label appearing after each position in the corrected transcription text is determined according to the probability of the non-punctuation label appearing after each character in the corrected transcription text.

[0129] Specifically, since the probability of the non-punctuation label appearing after each character in the transcription text is negatively correlated with the probability of the punctuation label appearing after each character, that is, if the probability of the non-punctuation label appearing after each character in the transcription text increases, the probability of the punctuation label appearing after each character in the transcription text decreases. If the probability of the non-punctuation label appearing after each character in the transcription text decreases, the probability of the punctuation label appearing after each character in the transcription text increases. Through the intervention of audio information, when there are more empty characters, the third normalized position weight δ i is reduced, so that the probability of the punctuation label appearing after each position in the corrected transcription text increases. When there are fewer empty characters, the third normalized position weight δ i, so as to reduce the probability of punctuation tags appearing at each position after correction, thereby realizing the intervention of audio information on punctuation output.

[0130] Figure 5A It is a schematic flowchart of another punctuation prediction method provided by an embodiment of the present disclosure. This embodiment is further extended and optimized on the basis of Figure 4 . Optionally, in this embodiment, when the correction method is an additive intervention method, the process of step S431 (obtaining the first normalized position weight, the second normalized position weight, and the third normalized position weight) is described.

[0131] When the correction method is an additive intervention method, the first normalized position weight, the second normalized position weight, and the third normalized position weight can be obtained through the following method:

[0132] S5311. Obtain the first normalized position weight according to the total number of empty characters in the first text information, the average number of empty characters in the first text information, and the number of empty characters after each character in the first text information.

[0133] Specifically, the first normalized position weight is calculated by the following formula (3):

[0134]

[0135] where α i represents the first normalized position weight, represents the total number of empty characters in the first text information, represents the average number of empty characters in the first text information, represents the number of empty characters after each character in the first text information.

[0136] In order to increase the probability of punctuation appearance at positions where there are more occurrences in the automatic speech recognition transcription result , the automatic speech recognition transcription result can be fused by including but not limited to the following methods.

[0137] Exemplarily, the transcribed text is: "That's good. Have you got off work?" There are 7 Chinese characters in this transcribed text, and there are 7 positions in total. Since punctuation generally does not appear at the beginning and no intervention is required, the before the first Chinese character is ignored. The position after the first character is marked as position 1, and then the number of at each position is counted. For example, the optimal result of speech decoding is " That is good You get off No in The position encoding can be represented as: P a = 2162211. The sum in the transcription result is recorded as The number of spaces in each position is recorded as The number of spaces in each position is recorded as The number of spaces in each position is recorded as The number of spaces in each position is recorded as

[0138] The effect of normalization is to eliminate the influence of the speaker's speed; it is beneficial to more accurately control the probability of non-punctuation labels, and then effectively intervene in the probability of punctuation labels.

[0139] S5312, according to the sum of the number of spaces in the second text information, the average number of spaces in the second text information, the number of spaces after each character in the second text information, a second normalized position weight is obtained.

[0140] Specifically, the second normalized position weight is calculated by the following formula (4):

[0141]

[0142] Wherein, ε i represents the second normalized position weight, θ N represents the sum of the number of spaces in the second text information, θ avg represents the average number of spaces in the second text information, and θ i represents the number of spaces after each character in the second text information.

[0143] The principle of VAD processing is that when there is sound, the label of the frame is set to "1", and when there is no sound, the label of the frame is set to "0". Based on the position of the punctuation, it is a fact that the speaker pauses more. Start counting after the character is converted from ASR, and count the positions where the VAD output label is "0". Then, according to the position alignment of the ASR transcription text.

[0144] In order to realize the position where "0" appears more in the VAD result, and improve the probability of punctuation appearing, the VAD result can be fused by using the following methods, but not limited to. As shown in Figure 5B , the position where "0" appears more in Figure 5B , the probability of punctuation appearing is greater. The position vector obtained after counting the number of "0" in each position is: P v= 3353334. The sum of "0" in the VAD result is 3+3+5+3+3+3+4=24, θ N = 24, so θ avg = 24 / 7, θ1=3, θ2=3, θ3=5, θ4=3, θ5=3, θ6=3, θ7=4.

[0145] Since the output of the VAD will have various interference, and is not like the above idealized result, for this, the labels of the VAD are combined with the same items in combination with the ASR transcription text, so as to realize the alignment of the ASR transcription result and the VAD label result. Referring to Figure 5C , Figure 5C the bold label "1" in the above formula (3) may be caused by errors in transcoding or other situations, in combination with the ASR transcription result, when the ASR transcription result is , it is corrected to "0".

[0146] S5313, according to the sum of the first normalized position weight and the second normalized position weight, a third normalized position weight is obtained.

[0147] Specifically, the third normalized position weight is calculated by the following formula (5):

[0148] δ i = α i + ε i formula (5)

[0149] Wherein, α i represents the first normalized position weight, and ε i represents the second normalized position weight.

[0150] In order to better fuse the ASR transcription result and the VAD result information, the formula (5) is used to additively fuse the audio information weight of the ASR transcription and the audio information weight of the VAD result, and δ i is obtained.

[0151] Figure 6 is a flowchart of another punctuation prediction method provided by the embodiment of the disclosure. The embodiment is further extended and optimized on the basis of Figure 4 . Optionally, the embodiment is to explain the process of step S431 (obtaining the first normalized position weight, the second normalized position weight, and the third normalized position weight) when the correction mode is the multiplicative intervention mode.

[0152] S6311, according to the maximum number of space characters after each character in the N characters of the first text information, the number of space characters after each character in the first text information, and a scaling factor, a first normalized position weight is obtained.

[0153] Specifically, the first normalized position weight is calculated by the following formula (6):

[0154]

[0155] wherein, α i represents the first normalized position weight, max 1N represents the maximum number of empty characters after each character in the N characters of the first text information, represents the number of empty characters after each character in the first text information, and γ represents a scaling factor.

[0156] For example, γ is a scaling factor, and its optimal value can be 1.5 or other reasonable values, which are not specifically limited here.

[0157] S6312, according to the maximum number of empty characters after each character in the N characters of the second text information, the number of empty characters after each character in the second text information, and the scaling factor, the second normalized position weight is obtained.

[0158] Specifically, the first normalized position weight is calculated by the following formula (7):

[0159]

[0160] wherein, ε i represents the second normalized position weight, max 2N represents the maximum number of empty characters after each character in the N characters of the second text information, θ i represents the number of empty characters after each character in the second text information, and γ represents a scaling factor.

[0161] S6313, according to the sum of the first normalized position weight and the second normalized position weight, the third normalized position weight is obtained.

[0162] wherein, N is an integer greater than or equal to 1.

[0163] Specifically, the third normalized position weight is calculated by the above formula (5):

[0164] δ i = α i + ε i Formula (5)

[0165] wherein, α i represents the first normalized position weight, and ε i represents the second normalized position weight.

[0166] In the embodiments of the present disclosure, the probability of non-punctuation appearing after each character in the original audio is corrected by the non-speech feature information corresponding to the original audio and the literal character and the space character corresponding to the original audio. The probability of non-punctuation appearing after each character is predicted based on the context meaning of the text transcribed after speech recognition, combined with the first text information and the second text information, that is, two kinds of audio information. Then, the prediction probability information of the corrected non-punctuation label is obtained. Since the higher the probability of non-punctuation appearing after each character, the lower the probability of punctuation appearing, the prediction accuracy and rationality of punctuation can be improved by obtaining the prediction probability information of the corrected non-punctuation label, and the readability of the transcribed text is improved, and the efficiency of speech interaction is improved.

[0167] Figure 7 A structural diagram of a semantic understanding device is provided in the embodiments of the present disclosure. The device is configured in a smart device, and can implement the punctuation prediction method described in any embodiment of the present disclosure. The device 700 specifically includes the following:

[0168] The punctuation probability acquisition module 710 is configured to acquire the probability of punctuation label and the probability of non-punctuation label appearing after each character in the transcribed text based on a punctuation prediction model. The transcribed text is a text sequence obtained by performing speech recognition on the original audio, and the punctuation prediction model is used to output the probability of punctuation label and the probability of non-punctuation label appearing after each character in the transcribed text.

[0169] The audio text acquisition module 720 is configured to acquire the first text information corresponding to the original audio and the second text information corresponding to the original audio. The first text information is text information obtained by performing audio truncation processing on the original audio, and the first text information includes speech feature information and non-speech feature information corresponding to the original audio. The second text information is text information obtained by performing decoding processing on the original audio, and the second text information includes transcribed literal characters and transcribed space characters corresponding to the original audio.

[0170] The punctuation probability correction module 730 is configured to correct the probability of non-punctuation label appearing after each character in the transcribed text according to the first text information and the second text information, and acquire prediction probability information of the corrected non-punctuation label.

[0171] As an optional implementation of the present disclosure, the punctuation probability correction module 730 includes:

[0172] The weight acquisition unit is configured to acquire a first normalized position weight, a second normalized position weight, and a third normalized position weight.

[0173] The first normalized position weight is a normalized weight of the number of empty characters after each character in the first text information, the second normalized position weight is a normalized weight of the number of empty characters after each character in the second text information, and the third normalized position weight is a sum of the first normalized position weight and the second normalized position weight.

[0174] The character probability correction unit is configured to obtain a probability of a non-punctuation label appearing after each character in the corrected transcription text according to a probability of a non-punctuation label appearing after each character in the transcription text, the third normalized position weight, and an audio intervention adjustment parameter.

[0175] The punctuation probability correction unit is configured to obtain the corrected punctuation probability information according to the probability of the non-punctuation label appearing after each character.

[0176] As an optional implementation of the embodiment of the present disclosure, the probability of a punctuation label appearing after each character in the transcription text and the probability of a non-punctuation label appearing after each character in the transcription text are negatively correlated.

[0177] As an optional implementation of the embodiment of the present disclosure, the character probability correction unit is specifically configured to:

[0178] According to a preset correction mode, the probability of a non-punctuation label appearing after each character in the transcription text, the third normalized position weight, and the audio intervention adjustment parameter are calculated to obtain a probability of a non-punctuation label appearing after each character in the corrected transcription text. The preset correction mode includes a first preset correction mode and a second preset correction mode.

[0179] As an optional implementation of the embodiment of the present disclosure, when the preset correction mode is the first preset correction mode, the calculation of the probability of a non-punctuation label appearing after each character in the transcription text, the third normalized position weight, and the audio intervention adjustment parameter according to the preset correction mode to obtain the probability of a non-punctuation label appearing after each character in the corrected transcription text includes:

[0180] The sum of the probability of a non-punctuation label appearing after each character in the transcription text and the product of the audio intervention adjustment parameter and the third normalized position weight is determined as the probability of a non-punctuation label appearing after each character in the corrected transcription text.

[0181] As an optional implementation of the embodiment of the present disclosure, when the preset correction mode is the second preset correction mode, the calculation of the probability of a non-punctuation label appearing after each character in the transcription text, the third normalized position weight, and the audio intervention adjustment parameter according to the preset correction mode to obtain the probability of a non-punctuation label appearing after each character in the corrected transcription text includes:

[0182] determine a probability of a non-punctuation tag following each character in the transcribed text according to a product of the probability of a non-punctuation tag following each character in the transcribed text, the audio intervention adjustment parameter, and the third normalized position weight.

[0183] As an optional implementation of the embodiment of the present disclosure, when the correction mode is the additive intervention mode, the weight obtaining unit is specifically configured to:

[0184] obtain a first normalized position weight according to a sum of the number of empty characters in the first text information, an average of the number of empty characters in the first text information, and the number of empty characters following each character in the first text information;

[0185] obtain a second normalized position weight according to a sum of the number of empty characters in the second text information, an average of the number of empty characters in the second text information, and the number of empty characters following each character in the second text information;

[0186] obtain a third normalized position weight according to a sum of the first normalized position weight and the second normalized position weight.

[0187] As an optional implementation of the embodiment of the present disclosure, when the correction mode is the multiplicative intervention mode, the weight obtaining unit is specifically configured to:

[0188] obtain a first normalized position weight according to a maximum number of empty characters following each of the N characters in the first text information, the number of empty characters following each character in the first text information, and a scaling factor;

[0189] obtain a second normalized position weight according to a maximum number of empty characters following each of the N characters in the second text information, the number of empty characters following each character in the second text information, and a scaling factor;

[0190] obtain a third normalized position weight according to a sum of the first normalized position weight and the second normalized position weight.

[0191] wherein N is an integer greater than or equal to 1.

[0192] In the embodiment of the present disclosure, the probability of a punctuation label and the probability of a non-punctuation label after each character in the transcription text are obtained based on a punctuation prediction model, the first text information corresponding to the original audio and the second text information corresponding to the original audio are obtained, the probability of the non-punctuation label after each character in the transcription text is corrected according to the first text information corresponding to the original audio and the second text information corresponding to the original audio, and the corrected prediction probability information of the non-punctuation label is obtained. The transcription text is a text sequence obtained by performing speech recognition on the original audio, the punctuation prediction model is used to output the probability of the non-punctuation label after each character in the transcription text, the first text information is text information obtained by performing audio truncation on the original audio, the first text information includes speech feature information and non-speech feature information corresponding to the original audio, and the second text information is text information obtained by decoding the original audio, the second text information includes transcription character and transcription null character corresponding to the original audio. The probability of the non-punctuation label after each character in the original audio is corrected by using the non-speech feature information corresponding to the original audio and the character and null character corresponding to the original audio. The probability of the non-punctuation label after each character is predicted based on the context meaning of the text after speech recognition and in combination with the first text information and the second text information, and then the corrected prediction probability information of the non-punctuation label is obtained. Since the higher the probability of the non-punctuation label after each character is, the lower the probability of the punctuation label is, the accuracy and rationality of punctuation prediction can be improved by obtaining the corrected prediction probability information of the non-punctuation label, and the readability of the transcription text is improved and the efficiency of speech interaction is improved.

[0193] The semantic understanding device provided in the embodiments of the present disclosure can perform the punctuation prediction method provided in any of the embodiments of the present disclosure, has the corresponding function modules and beneficial effects of the execution method, and to avoid repetition, details are not repeated here.

[0194] The embodiment of the present disclosure provides a computer device, comprising: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the punctuation prediction method in any of the embodiments of the present disclosure.

[0195] Figure 8 is a structural schematic diagram of a computer device provided by the embodiments of the present disclosure. As shown in Figure 8 , the computer device includes a processor 810 and a storage device 820; the number of processors 810 in the computer device can be one or more, Figure 8 , taking a processor 810 as an example; the processor 810 and the storage device 820 in the computer device can be connected through a bus or other means,Figure 8 The bus connection is taken as an example.

[0196] The storage 820, as a kind of computer readable storage medium, can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the punctuation prediction method in the embodiments of the present disclosure. The processor 810 performs various functional applications and data processing of the computer device by running the software programs, instructions and modules stored in the storage 820, that is, realizes the punctuation prediction method provided by the embodiments of the present disclosure.

[0197] The storage 820 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and application programs required by at least one function; the data storage area can store data created according to the use of the terminal and the like. In addition, the storage 820 can include a high-speed random access storage device, and can also include a non-volatile storage device, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some examples, the storage 820 can further include storage devices arranged remotely with respect to the processor 810, which can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0198] The computer device provided by the embodiments of the present disclosure can be used to execute the punctuation prediction method provided by any of the above embodiments, and has the corresponding functions and advantages.

[0199] The embodiments of the present disclosure also provide a storage medium containing computer executable instructions, which realize various processes performed by the method provided by any of the above embodiments when executed by a computer processor, and can achieve the same technical effects. To avoid repetition, details are not repeated here.

[0200] The computer readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0201] The above description has been made in conjunction with specific embodiments for the convenience of explanation. However, the above description in some embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed. Various modifications and variations can be derived from the above teachings. The selection and description of the above embodiments are to better explain the principles and practical applications, so that those skilled in the art can better use the embodiments and adapt various different modified embodiments according to specific use considerations.

Claims

1. A punctuation prediction method, characterized in that, The method includes: The probability of a punctuation mark appearing after each character in the transcribed text is obtained based on a punctuation prediction model; the transcribed text is a text sequence obtained by speech recognition processing of the original audio; the punctuation prediction model is used to output the probability of a punctuation mark appearing after each character in the transcribed text. Obtain first text information corresponding to the original audio and second text information corresponding to the original audio; the first text information is text information obtained by audio truncation processing of the original audio, and the first text information includes speech feature information and non-speech feature information corresponding to the original audio; the second text information is text information obtained by decoding the original audio, and the second text information includes transcribed text characters and transcribed empty characters corresponding to the original audio. Based on the first text information and the second text information, the probability of non-punctuation tags appearing after each character in the transcribed text is corrected, and the predicted probability information of the corrected non-punctuation tags is obtained. The step of correcting the probability of non-punctuation tags appearing after each character in the transcribed text based on the first text information and the second text information, and obtaining the corrected predicted probability information of non-punctuation tags, includes: Obtain the first normalized position weight, the second normalized position weight, and the third normalized position weight; Wherein, the first normalized position weight is the normalized weight of the number of empty characters after each character in the first text information, the second normalized position weight is the normalized weight of the number of empty characters after each character in the second text information, and the third normalized position weight is the sum of the first normalized position weight and the second normalized position weight. Based on the probability of a non-punctuation tag appearing after each character in the transcribed text, the third normalized position weight, and the audio intervention adjustment parameters, the probability of a non-punctuation tag appearing after each character after correction is obtained. Based on the probability of non-punctuation tags appearing after each corrected character, the predicted probability information of the corrected non-punctuation tags is obtained.

2. The method according to claim 1, characterized in that, The method further includes: The probability of a punctuation mark appearing after each character in the transcribed text is negatively correlated with the probability of a non-punctuation mark appearing after each character in the transcribed text.

3. The method according to claim 1, characterized in that, The step of obtaining the probability of a non-punctuation tag appearing after each character in the transcribed text, based on the probability of a non-punctuation tag appearing after each character, the third normalized position weight, and the audio intervention adjustment parameters, includes: According to the preset correction method, the probability of non-punctuation tags appearing after each character in the transcribed text, the third normalized position weight, and the audio intervention adjustment parameters are calculated to obtain the probability of non-punctuation tags appearing after each character after correction; the preset correction method includes: a first preset correction method and a second preset correction method.

4. The method according to claim 3, characterized in that, When the preset correction method is the first preset correction method, the step of calculating the probability of a non-punctuation tag appearing after each character in the transcribed text, the third normalized position weight, and the audio intervention adjustment parameters according to the preset correction method to obtain the probability of a non-punctuation tag appearing after each character after correction includes: The probability of a non-punctuation tag appearing after each character after correction is determined by multiplying the audio intervention adjustment parameter with the third normalized position weight and summing the probability of a non-punctuation tag appearing after each character in the transcribed text.

5. The method according to claim 3, characterized in that, When the preset correction method is the second preset correction method, the step of calculating the probability of a non-punctuation tag appearing after each character in the transcribed text, the third normalized position weight, and the audio intervention adjustment parameters according to the preset correction method to obtain the probability of a non-punctuation tag appearing after each character after correction includes: The probability of a non-punctuation tag appearing after each character in the transcribed text is determined by multiplying the probability of a non-punctuation tag appearing after each character in the transcribed text, the audio intervention adjustment parameters, and the third normalized position weights.

6. The method according to claim 3, characterized in that, When the preset correction method is the first preset correction method, obtaining the first normalized position weight, the second normalized position weight, and the third normalized position weight includes: The first normalized position weight is obtained based on the sum of the number of empty characters in the first text information, the average number of empty characters in the first text information, and the number of empty characters after each character in the first text information. The second normalized position weight is obtained based on the sum of the number of empty characters in the second text information, the average number of empty characters in the second text information, and the number of empty characters after each character in the second text information. The third normalized position weight is obtained by summing the first normalized position weight and the second normalized position weight.

7. The method according to claim 3, characterized in that, When the preset correction method is the second preset correction method, obtaining the first normalized position weight, the second normalized position weight, and the third normalized position weight includes: The first normalized position weight is obtained based on the maximum number of empty characters after each character in the N characters of the first text information, the number of empty characters after each character in the first text information, and the scaling factor. The second normalized position weight is obtained based on the maximum number of empty characters after each character in the N characters of the second text information, the number of empty characters after each character in the second text information, and the scaling factor; The third normalized position weight is obtained by summing the first normalized position weight and the second normalized position weight. Where N is an integer greater than or equal to 1.

8. A punctuation prediction device, characterized in that, The device includes: The punctuation probability acquisition module is used to obtain the probability of a punctuation tag appearing after each character in the transcribed text based on a punctuation prediction model; the transcribed text is a text sequence obtained by speech recognition processing of the original audio, and the punctuation prediction model is used to output the probability of a punctuation tag appearing after each character in the transcribed text and the probability of a non-punctuation tag appearing. The audio text acquisition module is used to acquire first text information corresponding to the original audio and second text information corresponding to the original audio; the first text information is text information obtained by performing audio truncation processing on the original audio, and the first text information includes speech feature information and non-speech feature information corresponding to the original audio; the second text information is text information obtained by decoding processing the original audio, and the second text information includes transcribed text characters and transcribed empty characters corresponding to the original audio. The punctuation probability correction module is used to correct the probability of non-punctuation tags appearing after each character in the transcribed text based on the first text information and the second text information, and to obtain the predicted probability information of the corrected non-punctuation tags. The punctuation probability correction module is specifically used to obtain the first normalized position weight, the second normalized position weight, and the third normalized position weight. Wherein, the first normalized position weight is the normalized weight of the number of empty characters after each character in the first text information, the second normalized position weight is the normalized weight of the number of empty characters after each character in the second text information, and the third normalized position weight is the sum of the first normalized position weight and the second normalized position weight. Based on the probability of a non-punctuation tag appearing after each character in the transcribed text, the third normalized position weight, and the audio intervention adjustment parameters, the probability of a non-punctuation tag appearing after each character after correction is obtained. Based on the probability of non-punctuation tags appearing after each corrected character, the predicted probability information of the corrected non-punctuation tags is obtained.

9. A voice recognition device, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the punctuation prediction method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Punctuation mark adding method and device

    CN106653030A

  • A text punctuation adjustment method and device

    CN109255115A