Speech recognition text punctuation determination method and device, electronic equipment and storage medium

By utilizing the acoustic features of speech data to correct punctuation marks in speech-recognized text, the problems of time-consuming and laborious manual annotation and inaccurate model predictions are solved, enabling more accurate addition and removal of punctuation marks and improving the readability and accuracy of speech-recognized text.

CN121835667APending Publication Date: 2026-04-10DATABAKER (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing speech recognition technologies, manually annotating punctuation marks is time-consuming, labor-intensive, and inefficient. When using natural language processing models to predict punctuation marks, only the semantic information of the text is considered, which may lead to prediction results that do not match the speaker's intended expression.

Method used

By acquiring the acoustic features of the speech data, the second punctuation mark in the initial text is determined, and the first punctuation mark in the initial text is corrected using the second punctuation mark, including adding punctuation marks in missing positions, removing punctuation marks in misjudged positions, and selecting appropriate punctuation marks based on differences in acoustic features.

Benefits of technology

It improves the accuracy of punctuation marks in speech-recognized text, making it more consistent with the user's expected expression, and enhances the readability and accuracy of speech-recognized text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121835667A_ABST
    Figure CN121835667A_ABST
Patent Text Reader

Abstract

The invention provides a speech recognition text punctuation determination method and device, electronic equipment and a storage medium. The speech recognition text punctuation determination method comprises the following steps: acquiring an initial text which corresponds to speech data and is added with a first punctuation mark; acquiring a second acoustic feature of the voice data; determining a second punctuation mark of the initial text according to the second acoustic feature; and correcting the first punctuation mark in the initial text by using the second punctuation mark to obtain a corrected text. According to the scheme, the punctuation marks of the speech recognition text can be determined more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech recognition text processing, in particular to a speech recognition text punctuation determination method, a speech recognition text punctuation determination device, an electronic device, a storage medium and a computer program product. BACKGROUND

[0002] Speech recognition technology aims to automatically and accurately convert human spoken language speech into corresponding text. Moreover, in order to better enhance the readability and expression accuracy of the text, corresponding punctuation marks are usually added to the recognized text.

[0003] After obtaining the text of the speech data, the punctuation marks in the text can be manually annotated. However, the manual annotation method is time-consuming and laborious, and the efficiency is not high. The punctuation marks of the text can also be predicted by using a trained natural language processing model to determine the punctuation marks of the text according to context information, word segmentation, etc. This method only considers the semantic information of the text itself, and the text containing the predicted punctuation marks may not be the content that the speaker wants to express. SUMMARY

[0004] The present application is proposed in consideration of the above problems.

[0005] According to a first aspect of the present application, a speech recognition text punctuation determination method is provided. The method comprises: obtaining an initial text to which first punctuation marks corresponding to speech data have been added; obtaining second acoustic features of the speech data; determining second punctuation marks of the initial text according to the second acoustic features; and correcting the first punctuation marks in the initial text by using the second punctuation marks to obtain a corrected text.

[0006] Exemplarily, the step of correcting the first punctuation marks in the initial text by using the second punctuation marks to obtain a corrected text comprises: determining a second text interval in which the second punctuation marks are located in the initial text according to a correspondence between the second acoustic features and the initial text; when the second text interval includes a first target text interval, for each first target text interval, adding a corresponding second punctuation mark at the first target text interval to obtain the corrected text, wherein the first target text interval is a text interval in which the first punctuation marks are not present.

[0007] According to an example, the correcting the first punctuation mark in the initial text by the second punctuation mark to obtain a corrected text comprises: determining a second character interval in which the second punctuation mark is located in the initial text according to a correspondence between the second acoustic feature and the initial text; and removing the first punctuation mark at a second target character interval in a first character interval in which the first punctuation mark exists in the initial text to obtain the corrected text, wherein the second target character interval is a character interval in which the second punctuation mark does not exist.

[0008] According to an example, the correcting the first punctuation mark in the initial text by the second punctuation mark to obtain a corrected text comprises: determining a second character interval in which the second punctuation mark is located in the initial text according to a correspondence between the second acoustic feature and the initial text; determining a first acoustic feature corresponding to the first punctuation mark according to a preset mapping relationship, wherein the first acoustic feature and the second acoustic feature are acoustic features of a same type; and selecting one of the second punctuation mark and the first punctuation mark as a punctuation mark at a character interval in which the second punctuation mark and the first punctuation mark are located according to a parameter difference between the second acoustic feature and the first acoustic feature corresponding to the second punctuation mark and the first punctuation mark, respectively.

[0009] According to an example, the first acoustic feature and the second acoustic feature are acoustic features of a single type, and the selecting one of the second punctuation mark and the first punctuation mark as a punctuation mark at a character interval in which the second punctuation mark and the first punctuation mark are located according to a parameter difference between the second acoustic feature and the first acoustic feature corresponding to the second punctuation mark and the first punctuation mark, respectively, comprises: determining the parameter difference between the second acoustic feature and the first acoustic feature corresponding to the second punctuation mark and the first punctuation mark, respectively; retaining the first punctuation mark at the character interval in which the second punctuation mark and the first punctuation mark are located when the parameter difference is less than or equal to a preset parameter threshold; and replacing the first punctuation mark by the second punctuation mark at the character interval in which the second punctuation mark and the first punctuation mark are located when the parameter difference is greater than the preset parameter threshold.

[0010] Exemplarily, the second acoustic feature and the first acoustic feature both include a plurality of types of acoustic features, and the selecting one of the pair of second punctuation mark and first punctuation mark as the punctuation mark at the interval of the pair of second punctuation mark and first punctuation mark according to the parameter difference between the second acoustic feature and the first acoustic feature corresponding to the pair of second punctuation mark and first punctuation mark respectively includes: determining the parameter difference between each type of second acoustic feature and first acoustic feature corresponding to the pair of second punctuation mark and first punctuation mark respectively; determining a difference evaluation score according to the weight coefficient and the parameter difference corresponding to the plurality of types of acoustic features respectively; when the difference evaluation score is less than or equal to a preset score threshold, retaining the first punctuation mark at the interval of the pair of second punctuation mark and first punctuation mark; when the difference evaluation score is greater than the preset score threshold, replacing the first punctuation mark with the second punctuation mark at the interval of the pair of second punctuation mark and first punctuation mark.

[0011] Exemplarily, the second acoustic feature includes a silence segment, and the determining the second punctuation mark of the initial text according to the second acoustic feature includes: determining the second punctuation mark corresponding to each silence segment according to the time length interval to which the time length of the silence segment belongs.

[0012] Exemplarily, the second acoustic feature further includes a prosody feature, and the determining the second punctuation mark corresponding to each silence segment according to the time length interval to which the time length of the silence segment belongs includes: determining the initial punctuation mark at the initial interval of characters corresponding to each silence segment according to the time length interval to which the time length of the silence segment belongs; determining the second punctuation mark at the initial interval of characters from the initial punctuation mark according to the prosodic change trend of the prosodic data of the preset time length corresponding to the characters before the initial interval of characters, wherein the prosodic change trend includes falling and rising.

[0013] Exemplarily, the determining the second punctuation mark at the initial interval of characters from the initial punctuation mark according to the prosodic change trend of the prosodic data of the preset time length corresponding to the characters before the initial interval of characters includes: determining the preliminary selected punctuation mark at the initial interval of characters from the initial punctuation mark according to the prosodic change trend; when the preliminary selected punctuation mark includes only one punctuation mark, taking the preliminary selected punctuation mark as the second punctuation mark at the initial interval of characters; when the preliminary selected punctuation mark includes a plurality of punctuation marks, determining the second punctuation mark at the initial interval of characters from the preliminary selected punctuation mark according to the prosodic change amount within the preset time length under the prosodic change trend.

[0014] Exemplarily, the prosodic features include pitch and / or energy, and determining the second punctuation mark at the initial text interval from the preliminary punctuation marks according to the prosodic variation amount within the preset time length under the prosodic variation trend comprises: determining the second punctuation mark at the initial text interval from the preliminary punctuation marks according to the interval to which the variation amount of the pitch and / or the variation amount of the energy within the preset time length under the prosodic variation trend belongs.

[0015] Exemplarily, the obtaining the initial text with the added first punctuation marks corresponding to the speech data comprises: inputting the speech data into a speech recognition model to predict the initial text.

[0016] According to the second aspect of the present application, a speech recognition text punctuation determination device is further provided, comprising: a text obtaining module configured to obtain an initial text with added first punctuation marks corresponding to speech data; a feature extracting module configured to obtain second acoustic features of the speech data; a symbol calculating module configured to determine second punctuation marks of the initial text according to the second acoustic features; and a revising module configured to revise the first punctuation marks in the initial text by using the second punctuation marks to obtain a revised text.

[0017] According to the third aspect of the present application, an electronic device is further provided, comprising: a processor and a memory, wherein the memory stores computer program instructions, and the computer program instructions are used to execute the speech recognition text punctuation determination method described above when executed by the processor.

[0018] According to the fourth aspect of the present application, a storage medium is further provided, and program instructions are stored on the storage medium, and the program instructions are used to execute the speech recognition text punctuation determination method described above when executed.

[0019] According to the fifth aspect of the present application, a computer program product is further provided, comprising computer program instructions, and the computer program instructions are used to execute the speech recognition text punctuation determination method described above when executed.

[0020] In the technical solution described above, the initial text with added first punctuation marks corresponding to the speech data is obtained, and the second acoustic features of the speech data are obtained, then the second punctuation marks of the initial text are determined according to the second acoustic features, and the first punctuation marks in the initial text are revised by using the second punctuation marks to obtain a revised text. In this way, the speech recognition text with added punctuation marks can be revised according to the acoustic features of the speech data, so that the punctuation marks added in the speech recognition text of the speech data can be more accurate and more consistent with the content expected to be expressed by the user.

[0021] The above description is only a summary of the technical solutions of the present application. In order to make the technical solutions of the present application more clearly understood and implemented, and to make the above and other purposes, features and advantages of the present application more apparent and easy to understand, the following will specifically describe the embodiments of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0022] The above and other purposes, features and advantages of the present application will become more apparent from the following detailed description of the embodiments of the present application, taken in conjunction with the accompanying drawings. The drawings are provided to assist in the understanding of the embodiments of the present application and form a part of the specification. The drawings together with the present application are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally refer to the same components or steps.

[0023] Figure 1 A schematic flow chart of a method for determining punctuation of speech recognition text according to an embodiment of the present application is shown;

[0024] Figure 2 A schematic flow chart of correcting a first punctuation symbol in an initial text with a second punctuation symbol to obtain a corrected text according to an embodiment of the present application is shown;

[0025] Figure 3 A schematic flow chart of correcting a first punctuation symbol in an initial text with a second punctuation symbol to obtain a corrected text according to another embodiment of the present application is shown;

[0026] Figure 4 A schematic flow chart of correcting a first punctuation symbol in an initial text with a second punctuation symbol to obtain a corrected text according to another embodiment of the present application is shown;

[0027] Figure 5 A schematic flow chart of filtering punctuation symbols according to parameter difference of acoustic features according to an embodiment of the present application is shown;

[0028] Figure 6 A schematic flow chart of filtering punctuation symbols according to parameter difference of acoustic features according to another embodiment of the present application is shown;

[0029] Figure 7 A schematic flow chart of determining a second punctuation symbol according to prosodic features and silence segments is shown;

[0030] Figure 8 A schematic flow chart of further determining a second punctuation symbol according to prosodic variation trend according to an embodiment of the present application is shown;

[0031] Figure 9 A schematic block diagram of a device for determining punctuation of speech recognition text according to an embodiment of the present application is shown;

[0032] Figure 10 A schematic block diagram of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0033] In order to make the objectives, technical solutions and advantages of the present application more apparent, the following will describe exemplary embodiments of the present application in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein. Based on the embodiments of the present application described in the present application, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present application.

[0034] In order to at least partially solve the above problems, a speech recognition text punctuation determination method is proposed. The method corrects the punctuation-added speech recognition text of the speech data according to the acoustic features of the speech data, so that the punctuation added in the speech recognition text of the speech data is more accurate and more consistent with the content expressed by the user.

[0035] Figure 1 A schematic flow chart of a speech recognition text punctuation determination method according to an embodiment of the present application is shown. As shown in Figure 1 The speech recognition text punctuation determination method can include steps S110 to S140.

[0036] In step S110, an initial text with added first punctuation symbols corresponding to the speech data is obtained.

[0037] The text sequence of the speech data can be determined first, and then the context information of the text sequence is analyzed by using a natural language processing model to determine the first punctuation symbols at each character interval of the text sequence. These first punctuation symbols are added to the corresponding positions in the text sequence to obtain the initial text. The characters in the text sequence are the characters in the initial text.

[0038] Optionally, the speech data can be retrieved from a speech database to determine the text sequence of the speech data.

[0039] Optionally, the acoustic features of the speech data can be recognized by using an acoustic model, and the text sequence corresponding to the acoustic features can be determined by using a language model to obtain the text sequence of the speech data.

[0040] Exemplarily, the speech data can be input into a speech recognition model to predict the initial text.

[0041] The voice data can be predicted by using a modular or end-to-end speech recognition model to predict the initial text. The modular speech recognition model can include BiLSTM / Transformer, and the end-to-end speech recognition model can include RNN-T, Transformer-based ASR, etc. The modular or end-to-end speech recognition model can automatically predict the text sequence of the voice data, and add the corresponding first punctuation symbol in the text sequence according to the context information of the text sequence.

[0042] In step S120, the second acoustic feature of the voice data is obtained.

[0043] The second acoustic feature can include one or more of prosodic features, silence duration, intensity, and other types of acoustic features. For each type of second acoustic feature, the corresponding algorithm of the second acoustic feature can be used to extract the second acoustic feature, and the acoustic model can also be used to extract the algorithm corresponding to the second acoustic feature to extract the second acoustic feature.

[0044] In step S130, the second punctuation symbol of the initial text is determined according to the second acoustic feature.

[0045] Optionally, the mapping rule of the acoustic feature and the punctuation symbol can be determined in advance, and then the punctuation symbol corresponding to the second acoustic feature is matched from the mapping rule to serve as the second punctuation symbol of the initial text. For example, the second acoustic feature includes a silence segment, and the mapping rule is that the duration of the silence segment corresponding to the comma is 200-500 milliseconds. Therefore, according to the silence segment of 300 milliseconds, it can be determined that the second punctuation symbol corresponding to the silence segment is a comma. The matching method for other acoustic features is similar, and the corresponding second punctuation symbol can be matched from the mapping rule.

[0046] It can be understood that each voice data segment in the voice data has a corresponding second acoustic feature, so according to the second acoustic feature and the mapping rule, not only the type of the matched second symbol (for example, comma, period, etc.) can be determined, but also the position in the initial text to which the matched second symbol should be added can be determined according to the position of the voice data segment corresponding to the second acoustic feature in the voice data. For example, the text sequence (only containing characters) corresponding to the initial text can be converted into a phoneme sequence (such as using the CMU phoneme set or the International Phonetic Alphabet), and then the voice data and the phoneme sequence can be aligned to determine the start and end time stamps of each phoneme in the phoneme sequence. Then, according to the time information of the voice data segment corresponding to the second acoustic feature and the start and end time stamps of the phonemes, the correspondence between the second acoustic feature and the characters of the initial text can be determined. According to the correspondence, the position in the initial text to which the second symbol corresponding to the second acoustic feature should be added can be determined. After determining the position in the initial text to which the second symbol should be added and the type of the second symbol, the complete information of the second symbol can be obtained.

[0047] Optionally, the second acoustic feature and the corresponding second symbol can be used as training data to train an artificial intelligence model, and then the second acoustic feature is input into the trained artificial intelligence model for recognition to determine the second symbol.

[0048] In step S140, the first punctuation symbol in the initial text is corrected by using the second punctuation symbol to obtain a corrected text.

[0049] The first punctuation symbol in the initial text is determined based on the context information, but because the speaking habits of a speaker may vary in different tones, such as tone, pause, etc. Therefore, the first punctuation symbol determined only based on the context information may not be accurate, and the first punctuation symbol can be screened according to the second punctuation symbol to determine a more accurate punctuation symbol.

[0050] For example, the positions of the first punctuation and the second punctuation in the initial text can be compared, and then the second punctuation can be added at the position corresponding to the second punctuation but not corresponding to the first punctuation.

[0051] For example, the positions of the first punctuation and the second punctuation in the initial text can be compared, and then the first punctuation at the position corresponding to the first punctuation but not corresponding to the second punctuation can be deleted.

[0052] For example, the second punctuation can be added to the corresponding position in the initial text, and one of the second punctuation and the first punctuation symbol at the same position can be retained according to a preset rule. The preset rule can include deleting the existing first punctuation symbol at the same position, deleting the second punctuation symbol at the position of the existing first punctuation symbol, retaining the punctuation symbol with higher confidence, etc.

[0053] For example, the second punctuation mark can be added to the corresponding position in the initial text and replace the first punctuation mark at the corresponding position when the first punctuation mark exists at the corresponding position according to the difference between the second punctuation mark and the corresponding acoustic feature at the corresponding position in the initial text. When the difference between the acoustic features is too large, the second punctuation mark is added to the corresponding position and replaces the first punctuation mark at the corresponding position when the first punctuation mark exists at the corresponding position.

[0054] It can be understood that the punctuation mark can be added to the text interval of the initial text, and thus the corresponding acoustic feature of the second punctuation mark at the corresponding position in the initial text is the acoustic feature corresponding to the text interval where the second punctuation mark is located. It can be understood that there is a corresponding relationship between the acoustic feature and the punctuation mark, and the corresponding acoustic feature is different when the text interval has different punctuation marks or has no punctuation mark. When the difference between the acoustic features is too large, it indicates that the punctuation mark that should exist at the text interval can be misjudged as a wrong punctuation mark or misjudged as no punctuation mark, and thus the corresponding second punctuation mark can be added to the text interval or the first punctuation mark at the text interval can be replaced. When the difference between the acoustic features is too small, it indicates that the first punctuation mark at the text interval corresponding to the difference is not necessarily wrong. Because the second punctuation mark determined according to the acoustic feature is not necessarily completely accurate due to the influence of the speaking habit of the speaker, and the first punctuation mark combines the context information, the corresponding second punctuation mark does not need to be added to the text interval or the punctuation mark at the text interval is not replaced by the second punctuation mark.

[0055] For example, it can be determined in advance by using the punctuation dictionary whether there is a fixed text between the fixed format text and the punctuation mark in the initial text (for example, the call "Mr."). The first punctuation mark in the fixed text will not be corrected by the second mark.

[0056] In the technical solution, the acoustic feature of the speech data can be used to correct the speech recognition text of the speech data to which the punctuation mark has been added, so that the punctuation mark added to the speech recognition text of the speech data is more accurate and more consistent with the content expected to be expressed by the user.

[0057] Figure 2 A schematic flowchart of correcting the first punctuation mark in the initial text by using the second punctuation mark to obtain the corrected text according to an embodiment of the present application is shown. As shown in Figure 2 The step S140 can include steps S210 to S220.

[0058] At step S210, a second punctuation mark is determined to be located at a second text interval in the initial text according to a correspondence between the second acoustic feature and the initial text.

[0059] The text sequence (containing only texts) corresponding to the initial text can be converted into a phoneme sequence, and then the speech data and the phoneme sequence can be aligned to determine the start and end time stamps of each phoneme in the phoneme sequence. Then, the correspondence between the second acoustic feature and the texts of the initial text can be determined according to the time information of the speech data segment corresponding to the second acoustic feature and the start and end time stamps of the phonemes. The punctuation mark needs to be determined at the text interval of the initial text. Therefore, the text interval at which the second punctuation mark is located in the initial text can be determined according to the second acoustic feature corresponding to the initial text and the correspondence. For example, the texts corresponding to the start and end time of the silence segment of the speech data in the phoneme sequence can be determined according to the phonemes corresponding to the start and end time of the silence segment, respectively, and the text interval between the two texts can be used as the second text interval. For another example, whether the text interval is the second text interval can be determined according to the prosody feature corresponding to the text before the text interval. For another example, whether the text interval is the second text interval can be determined according to the speech intensity corresponding to the text interval. The second text interval at which the second punctuation mark is located is the text interval corresponding to the second punctuation mark.

[0060] At step S220, when the second text interval includes a first target text interval, for each first target text interval, a corresponding second punctuation mark is added at the first target text interval to obtain a corrected text, wherein the first target text interval is a text interval in which the first punctuation mark does not exist.

[0061] When the second text interval includes the first target text interval, it indicates that the first punctuation mark in the initial text is missing, and a corresponding second punctuation mark can be added at the first target text interval in which the first punctuation mark is not added.

[0062] In the above technical solution, a second punctuation mark is determined to be located at a second text interval in the initial text according to a correspondence between the second acoustic feature and the initial text, and when the second text interval includes a first target text interval, for each first target text interval, a corresponding second punctuation mark is added at the first target text interval to obtain a corrected text, wherein the first target text interval is a text interval in which the first punctuation mark does not exist. In this way, the previously missing punctuation mark can be supplemented in the speech recognition text according to the acoustic feature, and the accuracy and readability of the semantic recognition text content can be improved.

[0063] Figure 3A flowchart illustrating a method for correcting first punctuation marks in an initial text with second punctuation marks to obtain a corrected text according to another embodiment of the present application is shown in FIG. 4. As shown in FIG. 4, the method includes steps S410-S430. Figure 3 The step S140 can include steps S310-S320.

[0064] In step S310, a second text interval where the second punctuation mark is located in the initial text is determined according to a correspondence between the second acoustic feature and the initial text.

[0065] The step S310 is similar to the step S210, and thus is not described in detail.

[0066] In step S320, the first punctuation mark at the second target text interval is removed to obtain the corrected text, when the first text interval includes the second target text interval, where the first text interval is a text interval in the initial text where the first punctuation mark exists, and the second target text interval is a text interval where the second punctuation mark does not exist.

[0067] When the first text interval includes the second target text interval, it indicates that the first punctuation mark at the second target text interval is a misjudged punctuation mark, and the first punctuation mark at the second target text interval can be removed.

[0068] In the above technical solution, the second text interval where the second punctuation mark is located in the initial text is determined according to a correspondence between the second acoustic feature and the initial text, and the first punctuation mark at the second target text interval is removed to obtain the corrected text, when the first text interval includes the second target text interval, where the first text interval is a text interval in the initial text where the first punctuation mark exists, and the second target text interval is a text interval where the second punctuation mark does not exist. In this way, the misjudged punctuation mark can be removed from the speech recognition text according to the acoustic feature, and the semantic of the speech recognition text is prevented from being distorted by the redundant punctuation mark, and the accuracy and readability of the speech recognition text are improved.

[0069] Figure 4 A flowchart illustrating a method for correcting first punctuation marks in an initial text with second punctuation marks to obtain a corrected text according to another embodiment of the present application is shown in FIG. 4. As shown in FIG. 4, the method includes steps S410-S430. Figure 4 The step S140 can include steps S410-S430.

[0070] In step S410, a second text interval where the second punctuation mark is located in the initial text is determined according to a correspondence between the second acoustic feature and the initial text.

[0071] The step S410 is similar to the step S210, and thus is not described in detail.

[0072] At step S420, a first acoustic feature corresponding to the first punctuation mark is determined according to a preset mapping relationship, where the first acoustic feature and the second acoustic feature are acoustic features of the same type.

[0073] The first acoustic feature corresponding to the first punctuation mark can be matched from the preset mapping relationship. Taking the first acoustic feature including a silence segment as an example, the preset mapping relationship indicates that the duration of the silence segment corresponding to the comma is 200-500 milliseconds, and thus the first acoustic feature corresponding to the comma is determined to be a silence segment of 200-500 milliseconds. The matching manner of other acoustic features is similar, and the second punctuation mark corresponding to the acoustic feature can be matched from the preset mapping relationship. The preset mapping relationship can be the same as the mapping rule described above. In the preset mapping relationship, the acoustic feature corresponding to the punctuation mark can be an acoustic feature within a certain range or an acoustic feature with a fixed parameter value.

[0074] At step S430, for each pair of second punctuation marks and first punctuation marks in the same text interval, a parameter difference between the second acoustic feature and the first acoustic feature corresponding to the pair of second punctuation marks and first punctuation marks is determined, and one of the pair of second punctuation marks and first punctuation marks is selected as the punctuation mark in the text interval.

[0075] When the first acoustic feature corresponding to the first punctuation mark and the second acoustic feature corresponding to the second punctuation mark are acoustic features within a certain range, the parameter difference can be determined according to the difference between the ranges of the two acoustic features (such as the difference between the maximum value of the first acoustic feature and the maximum value of the second acoustic feature, the difference between the minimum value of the first acoustic feature and the minimum value of the second acoustic feature, etc.).

[0076] When the second acoustic feature and the first acoustic feature corresponding to the pair of second punctuation marks and first punctuation marks are acoustic features with fixed values, the difference between the two acoustic features can be taken as the parameter difference.

[0077] When the first acoustic feature and the second acoustic feature are acoustic features of multiple types, the parameter difference can be determined by comprehensively considering the differences between the acoustic features of the multiple types.

[0078] When the parameter difference is too large, it indicates that the first punctuation mark in the text interval can be misjudged due to the context information, and the second punctuation mark in the text interval can be used to replace the first punctuation mark in the text interval. When the parameter difference is relatively small, it indicates that the second punctuation mark in the text interval can be misjudged due to the speaking habit of the speaker, and the first punctuation mark in the text interval can be retained.

[0079] In the technical solution, according to the correspondence between the second acoustic feature and the initial text, the second text interval in which the second punctuation mark is located is determined in the initial text, and then according to a preset mapping relationship, the first acoustic feature corresponding to the first punctuation mark is determined, wherein the first acoustic feature and the second acoustic feature are acoustic features of the same type. Then, for each pair of second punctuation mark and first punctuation mark located in the same text interval, according to the parameter difference between the second acoustic feature and the first acoustic feature corresponding to the pair of second punctuation mark and first punctuation mark respectively, one of the pair of second punctuation mark and first punctuation mark is selected as the punctuation mark at the text interval where the pair of second punctuation mark and first punctuation mark is located. In this way, the punctuation marks of the speech recognition text of the speech data can be more accurately determined according to the acoustic features in combination with the speaking habits of the speaker, which is beneficial to improving the accuracy and readability of the speech recognition text.

[0080] Exemplarily, the first acoustic feature and the second acoustic feature are acoustic features of a single type. Figure 5 A schematic flowchart of screening punctuation marks according to a parameter difference of acoustic features according to one embodiment of the present application is shown. As shown in Figure 5 The step S430 can include steps S510 to S530.

[0081] In step S510, a parameter difference between the second acoustic feature and the first acoustic feature corresponding to the pair of second punctuation mark and first punctuation mark respectively is determined.

[0082] When the first acoustic feature corresponding to the first punctuation mark and the second acoustic feature corresponding to the second punctuation mark are acoustic features within a certain range, the parameter difference can be determined according to the difference between the ranges of the two acoustic features (such as the difference between the maximum value of the first acoustic feature and the maximum value of the second acoustic feature, the difference between the minimum value of the first acoustic feature and the minimum value of the second acoustic feature, etc.). Taking the case where the acoustic feature is a silence segment as an example, if the time range of the silence segment corresponding to the first acoustic feature corresponding to the first punctuation mark is 200-500 milliseconds, and the second acoustic feature corresponding to the second punctuation mark is 600-700 milliseconds. The parameter difference A1 can be determined according to the difference between the ranges of the two acoustic features, i.e. A1 = [(600-200)+(700-500)] / 2 = 300, which is the average of the difference between the maximum value of the first acoustic feature and the maximum value of the second acoustic feature and the difference between the minimum value of the first acoustic feature and the minimum value of the second acoustic feature. The above example is only for illustrative purposes. For this case, the parameter difference representing the difference between the first acoustic feature and the second acoustic feature can also be determined in other ways in combination with the range of the first acoustic feature and the range of the second acoustic feature. And similar calculation methods can also be applied to other types of first acoustic features and second acoustic features.

[0083] When the second acoustic feature and the first acoustic feature corresponding to the pair of second punctuation mark and first punctuation mark respectively are acoustic features with fixed numerical values, the difference between the two acoustic features can be taken as the parameter difference. Still taking the acoustic feature as the silence segment as an example, if the duration of the silence segment corresponding to the first acoustic feature corresponding to the first punctuation mark is 200 milliseconds, and the second acoustic feature corresponding to the second punctuation mark is 600. The parameter difference A2=600-200=400 can be determined according to the difference between the ranges of the two acoustic features. The above example is only used for illustrative purposes. For this case, the parameter difference for representing the difference between the first acoustic feature and the second acoustic feature can also be determined in other ways by combining the parameter value of the first acoustic feature and the parameter value of the second acoustic feature. And similar calculation methods can also be applied to other types of first acoustic features and second acoustic features.

[0084] In step S520, when the parameter difference is less than or equal to the preset parameter threshold, the first punctuation mark at the text interval where the pair of second punctuation mark and first punctuation mark is located is retained.

[0085] When the parameter difference is less than or equal to the preset parameter threshold, it indicates that the second punctuation mark at the text interval can be misjudged due to the speaking habit of the speaker, and the first punctuation mark at the text interval can be retained.

[0086] In step S530, when the parameter difference is greater than the preset parameter threshold, the first punctuation mark at the text interval where the pair of second punctuation mark and first punctuation mark is located is replaced by the second punctuation mark.

[0087] When the parameter difference is greater than the preset parameter threshold, it indicates that the first punctuation mark at the text interval can be misjudged due to the context information, and the first punctuation mark at the text interval can be replaced by the second punctuation mark at the text interval.

[0088] In the above technical solution, in the case where the first acoustic feature and the second acoustic feature are acoustic features of a single type, the parameter difference between the second acoustic feature and the first acoustic feature corresponding to the pair of second punctuation mark and first punctuation mark is determined, when the parameter difference is less than or equal to the preset parameter threshold, the first punctuation mark at the text interval where the pair of second punctuation mark and first punctuation mark is located is retained, and when the parameter difference is greater than the preset parameter threshold, the first punctuation mark at the text interval where the pair of second punctuation mark and first punctuation mark is located is replaced by the second punctuation mark. In this way, the misjudgment caused by the speaking habit of the speaker and the context information can be considered, and the erroneous punctuation mark that can exist in the speech recognition text can be corrected to obtain a more accurate speech recognition text with added punctuation marks.

[0089] Exemplarily, the second acoustic feature and the first acoustic feature both contain multiple types of acoustic features. Figure 6 A schematic flow chart of screening punctuation marks according to parameter difference of acoustic features is shown according to yet another embodiment of the present application. As shown in Figure 6 The step S430 can include steps S610-S640.

[0090] In step S610, parameter difference between each type of second acoustic feature and first acoustic feature corresponding to the pair of second punctuation mark and first punctuation mark respectively is determined.

[0091] For each type of second acoustic feature and first acoustic feature corresponding to the pair of second punctuation mark and first punctuation mark respectively, parameter difference between the second acoustic feature and the first acoustic feature of this type can be determined according to the step S530, which will not be described in detail here. The second acoustic feature and the first acoustic feature can contain multiple types of acoustic features such as silence segment, prosodic feature, speech intensity, etc.

[0092] In step S620, difference evaluation score is determined according to weight coefficient and parameter difference corresponding to multiple types of acoustic features respectively.

[0093] Weight coefficient corresponding to each type of acoustic feature can be determined in advance. Exemplarily, the longer the duration of silence segment, the longer the speaking pause of the speaker within the time range of the silence segment, and the greater the impact on punctuation marks of the speech recognition text. Therefore, when the multiple types of acoustic features contain at least silence segment and prosodic feature, the weight corresponding to the silence segment can be greater than the weight coefficient corresponding to the prosodic feature, and can also be greater than the weight coefficient of any other type of acoustic feature. The weight coefficient corresponding to the prosodic feature can also be greater than the weight coefficient of any other type of acoustic feature except the silence segment. The difference evaluation score comprehensively reflects the difference between multiple types of acoustic features, and can more accurately evaluate the difference between the first acoustic feature and the second acoustic feature.

[0094] In step S630, when the difference evaluation score is less than or equal to a preset score threshold, the first punctuation mark is retained at the text interval where the pair of second punctuation mark and first punctuation mark is located.

[0095] When the difference evaluation score is less than or equal to the preset score threshold, it indicates that the second punctuation mark in the text interval can be misjudged due to the speaking habit of the speaker, and the first punctuation mark in the text interval can be retained.

[0096] In step S640, when the difference evaluation score is greater than the preset score threshold, the first punctuation mark is replaced by the second punctuation mark at the text interval where the pair of second punctuation mark and first punctuation mark is located.

[0097] When the difference evaluation score is greater than the preset score threshold, it indicates that the first punctuation mark in the text interval can be misjudged due to the context information, and the second punctuation mark in the text interval can be used to replace the first punctuation mark in the text interval.

[0098] In the technical solution, the parameter difference between each type of second acoustic feature and first acoustic feature corresponding to the pair of second punctuation mark and first punctuation mark is determined, then the difference evaluation score is determined according to the weight coefficient and the parameter difference corresponding to the acoustic feature of multiple types, and then when the difference evaluation score is less than or equal to the preset score threshold, the first punctuation mark is retained in the text interval where the pair of second punctuation mark and first punctuation mark is located, and when the difference evaluation score is greater than the preset score threshold, the second punctuation mark is used to replace the first punctuation mark in the text interval where the pair of second punctuation mark and first punctuation mark is located. The difference evaluation score comprehensively considers the difference between multiple types of acoustic features, which can more accurately evaluate the difference between the first acoustic feature and the second acoustic feature. Using the difference evaluation score can more accurately evaluate which one of the first punctuation and the second punctuation in the same text interval is more suitable as the punctuation mark of the text interval, so as to more accurately determine the punctuation mark of the speech recognition text and improve the accuracy of the speech recognition text.

[0099] Exemplarily, the second acoustic feature includes a silence segment, and the step S130 of determining the second punctuation mark of the initial text according to the second acoustic feature includes a step S131 of determining the second punctuation mark corresponding to each silence segment according to the time length interval to which the time length of the silence segment belongs.

[0100] The speech data can be subjected to speech detection to determine the silence segment of the speech data.

[0101] In the mapping rule or the preset mapping relationship, each punctuation symbol corresponds to a time length range of a silence segment, and the time length range is an interval to which a time length of the silence segment belongs. The time length of the silence segment represents a time length of a pause of the speaker, and the time length of the silence segment can be searched in the mapping rule or the preset mapping relationship to match the second punctuation symbol corresponding to the silence segment. If the time length of the silence segment only belongs to one time length interval in the mapping rule or the preset mapping relationship, the second punctuation symbol corresponding to the silence segment can be directly determined. For example, the time length interval to which the time length of the silence segment corresponding to the comma belongs is 200 milliseconds to 500 milliseconds, the time length interval to which the time length of the silence segment corresponding to the colon belongs is 100 milliseconds to 200 milliseconds, and if the time length of the silence segment is 150 milliseconds, the second punctuation symbol corresponding to the silence segment is the colon. If there is an overlap among the time length intervals in the mapping rule or the preset mapping relationship, the time length of the silence segment can belong to multiple time length intervals in the mapping rule or the preset mapping relationship, and in this case, the determined punctuation symbol can be further screened in combination with other types of acoustic features to obtain the second punctuation symbol.

[0102] In the technical solution, the second punctuation symbol corresponding to each silence segment is determined according to the time length interval to which the time length of the silence segment belongs, the pause of the speaker can be considered, the punctuation symbol with reference value can be determined, and the punctuation symbol in the initial text can be corrected according to the punctuation symbol, so that the semantics of the corrected text is more accurate.

[0103] For example, the second acoustic feature further includes a prosody feature. Figure 7 An illustrative flowchart for determining the second punctuation symbol according to the prosody feature and the silence segment is shown. As shown in Figure 7 The step S131 can further include steps S710 to S720.

[0104] In step S710, the initial punctuation symbol at the initial text interval corresponding to each silence segment is determined according to the time length interval to which the time length of the silence segment belongs.

[0105] The time length interval to which the time length of the silence segment corresponding to different second punctuation symbols belongs can partially overlap. For example, the pause time length corresponding to the question mark and the period in the speaking process is close, and the period and the question mark can be determined according to the time length interval to which the time length of the silence segment belongs. In this case, the period and the question mark need to be further distinguished to determine the unique first punctuation symbol at each text interval. Similarly, if the initial punctuation symbol at the initial text interval corresponding to each silence segment is determined to be multiple according to the time length interval to which the time length of the silence segment belongs, the initial punctuation symbol needs to be further screened. The process of determining the initial text interval corresponding to the silence segment is similar to the step S210, and will not be described here.

[0106] At step S720, the second punctuation mark at the initial text interval is determined from the initial punctuation mark according to the prosodic variation trend of the preset duration of the speech data corresponding to the text before the initial text interval, wherein the prosodic variation trend comprises falling and rising.

[0107] In the case of close correspondence to the corresponding speaking pause, the prosodic variation trend corresponding to the text before different punctuation marks is not necessarily the same, and the initial punctuation mark can be further screened based on the correspondence between the punctuation mark and the prosodic variation trend. Taking a period and a question mark as an example, the prosodic variation trend corresponding to the text before the period should generally be a falling trend, and the prosodic variation trend corresponding to the text before the question mark should generally be a rising trend. Thus, the second punctuation mark at the initial text interval can be determined from the period and the question mark according to the prosodic variation trend of the preset duration of the speech data corresponding to the text before the initial text interval. Similar to other punctuation marks that meet the above conditions, details are not described here. For example, the prosodic variation trend can also include rising followed by falling, for example, the prosodic variation trend of the preset duration of the speech data corresponding to the text before the exclamation mark can be rising followed by falling.

[0108] For example, for speech data "Really? [silence 800ms] I don't think so" ([silence 800ms] represents that the text interval corresponds to an 800-millisecond silence segment), the text to which the first punctuation mark is added is "really i don't think so", and the question mark and the period that should be added are lost according to the context information only. After detection, it is determined that there is a rising prosodic variation trend of 300 milliseconds and an 800-millisecond silence at the end of "Really", and there is a 600ms silence and a falling prosodic variation trend at the end of "so", at this time, it can be determined that the second punctuation mark after "Really" is a question mark, and the second punctuation mark after "so" is a period. Thus, the corrected text "Really? I don't think so." can be obtained.

[0109] In the above technical solution, the initial punctuation mark at the initial text interval corresponding to each silence segment is determined according to the time duration interval to which the time duration of the silence segment belongs, and then the second punctuation mark at the initial text interval is determined from the initial punctuation mark according to the prosodic variation trend of the preset duration of the speech data corresponding to the text before the initial text interval, wherein the prosodic variation trend comprises falling and rising. Thus, the second punctuation mark can be more accurately determined, the first punctuation mark in the initial text can be more accurately corrected, and more accurate speech recognition text can be obtained.

[0110] Figure 8A schematic flowchart of further determining the second punctuation mark according to the prosodic variation trend is shown according to an embodiment of the present application. As shown in Figure 8 The step S720 can include steps S810-S830.

[0111] In step S810, a preliminary punctuation mark at the initial text interval is determined from the initial punctuation mark according to the prosodic variation trend.

[0112] The preliminary punctuation mark at the initial text interval can be determined from the initial punctuation mark based on the correspondence between the punctuation mark and the prosodic variation trend.

[0113] In step S820, when the preliminary punctuation mark includes only one punctuation mark, the preliminary punctuation mark is determined as the second punctuation mark at the initial text interval.

[0114] In step S830, when the preliminary punctuation mark includes multiple punctuation marks, the second punctuation mark at the initial text interval is determined from the preliminary punctuation mark according to the prosodic variation amount within the preset time length under the prosodic variation trend.

[0115] The prosodic variation trends corresponding to different initial punctuation marks can be the same. For example, because the preset time length for determining the prosodic variation trend is set too short, the prosodic variation trend of rising and then falling can be identified as a prosodic variation trend of falling. For example, a question mark that should have been added can be identified as an exclamation mark as the second punctuation mark. The exclamation mark is usually used to indicate that the speaker speaks with strong emotion, so the prosodic variation amount within the preset time length will be greater than that of the question mark. According to the prosodic variation amount within the preset time length under the prosodic variation trend and the correspondence between the punctuation marks, the second punctuation mark at the initial text interval can be determined from the exclamation mark and the question mark. Similarly, for other punctuation marks that meet the above conditions, the second punctuation mark at the initial text interval can be further distinguished according to the prosodic variation amount in a similar manner.

[0116] Exemplarily, the prosodic features can include pitch and / or energy. The second punctuation mark at the initial text interval can be determined from the preliminary punctuation mark according to the interval to which the variation amount of the pitch and / or the variation amount of the energy within the preset time length under the prosodic variation trend belong.

[0117] The second punctuation mark can be further screened according to one or more of the variation amount of the pitch and the variation amount of the energy. When the second punctuation mark is further screened using the variation amount of the pitch and the variation amount of the energy, the interval to which the variation amount of the pitch and the variation amount of the energy within the preset time length belong need to be matched from the mapping rules or the preset mapping relationship at the same time, and the second punctuation mark at the initial text interval is determined according to the interval to which each of the two belongs.

[0118] In the technical solution, the initial punctuation mark at the initial character interval is determined according to the rhythm change trend from the initial punctuation mark, when the initial punctuation mark includes only one punctuation mark, the initial punctuation mark is taken as the second punctuation mark at the initial character interval, and when the initial punctuation mark includes multiple punctuation marks, the second punctuation mark at the initial character interval is determined from the initial punctuation mark according to the rhythm change amount within the preset time length according to the rhythm change trend. The rhythm change amount is used to further screen the determined multiple punctuation marks, so as to determine a more accurate second punctuation mark, and it is beneficial to obtain a more accurate speech recognition text.

[0119] Figure 9 A schematic block diagram of a speech recognition text punctuation determination apparatus according to one embodiment of the present application is shown. As shown in the figure, the speech recognition text punctuation determination apparatus includes a text acquisition module 910, a feature extraction module 920, a symbol calculation module 930, and a correction module 940. Figure 9

[0120] The text acquisition module 910 is configured to acquire an initial text with added first punctuation marks corresponding to speech data.

[0121] The feature extraction module 920 is configured to acquire second acoustic features of the speech data.

[0122] The symbol calculation module 930 is configured to determine second punctuation marks of the initial text according to the second acoustic features.

[0123] The correction module 940 is configured to correct the first punctuation marks in the initial text by using the second punctuation marks to obtain a corrected text.

[0124] For example, the correction module 940 includes a first interval determination sub-module and a first symbol correction sub-module. The first interval determination sub-module is configured to determine a second character interval where the second punctuation mark is located in the initial text according to a corresponding relationship between the second acoustic features and the initial text. The first symbol correction sub-module is configured to, when the second character interval includes a first target character interval, add a corresponding second punctuation mark at the first target character interval to obtain the corrected text, wherein the first target character interval is a character interval where there is no first punctuation mark.

[0125] ​Exemplarily, the correction module 940 comprises a second interval determining submodule and a second symbol correcting submodule. The second interval determining submodule is configured to determine, according to the correspondence between the second acoustic features and the initial text, a second text interval in which the second punctuation symbol is located in the initial text. The second symbol correcting submodule is configured to remove the first punctuation symbol at a second target text interval when the first text interval comprises the second target text interval, to obtain the corrected text, wherein the first text interval is a text interval in which the initial text has the first punctuation symbol, and the second target text interval is a text interval in which the second punctuation symbol is absent.

[0126] Exemplarily, the correction module 940 comprises a third interval determining submodule, a first feature extracting submodule and a third symbol correcting submodule. The third interval determining submodule is configured to determine, according to the correspondence between the second acoustic features and the initial text, a second text interval in which the second punctuation symbol is located in the initial text. The first feature extracting submodule is configured to determine, according to a preset mapping relationship, a first acoustic feature corresponding to the first punctuation symbol, wherein the first acoustic feature and the second acoustic feature are acoustic features of the same type. The third symbol correcting submodule is configured to, for each pair of the second punctuation symbol and the first punctuation symbol in the same text interval, select one of the pair of the second punctuation symbol and the first punctuation symbol as a punctuation symbol in the text interval in which the pair of the second punctuation symbol and the first punctuation symbol is located, according to a parameter difference between the second acoustic feature and the first acoustic feature corresponding to the pair of the second punctuation symbol and the first punctuation symbol, respectively.

[0127] Exemplarily, the first acoustic feature and the second acoustic feature are acoustic features of a single type, and the third symbol correcting submodule comprises a first difference calculating submodule and a first screening submodule. The first difference calculating submodule is configured to determine a parameter difference between the second acoustic feature and the first acoustic feature corresponding to the pair of the second punctuation symbol and the first punctuation symbol, respectively. The first screening submodule is configured to retain the first punctuation symbol in the text interval in which the pair of the second punctuation symbol and the first punctuation symbol is located when the parameter difference is less than or equal to a preset parameter threshold, and replace the first punctuation symbol with the second punctuation symbol in the text interval in which the pair of the second punctuation symbol and the first punctuation symbol is located when the parameter difference is greater than the preset parameter threshold.

[0128] Exemplarily, the second acoustic feature and the first acoustic feature both comprise a plurality of types of acoustic features, and the third symbol correction submodule comprises a second difference calculation submodule, a score evaluation submodule, and a second screening submodule. The second difference calculation submodule is configured to determine a parameter difference between each type of the second acoustic feature and the first acoustic feature corresponding to the pair of second punctuation symbol and first punctuation symbol respectively. The score evaluation submodule is configured to determine a difference evaluation score according to a weight coefficient and the parameter difference corresponding to each type of acoustic feature. The second screening submodule is configured to retain the first punctuation symbol at the text interval where the pair of second punctuation symbol and first punctuation symbol is located when the difference evaluation score is less than or equal to a preset score threshold, and replace the first punctuation symbol with the second punctuation symbol at the text interval where the pair of second punctuation symbol and first punctuation symbol is located when the difference evaluation score is greater than the preset score threshold.

[0129] Exemplarily, the second acoustic feature comprises a silence segment, and the symbol calculation module 930 comprises a first calculation submodule. The first calculation submodule is configured to determine the second punctuation symbol corresponding to each silence segment according to a time length interval to which the time length of the silence segment belongs.

[0130] Exemplarily, the second acoustic feature further comprises a prosody feature, and the first calculation submodule comprises a second calculation submodule and a third calculation submodule. The first calculation submodule is configured to determine an initial punctuation symbol at an initial text interval corresponding to each silence segment according to a time length interval to which the time length of the silence segment belongs. The third calculation submodule is configured to determine the second punctuation symbol at the initial text interval from the initial punctuation symbol according to a prosody change trend of the prosody data of a preset time length corresponding to the text before the initial text interval, wherein the prosody change trend comprises a falling and a rising.

[0131] Exemplarily, the third calculation submodule comprises a fourth calculation submodule, a fifth calculation submodule, and a sixth calculation submodule. The fourth calculation submodule is configured to determine a preliminary punctuation symbol at the initial text interval from the initial punctuation symbol according to the prosody change trend. The fifth calculation submodule is configured to take the preliminary punctuation symbol as the second punctuation symbol at the initial text interval when the preliminary punctuation symbol comprises only one punctuation symbol. The sixth calculation submodule is configured to determine the second punctuation symbol at the initial text interval from the preliminary punctuation symbol according to a prosody change amount within a preset time length under the prosody change trend when the preliminary punctuation symbol comprises a plurality of punctuation symbols.

[0132] Exemplarily, the prosody feature comprises a pitch and / or energy, and the sixth calculation submodule comprises a seventh calculation submodule. The seventh calculation submodule is configured to determine the second punctuation symbol at the initial text interval from the preliminary punctuation symbol according to an interval to which a change amount of the pitch and / or a change amount of the energy within a preset time length under the prosody change trend belongs.

[0133] Exemplarily, the text obtaining module 910 comprises a text predicting sub-module. The text predicting sub-module is configured to input the voice data into a voice recognition model to predict the initial text.

[0134] According to another aspect of the present application, there is also provided an electronic device. Figure 10 A schematic block diagram of an electronic device according to an embodiment of the present application is shown. As shown, the electronic device comprises a processor and a memory, wherein the memory stores computer program instructions which, when executed by the processor, are configured to perform the voice recognition text punctuation determination method as described above. Figure 10

[0135] In addition, according to still another aspect of the present application, there is also provided a storage medium having stored thereon program instructions which, when executed by a computer or a processor, cause the computer or the processor to perform the above-mentioned voice recognition text punctuation determination method according to the embodiments of the present application, and to implement the corresponding modules in the above-mentioned voice recognition text punctuation determination apparatus according to the embodiments of the present application. The storage medium may, for example, include a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above-mentioned storage media. The computer readable storage medium can be any combination of one or more computer readable storage media.

[0136] According to another aspect of the present application, there is also provided a computer program product comprising computer program instructions which, when executed, are configured to perform the above-mentioned voice recognition text punctuation determination method.

[0137] Those skilled in the art can understand the specific implementation and advantages of the above-mentioned voice recognition text punctuation determination apparatus, electronic device, storage medium and computer program product by reading the above-mentioned specific description of the voice recognition text punctuation determination method, and for brevity, will not be repeated here.

[0138] Although the example embodiments have been described herein with reference to the accompanying drawings, it is to be understood that the above-described example embodiments are merely exemplary and are not intended to limit the scope of the present application thereto. Various changes and modifications can be made thereto by those skilled in the art without departing from the scope and spirit of the present application. All such changes and modifications are intended to be included within the scope of the present application as claimed in the appended claims.

[0139] ​Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0140] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the above-described device embodiments are merely illustrative, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0141] In the specification provided herein, a large number of specific details are described. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some examples, well-known methods, structures and techniques are not described in detail in order not to obscure the understanding of the present specification.

[0142] Similarly, it should be understood that, in order to simplify the present application and help understand one or more of the various inventive aspects, in the description of the exemplary embodiments of the present application, various features of the present application are sometimes grouped into a single embodiment, figure, or description thereof. However, this method of the present application should not be interpreted as reflecting an intention that the claimed application requires more features than those explicitly recited in each claim. Rather, as reflected by the corresponding claims, the inventive point is that the corresponding technical problem can be solved with fewer features than all the features of a certain disclosed single embodiment. Therefore, the claims following the specific embodiments are hereby expressly incorporated into the specific embodiments, wherein each claim itself is a separate embodiment of the present application.

[0143] Those skilled in the art can understand that, except for the mutual exclusion between features, all features disclosed in the specification (including the accompanying claims, abstract and drawings) and all processes or units of any method or device disclosed in this way can be combined in any combination. Unless explicitly stated otherwise, each feature disclosed in the specification (including the accompanying claims, abstract and drawings) can be replaced by an alternative feature that provides the same, equivalent or similar purpose.

[0144] Furthermore, those skilled in the art will understand that although some embodiments described herein include certain features but not others included in other embodiments, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments. For example, in the claims, any of the claimed embodiments can be used in any combination.

[0145] The various component embodiments of the present invention can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some modules in the speech recognition text punctuation determination apparatus according to embodiments of the present invention. The present invention can also be implemented as an apparatus program (e.g., a computer program and computer program product) for performing some or all of the methods described herein. Such programs implementing the present invention can be stored on a computer-readable medium or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.

[0146] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.

[0147] The above description is merely a specific embodiment of the present invention or an explanation of that embodiment. The scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. The scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for determining punctuation in speech recognition text, characterized in that, The method includes: Obtain the initial text corresponding to the voice data, with the first punctuation mark added; Obtain the second acoustic features of the speech data; Based on the second acoustic feature, determine the second punctuation mark of the initial text; and The first punctuation mark in the initial text is corrected using the second punctuation mark to obtain the corrected text.

2. The method according to claim 1, characterized in that, The step of correcting the first punctuation mark in the initial text using the second punctuation mark to obtain the corrected text includes: Based on the correspondence between the second acoustic feature and the initial text, the second character interval where the second punctuation mark is located is determined in the initial text; When the second text interval includes the first target text interval, for each first target text interval, a corresponding second punctuation mark is added at the first target text interval to obtain the corrected text, wherein the first target text interval is the text interval in which the first punctuation mark is not present.

3. The method according to claim 1, characterized in that, The step of correcting the first punctuation mark in the initial text using the second punctuation mark to obtain the corrected text includes: Based on the correspondence between the second acoustic feature and the initial text, the second character interval where the second punctuation mark is located is determined in the initial text; When the first text interval includes the second target text interval, the first punctuation mark at the second target text interval is removed to obtain the corrected text, wherein the first text interval is the text interval of the initial text that contains the first punctuation mark, and the second target text interval is the text interval in which the second punctuation mark is not present.

4. The method according to claim 1, characterized in that, The step of correcting the first punctuation mark in the initial text using the second punctuation mark to obtain the corrected text includes: Based on the correspondence between the second acoustic feature and the initial text, the second character interval where the second punctuation mark is located is determined in the initial text; According to a preset mapping relationship, the first acoustic feature corresponding to the first punctuation mark is determined, wherein the first acoustic feature and the second acoustic feature are acoustic features of the same type. For each pair of second and first punctuation marks located in the same text interval, one of the second and first punctuation marks is selected as the punctuation mark at the text interval based on the parameter difference between the second and first acoustic features corresponding to the pair of second and first punctuation marks, respectively.

5. The method according to claim 4, characterized in that, The first acoustic feature and the second acoustic feature are single-type acoustic features. The step of selecting one of the pair of second and first punctuation marks as the punctuation mark at the text interval corresponding to the pair of second and first punctuation marks, based on the parameter difference between the second and first acoustic features respectively, includes: Determine the parameter differences between the second acoustic feature and the first acoustic feature corresponding to the second punctuation mark and the first punctuation mark, respectively; When the parameter difference is less than or equal to a preset parameter threshold, the first punctuation mark is retained at the text interval where the second punctuation mark and the first punctuation mark are located; When the parameter difference is greater than the preset parameter threshold, the second punctuation mark is used to replace the first punctuation mark at the text interval where the second punctuation mark and the first punctuation mark are located.

6. The method according to claim 4, characterized in that, Both the second acoustic feature and the first acoustic feature contain multiple types of acoustic features. The step of selecting one of the pair of second and first punctuation marks as the punctuation mark at the text interval corresponding to the pair of second and first punctuation marks, based on the parameter difference between the second and first acoustic features respectively, includes: Determine the parameter differences between the second acoustic feature and the first acoustic feature corresponding to each type of the pair of second and first punctuation marks; The difference assessment score is determined based on the weighting coefficients and parameter differences corresponding to the various types of acoustic features; When the difference evaluation score is less than or equal to a preset score threshold, the first punctuation mark is retained at the text interval where the second punctuation mark and the first punctuation mark are located; When the difference assessment score is greater than the preset score threshold, the second punctuation mark is used to replace the first punctuation mark at the text interval where the second punctuation mark and the first punctuation mark are located.

7. A speech recognition text punctuation determination device, characterized in that, include: The text acquisition module is used to acquire the initial text corresponding to the voice data, with the first punctuation mark added. The feature extraction module is used to obtain the second acoustic features of the speech data; A symbol calculation module is used to determine the second punctuation mark of the initial text based on the second acoustic feature; as well as The correction module is used to correct the first punctuation mark in the initial text using the second punctuation mark to obtain the corrected text.

8. An electronic device comprising a processor and a memory, characterized in that, The memory stores computer program instructions, which, when executed by the processor, are used to perform the speech recognition text punctuation determination method as described in any one of claims 1 to 6.

9. A storage medium on which program instructions are stored, characterized in that, The program instructions, when executed, are used to perform the speech recognition text punctuation determination method as described in any one of claims 1 to 6.

10. A computer program product comprising computer program instructions, characterized in that, The computer program instructions, when executed, are used to perform the speech recognition text punctuation determination method as described in any one of claims 1 to 6.