Real-time punctuation recovery method based on efficient corpus screening
By constructing and optimizing training and validation sets, and combining open-source models for data weighting and fine-tuning, the accuracy and real-time performance issues of existing real-time punctuation recovery models have been resolved, achieving more efficient punctuation recovery results, which are suitable for corpus processing after speech recognition.
Patent Information
- Application Number
- CN202511051615.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-07-29
AI Technical Summary
Existing real-time punctuation recovery models are not performing well, especially when dealing with incomplete sentences and multi-sentence inputs. They cannot efficiently and accurately recover punctuation marks, and traditional models such as the BART-large-Chinese model have shortcomings in real-time performance and accuracy.
By using an efficient corpus selection method, training, validation, and test sets were constructed. Combined with Chinese monolingual error-corrected corpus and speech recognition corpus, data cleaning and splicing were performed. Punctuation restoration experiments were conducted using the open-source CT-transformer and Qwen2.5 models. The punctuation restoration effect was optimized through data weighting and model fine-tuning.
It significantly improves the accuracy and efficiency of the real-time punctuation recovery model, shortens the inference time, adapts to the real-time requirements of speech recognition, and enhances the readability and semantic integrity of the text after speech recognition processing.
Smart Images

Figure CN120998203A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a real-time punctuation recovery method based on efficient corpus screening, belonging to the technical field of natural language processing. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, speech recognition technology as an important means of human-computer interaction, its application range is increasingly wide, such as intelligent customer service, speech translation, speech input, etc. However, the current mainstream automatic speech recognition system usually directly transcribes the input speech into a text sequence without punctuation. Such a text block without punctuation is not only difficult to read, but also causes an unavoidable performance loss to the downstream natural language processing tasks, such as text classification, sentiment analysis, machine translation, etc.
[0003] In the post-processing stage of speech recognition, punctuation recovery is one of the key tasks. Punctuation recovery aims to add appropriate punctuation marks, such as commas, periods, question marks, etc. to the recognized text to improve the readability and semantic integrity of the text. Early punctuation recovery work mainly focuses on the prediction of sequence sentence position, which cannot efficiently and accurately determine the specific punctuation marks at the sequence boundary. In addition, because of the specific task requirements of speech recognition, the model is often required to implement inference in a very short time, which provides a place for real-time punctuation recovery model. Real-time punctuation recovery model has a great advantage in inference time compared to pre-training language model and large language model, so it is very suitable for post-processing operation of speech recognition.
[0004] In addition, speech recognition systems often encounter situations where the recognized sentence is not a complete sentence when facing actual speech scenarios. This makes the traditional error correction model, such as the pre-training language model BART-large-Chinese model, unsuitable for this application scenario. Although the pre-training language model BART-large-Chinese model has excellent recovery effect for complete sentences, its results are not satisfactory when multiple sentences are input. Therefore, it is particularly necessary to use a real-time punctuation recovery model. However, the existing open-source real-time punctuation recovery model does not have satisfactory results, and improving its recovery effect is a necessary work. Therefore, in order to fine-tune the existing open-source real-time punctuation recovery model, the present application proposes a real-time punctuation recovery method based on efficient corpus screening. SUMMARY
[0005] The technical problem to be solved by the present application is that the present application provides a real-time punctuation recovery method based on efficient corpus screening to solve the problem of punctuation recovery in speech recognition post-processing, and the present application has achieved good experimental results in the real-time punctuation recovery method.
[0006] The technical scheme of the present application is: a real-time punctuation recovery method based on efficient corpus screening, the specific steps of the method comprising:
[0007] Step1, based on the punctuation examples in the Chinese single language correction corpus, combined with the commonly used data sets in the Chinese single language correction corpus, generate a training set; based on the Chinese corpus after speech recognition, generate a validation set and a test set;
[0008] Step2, post-processing operation is performed on the generated training set, validation set and test set, and rule-based data cleaning is performed, and at the same time, the splicing of sentences and data sets is completed;
[0009] Step3, different real-time punctuation recovery models, different punctuation weights and different training sets are used for punctuation recovery experiments, and the punctuation recovery results are evaluated, and the training set is adjusted accordingly through the evaluation results, and Step2 is repeated;
[0010] Step4, using the trained real-time punctuation recovery model to recover the punctuation of the corpus to be labeled.
[0011] Further, the Step1 comprises:
[0012] Step1.1, select Chinese single language correction corpus including news corpus, lang8 corpus and Wudao corpus as training corpus;
[0013] Step1.2, analyze the punctuation distribution in the Chinese single language correction corpus, focus on analyzing comma, period and question mark, including analyzing the proportion of sentences ending with period and question mark;
[0014] Step1.3, in the open source Chinese single language correction corpus, analyze the punctuation distribution of each corpus, and manually select the corpus similar to the punctuation distribution in the news corpus, lang8 corpus and Wudao corpus;
[0015] According to the punctuation distribution of the corpus selected by the artificial screening, the Chinese corpus after speech recognition is selected as the validation set and the test set of the experiment respectively;
[0016] After the test set is formed, it is segmented and randomly truncated to simulate the common interruption phenomenon in speech recognition.
[0017] Further, the Step2 comprises:
[0018] Step2.1, the Chinese single language error correction corpus includes news corpus, lang8 corpus and wudao corpus, and the special symbols, data lines with non-Chinese characters and useless texts in the news corpus, lang8 corpus and wudao corpus are deleted by manual screening and rule setting, so as to remove the potential low-quality corpus in the news corpus, lang8 corpus and wudao corpus, and the obtained corpus is used as a training set;
[0019] Step2.2, part of the corpus in the training set is spliced according to a dynamically changing proportion to form a one-sentence-per-line and multi-sentence-per-line mixed dataset or a pure multi-sentence dataset;
[0020] Step2.3, for the training set of the experiment, the operation of removing the punctuation is performed to simulate the result of speech recognition;
[0021] The one-sentence-per-line and multi-sentence-per-line mixed dataset or the pure multi-sentence dataset is checked again, and it is confirmed that the training set meets the application scenario.
[0022] Further, the Step3 comprises:
[0023] Step3.1, using the CT-transformer real-time punctuation recovery model in the open source Funasr framework to perform sentence punctuation recovery experiment on the sentence without punctuation, retaining multiple experimental prediction results, and calculating the F1 value of the experimental prediction results according to the label evaluation recovery effect;
[0024] Step3.2, using the open source large model Qwen2.5 to perform sentence punctuation recovery experiment on the sentence without punctuation, retaining the experimental prediction results, and calculating the F1 value of the experimental prediction results according to the label evaluation recovery effect;
[0025] Step3.3, by using multiple different number sizes, multiple different distributions of datasets as training sets, different punctuation data weighting operations with different weights are performed multiple times, multiple experimental prediction results are retained, the F1 value of the experimental prediction results is calculated, and the weight of the data weighting, the size and the ratio of the training set are adjusted according to the results;
[0026] Step3.4, Step2.2-Step2.3, Step3.1-Step3.3 are cycled, and the dynamically changing proportion mentioned in Step2.2 is repeatedly modified until the best model is trained.
[0027] The application also provides a real-time punctuation recovery system based on efficient corpus screening, the system comprising a module for executing the real-time punctuation recovery method based on efficient corpus screening.
[0028] The application further provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the real-time punctuation recovery method based on efficient corpus screening when executing the program.
[0029] The application further provides a non-transitory computer-readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the real-time punctuation recovery method based on efficient corpus screening.
[0030] The application further provides a computer program product, comprising a computer program, wherein the computer program is executed by a processor to implement the real-time punctuation recovery method based on efficient corpus screening.
[0031] The application has the following beneficial effects:
[0032] 1. The application provides a solution to the problem of real-time punctuation recovery of existing speech recognition corpus;
[0033] 2. The application provides the operation of splicing, the optimal proportion of one sentence per line and multiple sentences per line, and a new data enhancement and corpus screening method for subsequent corpus processing;
[0034] 3. The application provides the optimal weight proportion between text and punctuation;
[0035] 4. The application effectively utilizes Chinese error correction corpus and open-source CT-transformer model, and achieves good experimental results in speech recognition post-processing tasks. Through the data enhancement method of splicing multiple sentences into one line, mixing different corpora according to different proportions, and weighting different punctuation, the application solves the problem of lack of real speech recognition corpus after processing, effectively improves the recovery effect of the model, and thus achieves better performance in punctuation recovery tasks. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 The flowchart in the application. DETAILED DESCRIPTION
[0037] Embodiment 1: As shown in the following table, a real-time punctuation recovery method based on efficient corpus screening comprises the following specific steps: Figure 1
[0038] Step 1: Based on the punctuation examples in the Chinese monolingual error correction corpus, combined with the commonly used data set in the Chinese monolingual error correction corpus, a training set is generated; based on the Chinese corpus after speech recognition, a verification set and a test set are generated;
[0039] Step 2: Post-process the generated training set, validation set, and test set, and perform rule-based data cleaning. At the same time, complete the concatenation of sentences and data sets.
[0040] Step 3: Conduct punctuation recovery experiments with different real-time punctuation recovery models, different weights for different punctuation marks, and different training sets, and evaluate the punctuation recovery results. Adjust the training set accordingly based on the evaluation results, and repeat Step 2.
[0041] Step 4: Use the trained real-time punctuation recovery model to perform punctuation recovery on the corpus to be punctuated.
[0042] Further, Step 1 includes:
[0043] Step 1.1: Select Chinese monolingual error correction corpora, including news corpora, lang8 corpora, and Wudao corpora, as training corpora;
[0044] Step 1.2: Analyze the distribution of punctuation marks in the Chinese monolingual error correction corpus, focusing on commas, periods, and question marks, including the proportion of sentences ending with periods and question marks.
[0045] Step 1.3: In the open-source Chinese monolingual error correction corpus, analyze the distribution of punctuation in each corpus, and manually select corpora with similar punctuation distribution to the news corpus, lang8 corpus, and Wudao corpus;
[0046] Based on the distribution of punctuation marks in the manually selected corpora, the Chinese corpora after speech recognition were selected as the validation set and test set for the experiment, respectively.
[0047] After forming the test set, it is segmented into words and randomly truncated to simulate the interruption phenomenon commonly seen in speech recognition.
[0048] Furthermore, Step 2 includes:
[0049] Step 2.1: The Chinese monolingual error correction corpus, including news corpus, lang8 corpus, and Wudao corpus, is manually screened and cleaned. By formulating rules, special symbols, data lines containing non-Chinese characters, and useless text in the news corpus, lang8 corpus, and Wudao corpus are deleted to remove potentially low-quality corpus from the news corpus, lang8 corpus, and Wudao corpus. The resulting corpus is used as the training set.
[0050] Step 2.2: Combine a portion of the corpus in the training set according to the dynamically changing proportions to form a dataset consisting of one sentence per line, a mix of multiple sentences per line, or a dataset consisting of only multiple sentences.
[0051] Step2.3, for the training set of the experiment, remove the punctuation, simulate the result of speech recognition;
[0052] Again, check the mixed data set of one line and one sentence and the mixed data set of one line and multiple sentences or pure multiple sentence data set, confirm that the training set meets the application scenario.
[0053] Further, the Step3 includes:
[0054] Step3.1, using the CT-transformer real-time punctuation recovery model in the open source Funasr framework to perform sentence punctuation recovery experiment on the unpunctuated sentence, retaining multiple experimental prediction results, and according to the label evaluation recovery effect, calculating the F1 value of the experimental prediction result;
[0055] Step3.2, using the open source large model Qwen2.5 to perform sentence punctuation recovery experiment on the unpunctuated sentence, retaining the experimental prediction result, and according to the label evaluation recovery effect, calculating the F1 value of the experimental prediction result;
[0056] Step3.3, by using multiple different number sizes, multiple different distribution data sets as training sets, multiple different punctuation data weighting operations with different weights are performed multiple times, multiple experimental prediction results are retained, the F1 value of the experimental prediction result is calculated, and the weight of the data weighting, the scale and the ratio of the training set are adjusted according to the result;
[0057] Step3.4, Step2.2-Step2.3, Step3.1-Step3.3 are cycled, and the dynamic change ratio mentioned in Step2.2 is repeatedly modified until the best model is trained.
[0058] The application also provides a real-time punctuation recovery system based on efficient corpus screening, which comprises a module for executing the real-time punctuation recovery method based on efficient corpus screening.
[0059] The application also provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the real-time punctuation recovery method based on efficient corpus screening.
[0060] The application also provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the real-time punctuation recovery method based on efficient corpus screening.
[0061] The application further provides a computer program product comprising a computer program which, when executed by a processor, implements the real-time punctuation recovery method based on efficient corpus screening.
[0062] Embodiment 2, a real-time punctuation recovery method based on efficient corpus screening, the method comprising:
[0063] a1, download open-source Chinese error correction corpus news2016zh, lang8 corpus, and Wudao corpus from websites such as GitHub, and based on the punctuation examples in the Chinese monolingual error correction corpus, generate a training set in combination with commonly used datasets in the Chinese monolingual error correction corpus; obtain real Chinese corpus after speech recognition from news websites, generate a validation set and a test set;
[0064] a2, data cleaning of the downloaded training corpus by rules;
[0065] 1) screen out sentences involving garbled characters, non-Chinese, etc.;
[0066] 2) screen out completely identical sentences;
[0067] a3, remove punctuation from the cleaned training corpus by rules to form a simulated speech recognition result.
[0068] a4, use the CT-Transformer model open-sourced by Alibaba to perform punctuation recovery tasks on the test set and give the prediction results. This is the best real-time punctuation recovery model among the currently open-sourced models, and the accuracy, recall rate, and F1 value of the prediction results and labels are calculated. These data will be used as the baseline of this invention. The model only takes 30 seconds to reason 5000 rows of unpunctuated sentences.
[0069] a5, use the large model Qwen2.5 open-sourced by Alibaba to perform punctuation recovery tasks on the test set and give the prediction results. Calculate the accuracy, recall rate, and F1 value of the prediction results and labels. These data show the effect of large models in real-time punctuation recovery tasks. At the same time, record the inference time of the pre-training language model BART-large-Chinese model and the large model Qwen2.5. Using the same 5000-row test set, the BART-large-Chinese model and the large model Qwen2.5 respectively need 10 minutes and 30 minutes to complete the inference.
[0070] a6, fine-tune the CT-Transformer model open-sourced by Alibaba. First, divide the corpus into a one-sentence-per-line and a mixed one-sentence-per-line dataset, and at the same time, use the same corpus to create a one-sentence-per-line dataset. From this, several datasets can be obtained as follows:
[0071] (1) 200w rows of data sets with one sentence per row and multiple sentences per row mixed, of which lang8 single sentence accounts for 60w rows; Wudao corpus (single sentence) accounts for 40w rows; news corpus (single sentence) 20w rows; Wudao corpus (multiple sentences) accounts for 40w; news corpus (multiple sentences) accounts for 40w rows.
[0072] (2) 360w rows of data sets with one sentence per row and multiple sentences per row mixed, of which lang8 single sentence accounts for 60w rows; Wudao corpus (single sentence) accounts for 40w rows; news corpus (single sentence) 20w rows; Wudao corpus (multiple sentences) accounts for 200w; news corpus (multiple sentences) accounts for 40w rows.
[0073] (3) 440w rows of data sets with one sentence per row and multiple sentences per row mixed, of which lang8 single sentence accounts for 60w rows; Wudao corpus (single sentence) accounts for 40w rows; news corpus (single sentence) 20w rows; Wudao corpus (multiple sentences) accounts for 280w; news corpus (multiple sentences) accounts for 40w rows.
[0074] (4) 514w rows of data sets with one sentence per row and multiple sentences per row mixed, of which lang8 single sentence accounts for 60w rows; Wudao corpus (single sentence) accounts for 40w rows; news corpus (single sentence) 20w rows; Wudao corpus (multiple sentences) accounts for 354w; news corpus (multiple sentences) accounts for 40w rows.
[0075] (5) 394w of data sets with all multiple sentences per row, Wudao corpus (multiple sentences) accounts for 354w; news corpus (multiple sentences) accounts for 40w rows. This data set is the multiple sentences per row part of data set (4).
[0076] (6) 1007w rows of data sets with one sentence per row and multiple sentences per row mixed, of which lang8 single sentence accounts for 60w rows; Wudao corpus (single sentence) accounts for 40w rows; news corpus (single sentence) 20w rows; Wudao corpus (multiple sentences) accounts for 354w; news corpus (multiple sentences) accounts for 533w rows.
[0077] (7) 887w of data sets with all multiple sentences per row, Wudao corpus (multiple sentences) accounts for 354w; news corpus (multiple sentences) accounts for 533w rows. This data set is the multiple sentences per row part of data set (6).
[0078] At the same time, two different validation sets were used to adapt to different training sets. The weights of text and punctuation were set to 1:1 and 1:4, respectively, to obtain multiple prediction results, and the accuracy, recall and F1 value of each prediction result and label were calculated. In this way, a better model can be obtained according to the evaluation index. At the same time, the inference time of the real-time punctuation restoration model is recorded. Using the same 5000-line test set, the real-time punctuation restoration model before and after fine-tuning only needs 30 seconds to end the inference. It can be seen that the real-time punctuation restoration model has a huge advantage over the pre-training language model and the large language model in inference time, and long inference time is unacceptable in this task.
[0079] a7, verify whether the inference of punctuation restoration meets the requirements of time and accuracy.
[0080] The F1 value commonly used in the field of error correction is used as the evaluation index of the model. The specific calculation method is shown in formula (1):
[0081]
[0082] Among them, the definitions of Precision (precision) and Recall (recall) are shown in formula (2) and formula (3):
[0083] Precision (Precision): represents the proportion of truly positive samples in the samples predicted by the model as positive, and the calculation formula is:
[0084]
[0085] Recall (Recall): represents the proportion of samples correctly predicted by the model as positive among all truly positive samples, and the calculation formula is:
[0086] Among them:
[0087] TP (True Positive): True Positive, the number of samples correctly predicted by the model as positive.
[0088] FP (False Positive): False Positive, the number of samples incorrectly predicted by the model as positive.
[0089] FN (False Negative): False Negative, the number of samples incorrectly predicted by the model as negative.
[0090] 1) In order to verify the effect of the method proposed in the present application on the improvement of Chinese punctuation restoration, the present application selects a real-time punctuation restoration model as a benchmark model. The sequence-to-sequence model adopts the most commonly used Transformer structure, and the goal is to add punctuation to the sentence without adding punctuation in real time. In this verification, the present application fine-tunes the CT-transformer model for punctuation restoration.
[0091] The present application first constructs a training set through an open-source Chinese error correction data set, constructs a verification set and a test set through real speech recognition results; secondly, it uses rules to clean up the corpus, selects high-quality corpus from it, and realizes the simulation of the speech recognition post-processing results by the training set through truncation and other methods; finally, through the fine-tuning of the CT-transformer model by comparing multiple data sets and data weighting methods, the best data enhancement method for punctuation restoration accuracy is found.
[0092] The final experimental results are shown in Table 1. On the test set used by the real speech recognition, the pre-trained language model and the large language model do not perform well, and their inference time is too long, so only their results are used as a reference. The specific analysis is as follows:
[0093] Firstly, comparing the data in the first row with the data in the fourth, fifth and sixth rows, simply relying on the use of data sets for fine-tuning cannot guarantee the improvement of punctuation restoration effect. Secondly, comparing the data in the fourth, fifth and sixth rows with the data in the seventh, eighth, ninth and tenth rows, it can be seen that with the increase of the size of the data set, the punctuation restoration effect of the model improves. Finally, comparing the data in the seventh row with the data in the ninth row and the data in the eighth row with the data in the tenth row, it is found that using a data set with all one-line multiple sentences for fine-tuning can achieve better punctuation restoration effect than using a data set with a mixture of one-line single sentences and one-line multiple sentences. Moreover, the data set with all one-line multiple sentences is smaller in size, but it can achieve better real-time punctuation restoration effect. In the past, when fine-tuning real-time punctuation restoration models, a data set with a mixture of one-line single sentences and one-line multiple sentences was often used, but according to the present research, by using an efficient corpus screening method, the model effect can be significantly improved.
[0094] For the existing problem of real-time punctuation recovery of speech recognition corpus, the application proposes an effective solution. Through in-depth analysis of the performance of different data sets on the sequence-to-sequence model, the application finds that simply relying on data set fine-tuning cannot necessarily improve punctuation recovery effect. However, when the size of the data set increases, the punctuation recovery effect of the model improves. Further research finds that using a data set of all one-line multi-sentence for fine-tuning can not only achieve better punctuation recovery effect, but also requires smaller data set size compared to using a mixed data set of one-line one-sentence and one-line multi-sentence. This finding shows that through an efficient corpus screening method, the punctuation recovery performance of the model can be significantly improved under limited corpus conditions. Therefore, the application optimizes the data set structure and fine-tuning strategy to provide an innovative solution to the existing problem of real-time punctuation recovery of speech recognition corpus, effectively improving the performance of the model in real-time punctuation recovery tasks.
[0095] Table 1 Performance of different data sets on sequence-to-sequence model
[0096]
[0097] a8, using the punctuation recovery model obtained by using the 1007w one-line one-sentence and one-line multi-sentence mixed data set for fine-tuning and data weighting to perform punctuation recovery on the corpus to be numbered, realizes the application of the application.
[0098] The specific embodiments of the application are described in detail above in conjunction with the accompanying drawings, but the application is not limited to the above embodiments, and various changes can be made within the knowledge of those skilled in the art without departing from the spirit of the application.
Claims
1. A real-time punctuation restoration method based on efficient corpus screening, characterized in that: Specific steps of the method include: Step1, based on the punctuation examples in the Chinese single language correction corpus, combined with the commonly used data set in the Chinese single language correction corpus, generate the training set; based on the Chinese corpus after voice recognition, generate the validation set and test set; Step2, post-processing operation is performed on the generated training set, validation set and test set, and rule-based data cleaning is performed, and at the same time, the splicing of sentences and data sets is completed; Step3, different real-time punctuation restoration models, different punctuation and different training sets are used for punctuation restoration experiments, and the punctuation restoration results are evaluated, and the training set is adjusted accordingly through the evaluation results, and Step2 is repeated; Step4, using the trained real-time punctuation restoration model to restore the punctuation of the corpus to be restored.
2. The method of claim 1, wherein the method is based on efficient corpus filtering. The Step1 includes: Step1.1, select Chinese single language correction corpus including news corpus, lang8 corpus and Wudao corpus as training corpus; Step1.2, analyze the punctuation distribution in the Chinese single language correction corpus, focus on analyzing comma, period and question mark, including analyzing the proportion of sentences ending with period and question mark; Step1.3, in the open source Chinese single language correction corpus, analyze the punctuation distribution of each corpus, and manually select the corpus similar to the punctuation distribution in the news corpus, lang8 corpus and Wudao corpus; According to the punctuation distribution of the corpus selected by manual screening, the Chinese corpus after voice recognition is selected as the validation set and test set of the experiment respectively; After the test set is formed, it is segmented and randomly truncated to simulate the common interruption phenomenon in voice recognition.
3. The method of claim 1, wherein the method is characterized by: The Step2 includes: Step2.1, manually screen the Chinese single language correction corpus including news corpus, lang8 corpus and Wudao corpus, clean it, delete special symbols, data lines with non-Chinese characters and useless text in news corpus, lang8 corpus and Wudao corpus by formulating rules, which is used to remove potential low-quality corpus in news corpus, lang8 corpus and Wudao corpus, and the obtained corpus is used as the training set; Step2.2, part of the training set is spliced according to the dynamically changing proportion to form a mixed data set of one sentence per line and multiple sentences per line or a pure multiple sentence data set; Step2.3, for the training set of the experiment, remove the punctuation to simulate the result of voice recognition; Again, check the mixed data set of one sentence per line and multiple sentences per line or the pure multiple sentence data set to confirm that the training set meets the application scenario.
4. The method of claim 3, wherein the method further comprises: The Step3 includes: Step3.1, use the CT-transformer real-time punctuation restoration model in the open source Funasr framework to perform sentence punctuation restoration experiment on the unpunctuated sentence, retain multiple experimental prediction results, and calculate the F1 value of the experimental prediction results according to the label evaluation of the restoration effect; Step3.2, using open source large model Qwen2.5 to restore the sentence punctuation of the sentence without punctuation, retaining the experimental prediction results, and according to the label evaluation to restore the effect, calculating the F1 value of the experimental prediction results; Step3.3, through multiple experiments on different weight values of different punctuation on multiple different size and distribution of data sets as training set, retaining multiple experimental prediction results, calculating the F1 value of the experimental prediction results, adjusting the weight value of the data weighting operation, the scale and proportion of the training set according to the results; Step3.4, Step2.2-Step2.3, Step3.1-Step3.3 are cycled, and the dynamic change ratio mentioned in Step2.2 is repeatedly modified until the best model is trained.
5. A real-time punctuation restoration system based on efficient corpus screening, characterized in that, The system comprises a module for performing a real-time punctuation restoration method based on efficient corpus screening according to any one of claims 1 to 4.
6. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to realize the real-time punctuation restoration method based on efficient corpus screening according to any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the real-time punctuation restoration method based on efficient corpus screening according to any one of claims 1 to 4.
8. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to realize the real-time punctuation restoration method based on efficient corpus screening according to any one of claims 1 to 4.
Citation Information
Patent Citations
Text punctuation recovery method based on pre-training fused speech features
CN115017883A
Automated medical report formatting system
US20190065462A1