Text Punctuation Detection Method, Computer Device, and Storage Medium

By training the target training samples of the preset language model, a target language model that integrates context information and part of speech is formed, which solves the problem of low accuracy in text punctuation detection in the prior art, and achieves more accurate detection of sentence breaking errors.

CN114298032BActive Publication Date: 2025-06-17IFLYTEK CO LTD +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111547437.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-16
Publication Date
2025-06-17
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

The prior art has low accuracy in text punctuation detection, making it difficult to effectively solve semantic errors. The clause segmentation technology can only solve the misuse errors of periods or commas, and cannot handle redundant or missing errors.

Method used

By training the preset language model based on the target training sample, the context information of characters and the network layer of part-of-speech is fused to form a target language model, which is used to analyze the text to be identified and generate a sequence of punctuation labels, thereby performing punctuation detection.

Benefits of technology

It improves the accuracy of text punctuation detection, can effectively detect redundancy and missing errors of periods or commas in sentence breaking errors, and improves the ability to identify text punctuation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114298032B_ABST
    Figure CN114298032B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of language processing, and discloses a text punctuation detection method, a computer device, and a storage medium. The method includes: obtaining a text to be recognized, and inputting the text to be recognized into a pre-trained target language model, where the target language model is a network layer that fuses context information and part-of-speech information for analyzing characters in the text after training a preset language model based on target training samples, and the target training samples are text data obtained by correcting punctuation for text data based on a back-translation data augmentation strategy; analyzing the context information and part-of-speech information of the characters in the text to be recognized based on the target language model to obtain a punctuation label sequence of the text to be recognized; and performing punctuation detection on the text to be recognized based on the punctuation label sequence. The aim is to improve the accuracy of text punctuation detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of language processing technologies, and in particular, to a method for detecting text punctuation, a computer device, and a storage medium. Background Art

[0002] With the continuous development of language processing technologies, language processing technologies have been widely applied to text punctuation recognition. In the prior art, commonly used text punctuation recognition technologies include rule-based punctuation error detection technologies, clause segmentation technologies, and punctuation restoration technologies. Among them, the punctuation error detection technology can only detect normative punctuation errors such as continuous use of punctuation marks and mismatched paired punctuation marks, and it is difficult to effectively solve the sentence segmentation errors involving semantics. Compared with the punctuation error detection technology, the clause segmentation technology can be applied to the detection of sentence segmentation errors, but the clause segmentation technology can only solve the misuse errors of full stops or commas in sentence segmentation errors, and cannot solve the redundancy errors and missing errors of full stops or commas. The punctuation restoration technology is proposed to address the problems existing in the clause segmentation technology and can handle sentence segmentation errors caused by misuse, redundancy, or missing of full stops or commas. However, since the punctuation restoration technology mainly relies on a pre-trained language model, on the one hand, it is greatly affected by the quality of training samples, and on the other hand, it often ignores the influence of word information in clauses in the text on punctuation marks, resulting in the inability of the punctuation restoration technology to accurately detect punctuation errors in the text.

[0003] Therefore, the commonly used language processing technologies in the prior art still have the problem of low accuracy in detecting text punctuation. Summary of the Invention

[0004] This application provides a method for detecting text punctuation, a computer device, and a storage medium. By training a preset language model that integrates a network layer for analyzing the context information and part-of-speech of characters in the text based on target training samples, and performing text punctuation detection using the trained target language model, the accuracy of text punctuation detection is improved.

[0005] In a first aspect, this application provides a method for detecting text punctuation, and the method includes:

[0006] Obtain the text to be recognized, and input the text to be recognized into a pre-trained target language model, where the target language model is a preset language model that has been trained based on target training samples and integrates a network layer for analyzing the context information and part-of-speech of characters in the text, and the target training samples are text data obtained by correcting the punctuation of text data based on a back-translation data augmentation strategy;

[0007] Analyze the context information and part-of-speech of the characters in the text to be recognized based on the target language model, and obtain the punctuation label sequence of the text to be recognized;

[0008] Perform punctuation detection on the text to be recognized based on the punctuation label sequence.

[0009] In a second aspect, the present application further provides a computer device, including:

[0010] A memory and a processor;

[0011] The memory is used to store a computer program;

[0012] The processor is configured to execute the computer program and, when executing the computer program, implement the steps of the text punctuation detection method described in the first aspect above.

[0013] In a third aspect, the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to implement the steps of the text punctuation detection method described in the first aspect above.

[0014] The present application discloses a text punctuation detection method, a computer device, and a storage medium. First, analyze the context information and part-of-speech of characters in the text to be recognized through a target language model to obtain the punctuation label sequence of the text to be recognized; then perform punctuation detection on the text to be recognized based on the punctuation label sequence. Since the target language model is a network layer that fuses the context information and part-of-speech for analyzing characters in the text after training a preset language model based on target training samples, and the corresponding target training samples are text data obtained by punctuation correction of text data based on a back-translation data augmentation strategy, therefore, through the target language model, the context information and part-of-speech of characters in the text to be recognized can be analyzed, and then the punctuation label sequence of the text to be recognized can be predicted according to the context information and part-of-speech of characters in the text to be recognized, improving the accuracy of text punctuation detection. Description of the Drawings

[0015] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0016] Figure 1 It is a schematic diagram of the application scenario architecture of the text punctuation detection method disclosed in an embodiment of the present application;

[0017] Figure 2 It is a schematic diagram of the structure of the text punctuation detection system provided by the embodiments of the present application;

[0018] Figure 3 It is a schematic diagram of an application scenario of the text punctuation detection method provided by another embodiment of the present application;

[0019] Figure 4 It is a schematic diagram of the implementation process of the text punctuation detection method provided by an embodiment of the present application;

[0020] Figure 5 It is a schematic diagram of the structure of the preset language model provided by an embodiment of the present application;

[0021] Figure 6 It is Figure 4 The specific implementation flowchart of S402 in

[0022] Figure 7 It is Figure 4 The specific implementation flowchart of S404 in

[0023] Figure 8 It is a schematic diagram of the implementation process of the text punctuation detection method provided by another embodiment of the present application;

[0024] Figure 9 It is a schematic diagram of the structure of the computer device provided by an embodiment of the present application. Detailed implementation manners

[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0026] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged, so the actual execution order may be changed according to the actual situation.

[0027] It should be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0028] It should also be understood that the term " / and" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.

[0029] The embodiments of the present application provide a text punctuation detection method, a computer device and a storage medium. The text punctuation detection method provided in the embodiments of the present application can be used to detect punctuation points on a text, and can improve the accuracy of text punctuation detection.

[0030] Before describing the text punctuation detection method, computer device and storage medium provided by the embodiment of the present application, the problems existing in the prior art are explained exemplarily. It should be understood that sentence segmentation errors are a type of punctuation errors that often occur in articles, especially student compositions. For example, when primary school students learn to write essays, they often ignore the correct use of punctuation marks such as periods and commas, which makes the syntactic structure of clauses unclear and the semantic logic confusing, seriously affecting the expression of the composition content and reducing the overall quality of the composition.

[0031] Generally, punctuation errors can be divided into the following three types of errors: the first is the misuse of commas, that is, a comma is used where a period should be used, or a period is used where a comma should be used; the second is redundant commas, that is, a period or comma is used where a period or comma should not be used; the third is missing commas, that is, a period or comma is not used where a period or comma should be used. Examples of these three types of errors are shown in Table 1:

[0032]

[0033] Table 1 Sentence segmentation error example table

[0034] At present, there are three common natural language processing technologies: one is rule-based punctuation error detection technology; the second is sentence segmentation technology; and the third is punctuation restoration technology.

[0035] Among them, rule-based punctuation error detection technology usually relies on manually designed punctuation error patterns and uses regular matching technology to match the rules of the text to be detected. If the text to be detected can match the manually designed error pattern, it can be considered that the text to be detected contains punctuation errors.

[0036] Clause segmentation technology is a fundamental technology in the field of natural language processing. Its purpose is to segment the text of a paragraph into clause fragments that are relatively complete syntactically and semantically, thus facilitating other more advanced analysis and processing of the text, such as word segmentation, part-of-speech tagging, named entity recognition, and syntactic analysis. In English, the punctuation marks for clause segmentation usually have ambiguity. For example, the English full stop "." can be both the boundary of a clause and used in abbreviations, such as "U.S.". Therefore, traditional rule-based clause segmentation technology is difficult to perfectly solve the clause segmentation problem. Some researchers have started to model the clause segmentation problem as a binary classification problem and use machine learning technology and deep learning technology to predict whether the punctuation marks that may be clause boundaries in the text are real clause boundaries. In Chinese, the punctuation marks serving as clause boundaries basically do not have ambiguity. However, in many non-standard Chinese texts, the use of full stops and commas is not paid much attention to, and the phenomenon of "using only commas throughout" often occurs. These problems will also have a certain impact on other advanced natural language processing tasks such as syntactic analysis, machine translation, and discourse analysis. Therefore, clause segmentation technology also has a very broad application space in the field of Chinese natural processing, and some machine learning algorithms such as decision trees, support vector machines, Bayesian classification, and maximum entropy classification have also been successively applied by researchers.

[0037] Punctuation restoration technology is a post-processing technology often used in the field of speech recognition. Its purpose is to add the correct punctuation marks to the text string without any punctuation after speech recognition, so that the text after speech recognition is easier for humans to read. The punctuation restoration task is usually regarded as a sequence labeling task, that is, to predict the punctuation marks to be added to each character in the text string. If no punctuation mark needs to be added, only a special marker needs to be predicted to indicate that no punctuation mark needs to be added. With the development of pre-trained language model technology, current punctuation restoration technology usually relies on using large-scale pre-trained language models such as BERT and ELECTRA for sequence labeling to achieve punctuation restoration.

[0038] It should be understood that sentence segmentation error detection is a natural language processing task at the semantic level. Due to the complexity and ambiguity of natural language, rule-based punctuation error detection technology is difficult to effectively solve the sentence segmentation errors involving semantics. Therefore, existing punctuation error detection systems can usually only detect normative punctuation errors such as continuous use of punctuation marks and mismatched paired punctuation marks, greatly limiting the capabilities and application scenarios of punctuation error detection systems.

[0039] Although clause segmentation technology can be applied to sentence segmentation error detection, it can only solve the misuse of commas in sentence segmentation errors, but cannot solve comma redundancy errors and comma missing errors. In addition, training a clause segmentation binary classifier requires manually annotated training data, that is, the training data should mark which periods and commas are used correctly and which are used incorrectly. These factors limit the application of clause segmentation technology in sentence segmentation error detection tasks to a certain extent.

[0040] Punctuation recovery technology can handle three types of punctuation errors: misuse of punctuation, redundant punctuation, and missing punctuation. However, there are still two challenges in applying punctuation recovery technology to article punctuation, especially in the task of detecting punctuation errors in primary school essays: first, there is a lack of training data suitable for the field of primary school essays that can be used to train punctuation recovery models; second, existing punctuation recovery technology only relies on the use of pre-trained language models, ignoring the use of clause information, part of speech information, and other text information, which is very valuable for the task of punctuation error detection. Through clause position information, the model can learn the relationship between the position and length of the clause and the use of punctuation marks. For example, a shorter clause at the beginning of a paragraph is more likely to be followed by a comma. Through part of speech information, the model can learn the relationship between part of speech and the use of punctuation marks. For example, there is a higher probability of using a period before a personal pronoun and a higher probability of using a comma after a conjunction.

[0041] In view of the above points, we proposed a sentence segmentation error detection method with punctuation recovery technology as the core. On the one hand, we used back-translation technology to correct some punctuation in the text and obtained a batch of high-quality punctuation data in the text field. On the other hand, in order to strengthen the use of text information, we designed a punctuation recovery model that integrates clause information and part-of-speech tagging.

[0042] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0043] In order to better understand the text punctuation detection method, computer device and storage medium disclosed in the embodiments of the present application, the following first combines Figure 1 The application scenario and scenario architecture of the text punctuation detection method provided in the embodiments of the present application are described exemplarily.

[0044] See also Figure 1 , Figure 1 Schematic diagram of the application scenario architecture of the text punctuation detection method disclosed in an embodiment of the present application. Figure 1As shown in the figure, the text punctuation detection method can be applied to a computer device 10, in which a text punctuation detection system 11 is integrated. Among them, the computer device 10 can be a server or a terminal device. The server can be a remote server, a cloud server, or a server cluster, etc., which can be used to run the text punctuation detection system 11. The terminal device can be a personal computer, a notebook, a PAD, a robot, or a handheld intelligent terminal device with certain computing capabilities. The terminal device can also be used to run the text punctuation detection system 11. The text punctuation detection system 11 is an application program integrated on the computer device 10 and having the function of text punctuation detection.

[0045] It should be understood that the text punctuation detection method described in the embodiments of the present application can be applied to all application scenarios where a text punctuation detection system 11 is integrated in a computer device 10 and the text is punctuated by the text punctuation detection system 11.

[0046] Exemplarily, please refer to Figure 2 as shown in the figure, Figure 2 is a schematic structural diagram of the text punctuation detection system provided by the embodiments of the present application. It can be seen from Figure 2 that the text punctuation detection system 11 includes:

[0047] An acquisition module 111, configured to acquire a preset number of target training samples, where the target training samples are text data obtained by punctuating the text data based on a back-translation data augmentation strategy;

[0048] A training module 112, configured to train a preset language model based on the target training samples to obtain a target language model, where the preset language model integrates a network layer for analyzing the context information and part-of-speech of characters in the text;

[0049] An analysis module 113, configured to analyze the context information and part-of-speech of characters in the text to be recognized based on the target language model to obtain a punctuation label sequence of the text to be recognized;

[0050] A recognition module 114, configured to perform punctuation detection on the text to be recognized based on the punctuation label sequence.

[0051] Due to the functions of the above modules, the text punctuation detection system provided by the embodiments of the present application can improve the accuracy of text punctuation detection, and further can be applied to the punctuation detection of articles, such as student compositions, to achieve the accuracy of composition punctuation detection.

[0052] In addition, the text punctuation detection system 11 may be composed of multiple subsystems. For example, the text punctuation system 11 includes a training subsystem 101 and a detection subsystem 102. It should be understood that both the training subsystem 101 and the detection subsystem 102 may be integrated into a computer program. In the computer device 10, the text punctuation detection system 11 is used as an application program to complete the detection of text punctuation. Alternatively, in the computer device 10, the training subsystem 101 and the detection subsystem 102 are integrated into the same application program to complete the detection of text punctuation.

[0053] It should be understood that the training subsystem 101 and the detection subsystem 102 may also be integrated into the computer device 10 as two different application programs respectively. The computer device 10 completes its respective functions by calling the corresponding application programs respectively. Exemplarily, the computer device 10 completes the training of the target language model by calling the application program corresponding to the training subsystem 101, and the computer device 10 completes the punctuation detection of the text to be recognized based on the target language model by calling the application program corresponding to the detection subsystem 102.

[0054] It should be understood that when the computer device 10 has high computing power, such as a server or a server cluster, the corresponding training subsystem 101 and detection subsystem 102 may both be integrated into the computer device 10. When the computer device 10 is a terminal device with limited computing power, such as a handheld intelligent terminal device, it may be considered to only integrate the test subsystem 102 in the terminal device, and deploy the training subsystem 101 in the cloud that is communicatively connected to the terminal device.

[0055] Exemplarily, as Figure 3 shown, Figure 3 is a schematic diagram of an application scenario of the text punctuation detection method provided by another embodiment of the present application. In Figure 3 it, the text punctuation detection method is jointly implemented by the cloud platform 20 and the computer device 10. It should be noted that the cloud platform 20 includes a cloud data center and a cloud service platform (not shown in the figure). The cloud data center includes a large number of basic resources owned by the cloud service provider, such as a preset language model; the computing resources included in the cloud data center may be a large number of computer devices, such as servers or server clusters.

[0056] In this embodiment, a preset language model is deployed on the cloud service platform, and a training subsystem is deployed on the cloud data center. The training subsystem can train the preset language model on the cloud service platform to obtain a target language model to ensure the training efficiency of the target language model.

[0057] In some embodiments, the training subsystem is used to train a preset language model in the cloud service platform according to target training samples to obtain a target language model. The target training samples are text data obtained by correcting punctuation marks in the text data based on the back-translation data augmentation strategy. Among the existing text data, the punctuation quality of some general domain data such as People's Daily, Wikipedia, and Baidu Encyclopedia is relatively high. However, for the articles written by writers, such as students' compositions, due to the uneven quality of the articles and the lack of data suitable for training the punctuation restoration model, the performance of the punctuation restoration model in the writing field decreases. To solve this problem, in this case, a domain data augmentation strategy based on the back-translation technology is used to correct some punctuation marks in the article to obtain target training samples, which can improve the quality of training data in the writing field.

[0058] Please refer to Figure 4 as shown Figure 4 FIG. is a schematic flowchart of the implementation of the text punctuation detection method provided by an embodiment of the present application. The text punctuation detection method provided by this embodiment can be completed by Figure 1 the computer device 10 shown as follows:

[0059] S401, obtain the text to be recognized, and input the text to be recognized into the pre-trained target language model.

[0060] The target language model is a network layer that integrates the context information and part-of-speech of characters in the text after training the preset language model based on the target training samples. The target training samples are text data obtained by correcting punctuation marks in the text data based on the back-translation data augmentation strategy.

[0061] The back-translation data augmentation strategy is a technique for correcting punctuation marks in text data by translating the text data through a translation engine and using different language punctuation annotation standards. For example, the text data can be input into a Chinese-to-English translation engine to obtain the corresponding English of the text data, and then the English-to-Chinese translation engine is used to translate the corresponding English back into Chinese. Through the back-translation of the text data, the punctuation correction of the text data is completed. Since the syntactic structure of English is relatively clear compared to that of Chinese and the punctuation usage rules are relatively simple, in English text, a full stop is usually added after a text fragment with a complete subject-predicate-object structure, and the phenomenon of only using commas, which is relatively common in Chinese, rarely appears. Therefore, first translating the text data into English and correcting some punctuation marks in the text can improve the quality of the text data. Especially for the problem of multiple commas in compositions, it can be effectively corrected to improve the quality of composition data.

[0062] Exemplarily, punctuation correction of text data is performed based on a back-translation data augmentation strategy, including: splitting the text data into paragraphs to obtain at least one first paragraph; respectively inputting each of the first paragraphs into a first translation engine to obtain the English corresponding to each of the first paragraphs; respectively inputting the English corresponding to each of the first paragraphs into a second translation engine to obtain the corrected text data for each of the first paragraphs; wherein, the corrected text data is the text data obtained after punctuation correction of each of the first paragraphs.

[0063] It should be understood that the first translation engine is an English translation engine from Chinese, and the second translation engine is a Chinese translation engine from English. After punctuation correction of text data through the back-translation data augmentation strategy, some punctuation in the text can be corrected and the semantics of the text basically remains unchanged. Using the target text data obtained after punctuation correction of text data based on the back-translation data augmentation strategy as a training sample can effectively improve the accuracy of the target language model in punctuation correction of text data. Especially in the process of punctuation correction of compositions, after punctuation correction of the original composition based on the back-translation data augmentation strategy, the trained target language model for punctuation correction of compositions can effectively correct the punctuation in the compositions, can be applied to the automatic correction of students' compositions, assist students in discovering sentence-breaking errors in the compositions, enhance students' awareness and ability to use punctuation marks, and thus help students write clearer and more fluent compositions. In addition, this technology can also be applied to the automatic scoring of students' compositions, provide valuable features for the automatic scoring of compositions, and make the results of automatic scoring of compositions more reasonable.

[0064] It should be understood that before inputting the text to be recognized into the pre-trained target language model, it is necessary to train a preset language model based on the target training sample to obtain a target language model, wherein the preset language model incorporates a network layer for analyzing the context information and part-of-speech of characters in the text.

[0065] Exemplarily, in the embodiments of the present application, the preset language model is a punctuation restoration model for joint part-of-speech tagging that integrates clause information. This model consists of three network layers (which can also be referred to as sub-modules), namely: the first network layer, the second network layer, and the third network layer. Among them, the first network layer is a character representation layer that integrates clause information. This first network layer can encode each character in the text string into a low-dimensional dense vector and integrate the clause information to which each character belongs into the vector of each character for representation; the second network layer is a context representation layer based on the pre-trained language model BERT. This second network layer can be composed of 12 Transformer encoders. The parameters of these encoders are initialized with pre-trained parameters and fine-tuned during the training of the punctuation restoration task; the third network layer is a punctuation prediction module. This third network layer is stacked by 2 linear layers. The bottom linear layer performs the part-of-speech tagging task, and the upper linear layer performs the punctuation restoration task.

[0066] Exemplarily, as Figure 5 shown, Figure 5 is a schematic structural diagram of the preset language model provided by the embodiments of the present application. It can be seen from Figure 5 that the first network layer 501 is used to represent each character in the target training sample 504 as a vector to obtain the first vector corresponding to each character; the second network layer 502 is used to analyze the first vector to obtain the second vector representing the context information of each character; the third network layer 503 is used to predict the punctuation label after the character corresponding to each context information according to the part-of-speech label of each context information in the second vector to obtain a punctuation label sequence.

[0067] It should be noted that Figure 5 the clause position encoding in represents the position of the clause where a certain character is located in the entire paragraph. Among them, a clause refers to a text segment separated by a comma. For example, in the text segment "is the darling of the new era, therefore,...", the first clause included is "is the darling of the new era"; the second clause is "therefore". In this text segment, the clause position encoding of all characters in the first clause is 0. The clause position encoding of all characters in the second clause is 1. It should be understood that if the above text segment also includes a third clause, then the clause position encoding of all characters in the third clause is 2, and so on, the clause position encoding of all characters in all clauses in the text segment can be represented. The position encoding refers to encoding the characters in the text segment in order. For example, the position encoding of the first character is 0, the position encoding of the second character is 1, and then it increases sequentially.

[0068] Furthermore, as Figure 5As shown, it should be noted that the v, n, nn, u, n, n, p, p output by the part-of-speech prediction layer represent the part of speech of the word where each character is located. Among them, V represents a verb, n represents a noun, u represents a particle, etc. The 0, 0, 0, 0, 0, 0, 0E, 0 after the punctuation prediction layer represents the punctuation mark that needs to be added after each character. For example, O means no punctuation mark needs to be added, and E means a full stop needs to be added, etc.

[0069] It should be understood that during the training process of the preset language model, a parameter is required to evaluate the performance of the model, and then determine whether the training of the preset language model is completed. In this embodiment, a loss function is used to evaluate the performance of the model. Specifically, the preset language model further includes an output layer and a loss function.

[0070] Exemplarily, as Figure 6 shown Figure 6 is the implementation flowchart of the target language model training provided by the embodiments of the present application. As Figure 6 can be seen, the process of training the preset language model based on the target training sample to obtain the target language model includes S4021 to S4024. Details are as follows:

[0071] S4021, input the target training sample into the first network layer for analysis to obtain each of the first vectors.

[0072] Among them, the first network layer includes a character embedding layer, a character position embedding layer, and a clause position embedding layer; the inputting the target training sample into the first network layer for analysis to obtain each of the first vectors includes: performing paragraph splitting processing on the target training sample to obtain a plurality of second paragraphs; for any one of the second paragraphs, removing the preset type of punctuation in the second paragraph to obtain the string of the second paragraph; inputting the string into the character embedding layer for analysis to obtain the character information of each character in the string; inputting the string into the character position embedding layer for analysis to obtain the first position information of each character in the second paragraph in the second paragraph; inputting the string into the clause position embedding layer for analysis to obtain the second position information of the clause to which each character belongs in the second paragraph in the second paragraph; generating the first vector of each character in the string based on the character information, the first position information, and the second position information.

[0073] Among them, the character information of each character in the string represents the semantic information of each character itself in the corresponding string.

[0074] It should be understood that the first network layer is used to represent the characters that fuse clause information, and its input is a text string without full stops and commas in paragraphs. For each character in the text string, each character is represented by using the information of the character itself, the position information of the character in the paragraph, and the position information of the clause where the character is located. For the information of the character itself and the position information of the character in the paragraph, the character embedding layer E char and the character position embedding layer E char-pos in the pre-trained language model BERT can be used to represent each character. For the position information of the clause where the character is located, we use a randomly initialized clause position embedding layer E clause-pos to encode the position information of the clause where the character is located. The character embedding layer, the character position embedding layer, and the clause position embedding layer where the character is located will be continuously optimized as the entire model is trained. For any character c i in the paragraph, its final vector representation is calculated as shown in the following formula:

[0075]

[0076] where E char (c i ) represents the encoded information obtained by encoding the information of the i-th character itself, and E char-pos (c i ) represents the encoded information obtained by encoding the position information of the i-th character in the text string, and E clause-pos (c i ) represents the encoded information obtained by encoding the position information of the clause where the i-th character is located.

[0077] S4022. Input each of the first vectors into the second network layer for analysis to obtain each of the second vectors.

[0078] Among them, the second network layer is an encoding layer based on the multi-head attention mechanism; exemplarily, inputting each of the first vectors into the second network layer for analysis to obtain each of the second vectors includes: for each of the first vectors, determining each attention weight of the first vector through the multi-head attention mechanism, and each of the attention weights is the attention weight between the character corresponding to the first vector and other characters in its affiliated string; performing a weighted sum on each of the attention weights to obtain the context information of the character corresponding to the first vector in its affiliated string; generating a corresponding second vector based on the context information.

[0079] It should be understood that the second network layer is an input of the context representation module based on the pre-trained language model BERT, which is the vector representation of each character The vector representation of each character obtains the vector representation of each character with context information through the multi-head self-attention mechanism. In the embodiments of the present application, through the multi-head self-attention mechanism, each character can learn the attention weights of all characters in the context, and perform weighted summation on the vectors of all characters in the context according to these attention weights, so as to obtain the vector representation of each character with context information. Its calculation process can be shown as the following formula:

[0080]

[0081] Wherein, represents the vector of the i-th character, represents the information of the i-th character in the current context.

[0082] S4023. Input each of the second vectors into the third network layer for analysis, output a punctuation label sequence through the output layer, and detect the value of the loss function.

[0083] Wherein, the third network layer includes a part-of-speech prediction layer and a punctuation prediction layer; exemplarily, inputting each of the second vectors into the third network layer for analysis includes: predicting the part of speech of each context information in the second vector based on the part-of-speech prediction layer to generate a part-of-speech label; analyzing the part-of-speech label based on the punctuation prediction layer to predict the punctuation label after the character corresponding to each context information, so as to obtain a punctuation label sequence.

[0084] It should be understood that the third network layer is a hierarchical punctuation prediction layer, and the corresponding input is the vector representation of the context information of each character. In the embodiments of the present application, specifically, 2 stacked linear layers Linear pos and Linear punc are used to sequentially complete the prediction of the part-of-speech tagging task and the punctuation restoration task from bottom to top when training a preset language model. Exemplarily, the prediction process of the part-of-speech tagging task can be represented by the following formula:

[0085]

[0086] The prediction process of the punctuation restoration task can be represented by the following formula:

[0087]

[0088] It should be understood that when performing the part-of-speech tagging task prediction, it can be assumed that the tag set of "BI + part-of-speech tag" is used. Among them, B represents the first character in a word, and I represents other characters except the first character. For example, if 4 consecutive characters in the text string can form a noun (correspondingly, the part-of-speech tag of the noun is denoted as "n"), the tags of these 4 characters can be expressed as "B-n", "I-n", "I-n", "I-n" in sequence. It should be understood that the part-of-speech tags of characters are output by the part-of-speech prediction layer.

[0089] When performing the punctuation restoration task, it is assumed that the tag set used is "MEO". If a comma needs to be added after a character, the tag of this character is "M". If a period needs to be added after a character, the tag of this character is "E". If no period or comma needs to be added after a character, the tag of this character is "O".

[0090] S4024, if the value of the loss function is less than the preset threshold, stop training the preset language model to obtain the target language model.

[0091] Exemplarily, the loss function can be a cross-entropy (CE) loss function based on multi-task learning. The calculation process of this loss function is as follows:

[0092]

[0093] Among them, λ is a hyperparameter that controls the importance of losses of different tasks denotes and the cross-entropy loss value, denotes and the cross-entropy loss value of, denotes the predicted part of speech of the i-th word, denotes the actual part of speech of the i-th word, denotes the predicted punctuation to be restored after the i-th word, denotes the actual punctuation to be restored after the i-th word.

[0094] It should be understood that when the value of the cross-entropy loss function is less than the preset threshold, such as 0.3, it means that the probability that the predicted part of speech is the same as the actual part of speech, and the predicted punctuation to be restored is the same as the actual punctuation to be restored is close to 99.7%, that is to say, the prediction probability of the model is close to 100%. Therefore, it can be determined that the training of the preset language model is completed based on the value of the cross-entropy loss function, and a target language model with high punctuation prediction accuracy is obtained.

[0095] S402. Analyze the context information and part-of-speech of the characters in the text to be recognized based on the target language model, and obtain the punctuation label sequence of the text to be recognized.

[0096] Exemplarily, analyze the text to be recognized based on the target language model, and obtain the punctuation label sequence of the text to be recognized predicted by the target language model.

[0097] S403. Perform punctuation detection on the text to be recognized based on the punctuation label sequence.

[0098] Compare the punctuation label sequence predicted by the target language model with the punctuation label sequence in the text to be recognized. If there is a punctuation label in the predicted punctuation label sequence that is inconsistent with the punctuation label at the corresponding position in the punctuation label sequence of the text to be recognized, it is considered that there is a punctuation error at the corresponding position. Since the punctuation label sequence includes the punctuation symbols at each position where punctuation needs to be added, by comparing the punctuation label sequences, the verification of all punctuation positions in the text can be realized, improving the accuracy of punctuation verification while ensuring the efficiency of punctuation verification.

[0099] Exemplarily, as Figure 7 shown, Figure 7 is Figure 4 the specific implementation flowchart of S403 in. It can be seen from Figure 7 that S403 includes S4031 and S4032. Details are as follows:

[0100] S4031. Compare each first punctuation in the punctuation label sequence with the corresponding second punctuation in the text to be recognized in turn.

[0101] S4032. If the second punctuation at the target position in the text to be recognized is different from the first punctuation at the target position in the punctuation label sequence, it is determined that there is a punctuation error at the target position.

[0102] From the above analysis, it can be seen that the text punctuation detection method provided by the embodiments of the present application first analyzes the context information and part-of-speech of the characters in the text to be recognized through the target language model to obtain the punctuation label sequence of the text to be recognized; then performs punctuation detection on the text to be recognized based on the punctuation label sequence. Since the target language model is a network layer that fuses the context information and part-of-speech for analyzing characters in the text after training the preset language model based on the target training samples, and its corresponding target training samples are the text data obtained after punctuation correction of the text data based on the back-translation data augmentation strategy, therefore, the context information and part-of-speech of the characters in the text to be recognized can be analyzed through the target language model, and then the punctuation label sequence of the text to be recognized can be predicted according to the context information and part-of-speech of the characters in the text to be recognized, improving the accuracy of text punctuation detection.

[0103] Please refer to Figure 8 as shown Figure 8 which is a schematic diagram of the implementation process of the text punctuation detection method provided by another embodiment of this application. The text punctuation detection method provided by this embodiment can be completed by Figure 3 the cloud platform and computer device shown as follows:

[0104] S801, the cloud platform obtains a preset number of target training samples, and the target training samples are text data obtained after punctuation correction of the text data based on the back-translation data augmentation strategy.

[0105] S802, the cloud platform trains a preset language model based on the target training samples to obtain a target language model, and the preset language model integrates a network layer for analyzing the context information and part-of-speech of characters in the text.

[0106] S803, the computer device analyzes the context information and part-of-speech of the characters in the text to be recognized based on the target language model to obtain a punctuation label sequence of the text to be recognized.

[0107] S804, the computer device performs punctuation detection on the text to be recognized based on the punctuation label sequence.

[0108] It should be noted that those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific implementation processes of the above steps can refer to Figure 4 the specific implementation processes of the corresponding steps in the embodiments shown, and will not be elaborated here.

[0109] It should be understood that using natural language processing technology to automatically identify text, especially the punctuation errors in primary school compositions, has become a very valuable technology. On the one hand, this technology can be applied to the automatic correction of primary school compositions, assisting students in discovering punctuation errors in their compositions, enhancing students' awareness and ability to use punctuation marks, and thus helping students write clearer and more fluent compositions. On the other hand, this technology can be applied to the automatic scoring of primary school compositions, providing valuable features for the automatic scoring of compositions, making the results of automatic composition scoring more reasonable.

[0110] As can be seen from the above analysis, the text punctuation detection method provided by the embodiments of the present application first obtains a preset number of target training samples, where the target training samples are text data obtained by correcting the punctuation of text data based on a back-translation data augmentation strategy; then trains a preset language model based on the target training samples to obtain a target language model, and the preset language model incorporates a network layer for analyzing the context information and part-of-speech of characters in the text; then analyzes the context information and part-of-speech of characters in the text to be recognized based on the target language model to obtain a punctuation label sequence of the text to be recognized; and finally performs punctuation detection on the text to be recognized based on the punctuation label sequence. By training a preset language model that incorporates a network layer for analyzing the context information and part-of-speech of characters in the text based on target training samples, and performing text punctuation detection using the trained target language model, the accuracy of text punctuation detection is improved.

[0111] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 10 includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory may include a non-volatile storage medium and an internal memory.

[0112] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can be caused to execute the text punctuation detection method.

[0113] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0114] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium, and when the computer program is executed by the processor, the processor can be caused to execute the text punctuation detection method.

[0115] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 9 the structure shown in

[0116] It should be understood that the processor may be a Central Processing Unit (CPU), and the processor may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0117] Among them, in one embodiment, the processor is used to run a computer program stored in a memory to implement the following steps:

[0118] Obtain the text to be recognized, and input the text to be recognized into a pre-trained target language model. Among them, the target language model is a network layer that fuses the context information and part-of-speech of characters in the text after training a preset language model based on target training samples. The target training samples are text data obtained by punctuation correction of text data based on a back-translation data augmentation strategy;

[0119] Analyze the context information and part-of-speech of characters in the text to be recognized based on the target language model to obtain a punctuation label sequence of the text to be recognized;

[0120] Perform punctuation detection on the text to be recognized based on the punctuation label sequence.

[0121] In one embodiment, the punctuation correction of text data based on the back-translation data augmentation strategy includes:

[0122] Perform paragraph splitting processing on the text data to obtain at least one first paragraph;

[0123] Input each of the first paragraphs into a first translation engine respectively to obtain the English corresponding to each of the first paragraphs;

[0124] Input the English corresponding to each of the first paragraphs into a second translation engine respectively to obtain corrected text data for each of the first paragraphs;

[0125] Among them, the corrected text data is the text data obtained after punctuation correction of each of the first paragraphs.

[0126] In one embodiment, the preset language model includes a first network layer, a second network layer, and a third network layer;

[0127] The first network layer is used to represent each character in the target training sample as a vector, obtaining a first vector corresponding to each character;

[0128] The second network layer is used to analyze the first vector, obtaining a second vector representing the context information of each character;

[0129] The third network layer is used to predict the punctuation label following the character corresponding to each context information according to the part-of-speech labels of the respective context information in the second vector, obtaining a punctuation label sequence.

[0130] In one embodiment, the preset language model further includes an output layer and a loss function;

[0131] Training the preset language model based on the target training sample to obtain a target language model includes:

[0132] Inputting the target training sample into the first network layer for analysis to obtain the respective first vectors;

[0133] Inputting the respective first vectors into the second network layer for analysis to obtain the respective second vectors;

[0134] Inputting the respective second vectors into the third network layer for analysis, outputting a punctuation label sequence through the output layer, and detecting the value of the loss function;

[0135] If the value of the loss function is less than a preset threshold, stop training the preset language model to obtain the target language model.

[0136] In one embodiment, the first network layer includes a character embedding layer, a character position embedding layer, and a clause position embedding layer;

[0137] The inputting the target training sample into the first network layer for analysis to obtain the respective first vectors includes:

[0138] Performing paragraph splitting processing on the target training sample to obtain a plurality of second paragraphs;

[0139] For any one of the second paragraphs, removing the preset type of punctuation in the second paragraph to obtain the string of the second paragraph;

[0140] Inputting the string into the character embedding layer for analysis to obtain the character information of each character in the string;

[0141] Inputting the string into the character position embedding layer for analysis to obtain the first position information of each character in the second paragraph in the string;

[0142] Inputting the character string into the clause position embedding layer for analysis to obtain second position information of the clause to which each character in the character string belongs in the second paragraph;

[0143] The first vector of each character in the character string is generated based on the character information, the first position information, and the second position information.

[0144] In one embodiment, the second network layer is an encoding layer based on a multi-head attention mechanism;

[0145] The step of inputting each of the first vectors into the second network layer for analysis to obtain each of the second vectors includes:

[0146] For each of the first vectors, determine each attention weight of the first vector by a multi-head attention mechanism, where each attention weight is an attention weight between a character corresponding to the first vector and other characters in the character string to which it belongs;

[0147] Performing a weighted summation on the attention weights to obtain context information of the character corresponding to the first vector in the character string to which it belongs;

[0148] A corresponding second vector is generated based on the context information.

[0149] In one embodiment, the third network layer includes a part-of-speech prediction layer and a punctuation prediction layer;

[0150] The step of inputting each of the second vectors into the third network layer for analysis includes:

[0151] Predicting the part of speech of each context information in the second vector based on the part of speech prediction layer to generate a part of speech tag;

[0152] The part-of-speech tags are analyzed based on the punctuation prediction layer, and the punctuation tags following the characters corresponding to each context information are predicted to obtain a punctuation tag sequence.

[0153] In one embodiment, the step of performing punctuation detection on the text to be recognized based on the punctuation tag sequence includes:

[0154] Comparing each first punctuation mark in the punctuation mark tag sequence with the corresponding second punctuation mark in the text to be recognized;

[0155] If a second punctuation mark at a target position in the text to be recognized is different from a first punctuation mark at the target position in the punctuation mark tag sequence, it is determined that a punctuation mark error exists at the target position.

[0156] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions, and the processor executes the program instructions to implement the present application Figure 4 The text punctuation detection method provided by the embodiment shown

[0157] Among them, the computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device

[0158] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or replacements, and these modifications or replacements should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims

Claims

1. A method for detecting text punctuation, characterized in that, The method includes: Obtain the text to be recognized, and input the text to be recognized into a pre-trained target language model. Among them, the target language model is a network layer that fuses the context information and part-of-speech of characters in the text after training a preset language model based on target training samples. The target training samples are text data obtained by correcting punctuation for text data based on a back-translation data augmentation strategy; Analyze the context information and part-of-speech of characters in the text to be recognized based on the target language model to obtain a punctuation label sequence of the text to be recognized; Perform punctuation detection on the text to be recognized based on the punctuation label sequence; Among them, the preset language model includes a first network layer, and the first network layer includes a character embedding layer, a character position embedding layer, and a clause position embedding layer; The first network layer is used to represent each character in the target training sample as a vector to obtain a first vector corresponding to each character, including: obtaining a character string, inputting the character string into the character embedding layer for analysis to obtain the character information of each character in the character string; inputting the character string into the character position embedding layer for analysis to obtain the first position information of each character in the second paragraph in the character string; inputting the character string into the clause position embedding layer for analysis to obtain the second position information of the clause to which each character in the character string belongs in the second paragraph; generating a first vector for each character in the character string based on the character information, the first position information, and the second position information.

2. The method according to claim 1, characterized in that, The punctuation correction of the text data based on the back-translation data augmentation strategy includes: Perform paragraph splitting on the text data to obtain at least one first paragraph; Input each of the first paragraphs into a first translation engine respectively to obtain the English corresponding to each of the first paragraphs; Input the English corresponding to each of the first paragraphs into a second translation engine respectively to obtain corrected text data for each of the first paragraphs; Among them, the corrected text data is text data obtained after punctuation correction for each of the first paragraphs.

3. The method according to claim 1, characterized in that, The preset language model further includes a second network layer and a third network layer; The second network layer is used to analyze the first vector to obtain a second vector representing the context information of each character; The third network layer is used to predict the punctuation label after the character corresponding to each context information based on the part-of-speech label of each context information in the second vector to obtain a punctuation label sequence.

4. The method according to claim 3, characterized in that, The preset language model further includes an output layer and a loss function; The training of the preset language model based on the target training samples includes: Input the target training samples into the first network layer for analysis to obtain each of the first vectors; Input each of the first vectors into the second network layer for analysis to obtain each of the second vectors; Input each of the second vectors into the third network layer for analysis, output a punctuation label sequence through the output layer, and detect the value of the loss function; If the value of the loss function is less than a preset threshold, stop training the preset language model to obtain the target language model.

5. The method according to claim 4, characterized in that, The step of inputting the target training sample into the first network layer for analysis to obtain each of the first vectors further includes: Performing paragraph splitting processing on the target training sample to obtain a plurality of second paragraphs; For any one of the second paragraphs, the preset type of punctuation in the second paragraph is removed to obtain the character string of the second paragraph.

6. The method according to claim 4, characterized in that, The second network layer is an encoding layer based on a multi-head attention mechanism; The step of inputting each of the first vectors into the second network layer for analysis to obtain each of the second vectors includes: For each of the first vectors, determine each attention weight of the first vector by a multi-head attention mechanism, where each attention weight is an attention weight between a character corresponding to the first vector and other characters in the character string to which it belongs; Performing a weighted summation on the attention weights to obtain context information of the character corresponding to the first vector in the character string to which it belongs; A corresponding second vector is generated based on the context information.

7. The method according to claim 4, characterized in that, The third network layer includes a part-of-speech prediction layer and a punctuation prediction layer; The step of inputting each of the second vectors into the third network layer for analysis includes: Predicting the part of speech of each context information in the second vector based on the part of speech prediction layer to generate a part of speech tag; The part-of-speech tags are analyzed based on the punctuation prediction layer, and the punctuation tags following the characters corresponding to each context information are predicted to obtain a punctuation tag sequence.

8. The method according to claim 1, characterized in that, The performing punctuation detection on the to-be-recognized text based on the punctuation tag sequence includes: Comparing each first punctuation mark in the punctuation mark tag sequence with the corresponding second punctuation mark in the text to be recognized; If a second punctuation mark at a target position in the text to be recognized is different from a first punctuation mark at the target position in the punctuation mark tag sequence, it is determined that a punctuation mark error exists at the target position.

9. A computer device, characterized in that, include: Memory and processor; The memory is used to store computer programs; The processor is used to execute the computer program and implement the steps of the text punctuation detection method according to any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the steps of the text punctuation detection method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method for adding punctuation marks to punctuation-free text

    CN108932226A

  • Text punctuation correction method and device, electronic equipment and storage medium

    CN112733530A

  • Text feature extraction method and device, computer equipment and storage medium

    CN113449081A