Ancient Chinese punctuation prediction method, system, electronic device and medium

Through the method of splitting the ancient text training data and building an index library with the minimum hash algorithm, the problem of low accuracy in ancient text punctuation prediction is solved, and higher punctuation prediction accuracy and ability to adapt to complex scenarios is achieved.

CN119150864BActive Publication Date: 2025-05-13SHANGHAI MIDU INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411640553.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-05-13
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

The prior art has little accuracy in the prediction of punctuation in ancient Chinese, especially when dealing with complex rhetorical techniques, nested long sentences or multiple meanings, it is difficult to correctly understand the context, resulting in inappropriate punctuation marks.

Method used

By obtaining training data, splitting processing is performed to obtain the data-enhanced training data set, using the minimum hash algorithm to build an index library, obtain the reference text of the training data set, and train the initial language model in combination with the original text to obtain the ancient text punctuation prediction model.

Benefits of technology

This method can avoid the missed and false alarm problems of continuous punctuation prediction, adapt to complex ancient text scenes, quickly complete the breaking sentences and punctuation of ancient texts, and improve the accuracy of punctuation prediction of ancient texts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119150864B_ABST
    Figure CN119150864B_ABST
Patent Text Reader

Abstract

The present application provides a method, system, electronic device and medium for predicting ancient Chinese punctuation, the method comprising: obtaining training data; splitting the training data and obtaining a training data set using the split data blocks; constructing an index library using a minimum hash algorithm to obtain a reference text of the training data set; training an initial language model using the reference text and the original text of the training data set to obtain an ancient Chinese punctuation prediction model; and predicting a text to be predicted using the ancient Chinese punctuation prediction model to obtain a prediction result. This method for predicting ancient Chinese punctuation can avoid the problem of underreporting in continuous punctuation prediction and improve the accuracy of ancient Chinese punctuation prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of natural language processing, and relates to a method for predicting punctuation marks in ancient Chinese, and in particular to a method, system, electronic device and medium for predicting punctuation marks in ancient Chinese. Background Art

[0002] With the development of the Internet, the amount of unstructured data is also gradually increasing. People have begun to use artificial intelligence to process data and extract effective information from it. As an important part of Chinese history, the language structure, grammatical rules and punctuation usage of ancient Chinese are significantly different from those of modern Chinese. Punctuation marks are often omitted in ancient Chinese, and the use of punctuation is not fixed. It often depends on the context and the author's personal habits, which makes it easy for machine learning models to be ambiguous and inaccurate when processing ancient Chinese punctuation. Most of the current punctuation prediction technologies are trained based on the punctuation rules of modern Chinese, and the grammatical structure of modern Chinese is quite different from that of ancient Chinese. For example, sentences in ancient Chinese often omit subject-predicate structures or use inverted sentences, which are relatively rare in modern Chinese. Although traditional rule-based models and deep learning methods have achieved certain success in modern Chinese punctuation prediction, in the processing of ancient Chinese, due to the lack of sufficient corpus and rule support, the accuracy of the prediction results is not high. In particular, when complex rhetoric, long sentence nesting or multiple meanings appear in ancient Chinese, the existing technology models are difficult to correctly understand the context, and thus give inappropriate punctuation marks.

[0003] Therefore, how to solve the inaccurate problem in punctuation prediction of ancient Chinese texts has become an important direction of current research. Summary of the invention

[0004] In view of the above-mentioned shortcomings of the prior art, the purpose of the present application is to provide a method, system, electronic device and medium for predicting ancient Chinese punctuation, so as to solve the problem of low accuracy of ancient Chinese punctuation prediction in the prior art.

[0005] In a first aspect, the present application provides a method for predicting punctuation in classical Chinese, which comprises: obtaining training data; splitting the training data and obtaining a training data set using the split data blocks; constructing an index library using a minimum hash algorithm to obtain a reference text of the training data set; training an initial language model using the reference text and the original text of the training data set to obtain a classical Chinese punctuation prediction model; and predicting a text to be predicted using the classical Chinese punctuation prediction model to obtain a prediction result.

[0006] In this application, the training data is split and processed to obtain a data-enhanced training data set, the reference text of the training data set is obtained using the minimum hash algorithm, the initial language model is trained using the reference text and the original text, and an ancient Chinese punctuation prediction model is obtained, and the prediction result of the text to be predicted is obtained using the ancient Chinese punctuation prediction model. This ancient Chinese punctuation prediction method can avoid the problems of missed reports and false reports in continuous punctuation prediction, adapt to complex ancient Chinese scenes, quickly complete the sentence segmentation and punctuation of ancient texts, and improve the accuracy of ancient Chinese punctuation prediction.

[0007] In an implementation of the first aspect, the training data is split and processed, and obtaining a training data set using the split data blocks includes: performing sentence splitting and text splitting on the training data to obtain a short sentence set and a text set; obtaining a maximum number of short sentences and a maximum number of texts within a window range based on the short sentence set and the text set; and obtaining the training data set based on the maximum number of short sentences, the maximum number of texts, the short sentence set, and the text set.

[0008] In an implementation of the first aspect, the training data set includes at least one text to be retrieved, and using a minimum hash algorithm to construct an index library to obtain a reference text of the text to be retrieved includes: obtaining a hash signature vector of the training data set based on the minimum hash algorithm; using the hash signature vector to construct a signature index library based on a local sensitive hash forest; obtaining a hash signature vector of the text to be retrieved based on the minimum hash algorithm; and using the hash signature vector of the text to be retrieved to search in the signature index library to obtain a reference text of the text to be retrieved.

[0009] In an implementation of the first aspect, searching in the signature index library using the hash signature vector of the text to be retrieved to obtain a reference text of the text to be retrieved includes: searching in the signature index library using the hash signature vector of the text to be retrieved to obtain at least one search result; screening the search results using Jaccard similarity to obtain a search result whose Jaccard similarity is greater than a set threshold as a reference text of the text to be retrieved.

[0010] In an implementation of the first aspect, the classical Chinese punctuation prediction method includes: when predicting the text to be predicted using the classical Chinese punctuation prediction model, limiting the next predicted character to be consistent with the character of the text to be predicted.

[0011] In an implementation of the first aspect, the ancient Chinese punctuation prediction method includes: when using the ancient Chinese punctuation prediction model to predict the text to be predicted, obtaining at least one candidate prediction result, and taking the candidate prediction result with the largest probability value as the prediction result.

[0012] In an implementation of the first aspect, the initial language model is trained using the reference text and the original text of the training data set to obtain an ancient Chinese punctuation prediction model, including: using the original text of the training data set and the reference text to input the initial language model for training to obtain an initial ancient Chinese punctuation prediction model; and using back propagation to adjust the parameters of the initial ancient Chinese punctuation prediction model until the model converges to obtain the ancient Chinese punctuation prediction model.

[0013] In a second aspect, the present application provides an ancient Chinese punctuation prediction system, which includes: a data acquisition module for acquiring training data; a data processing module for splitting the training data and obtaining a training data set using the split data blocks; a reference text module for constructing an index library using a minimum hash algorithm to obtain a reference text of the training data set; a model acquisition module for training an initial language model using the reference text and the original text of the training data set to obtain an ancient Chinese punctuation prediction model; and a result prediction module for predicting a text to be predicted using the ancient Chinese punctuation prediction model to obtain a prediction result.

[0014] In a third aspect, the present application provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program stored in the memory, so that the electronic device performs the ancient Chinese punctuation prediction method as described in any one of the first aspects.

[0015] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the ancient Chinese punctuation prediction method described in any one of the first aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1A Shown is a schematic diagram of an application scenario of the ancient Chinese punctuation prediction device described in this application.

[0017] Figure 1B Shown are structural diagrams of the client-cloud interaction scenarios in these implementations.

[0018] Figure 2 Shown is a flow chart of the ancient Chinese punctuation prediction method described in an embodiment of the present application.

[0019] Figure 3A Shown is a flow chart of the ancient Chinese punctuation prediction method described in an embodiment of the present application.

[0020] Figure 3B Shown is a schematic diagram of the splitting process described in an embodiment of the present application.

[0021] Figure 4Shown is a flow chart of the ancient Chinese punctuation prediction method described in an embodiment of the present application.

[0022] Figure 5 Shown is a structural schematic diagram of the ancient Chinese punctuation prediction system described in an embodiment of the present application.

[0023] Figure 6 Shown is a schematic diagram of the structure of an electronic device described in an embodiment of the present application. DETAILED DESCRIPTION

[0024] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0025] It should be noted that in the embodiments of the present application, the words "optionally" or "for example" represent examples, illustrations or descriptions. Any embodiment or design described as "optionally" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "optionally" or "for example" is intended to present related concepts in a specific way.

[0026] In the embodiments of the present application, "at least one" refers to one or more, and "plurality" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can represent: a, b, c, ab, ac, bc or abc, where a, b, c can be single or multiple.

[0027] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application, and thus the drawings only show components related to the present application rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed at will, and the component layout may also be more complicated.

[0028] The rule-based ancient Chinese punctuation prediction method and the deep learning-based ancient text punctuation repair method face two problems: First, the rule-based ancient Chinese punctuation prediction method divides the ancient book sentence segmentation and punctuation into two serial tasks, which will cause error transmission, and it is time-consuming to manually design a large number of rule judgments. Secondly, ancient Chinese deep learning models such as SikuBERT have poor accuracy and cannot handle complex sentences. The use of LLM (Large Language Model) large language model and retrieval enhancement can effectively solve this series of problems. In specific scenarios, because there are multiple consecutive punctuation marks in the sentence, if structural models such as BERT are used, punctuation marks will be missed. In more complex ancient Chinese scenarios, due to the limitation of model parameters, ordinary deep learning models will cause false positives and missed negatives in the prediction process of punctuation marks such as book titles and quotation marks.

[0029] At least to address the above-mentioned problems, an embodiment of the present application provides a method for predicting punctuation in classical Chinese, which comprises: obtaining training data; splitting the training data and obtaining a training data set using the split data blocks; constructing an index library using a minimum hash algorithm to obtain a reference text of the training data set; training an initial language model using the reference text and the original text of the training data set to obtain a classical Chinese punctuation prediction model; and predicting a text to be predicted using the classical Chinese punctuation prediction model to obtain a prediction result.

[0030] In the embodiment of the present application, the training data is split and processed to obtain a data-enhanced training data set, the reference text of the training data set is obtained using the minimum hash algorithm, the initial language model is trained using the reference text and the original text, and the ancient Chinese punctuation prediction model is obtained, and the prediction result of the text to be predicted is obtained using the ancient Chinese punctuation prediction model. This ancient Chinese punctuation prediction method can avoid the problems of missed reports and false reports in continuous punctuation prediction, adapt to complex ancient Chinese scenes, quickly complete the sentence segmentation and punctuation of ancient texts, and improve the accuracy of ancient Chinese punctuation prediction.

[0031] Figure 1A The diagram shows an application scenario of the ancient Chinese punctuation prediction device described in the present application. The ancient Chinese punctuation prediction device 1 can be used to implement the ancient Chinese punctuation prediction method provided in the present application embodiment, but the application scenario of the ancient Chinese punctuation prediction method provided in the present application embodiment is not limited to Figure 1A The ancient Chinese punctuation prediction device 1 shown in FIG. Figure 1A As shown, the ancient Chinese punctuation prediction device 1 includes a text storage device 11, a local processor 12 and a display terminal 13. The ancient Chinese punctuation prediction method provided in the embodiment of the present application can be applied to the local processor 12.

[0032] in, Figure 1AThe local processor 12 in the embodiment may be a local processor or a local processor cluster composed of multiple local processors or a cloud computing center, etc., which is not limited to the specifics herein. Figure 1A Only one text storage device 11, one local processor 12 and one display terminal 13 are shown, but it should be understood that Figure 1A The examples are only used to understand the present solution, and the specific numbers of local processors 12 and display terminals 13 should be flexibly determined based on actual conditions.

[0033] In some other implementations, the ancient Chinese punctuation prediction device 1 may not include the display terminal 13, but only include a local processor 12 with a display function and an ancient Chinese storage device 11. The ancient Chinese punctuation prediction method provided in the embodiment of the present application can be applied to the local processor 12. The local processor 12 with a display function can include a tablet computer, a laptop computer, a PDA, a mobile phone, a personal computer, and a voice interaction device, which are not limited here.

[0034] In still other implementations, the ancient Chinese punctuation prediction method described in this application can be applied to end-cloud interaction scenarios. Figure 1B The following is a schematic diagram of the structure of the client-cloud interaction scenario in these implementations. Figure 1B As shown, the terminal-cloud interaction system 2 includes a terminal 20 and a cloud server 21. The terminal 20 and the cloud server 21 can communicate with each other, and the communication method is not limited to wired or wireless.

[0035] Among them, the terminal 20 can be mobile or fixed, for example, the terminal 20 can be a wireless terminal or a wired terminal. The wireless terminal can refer to a device with wireless transceiver function, which can be deployed indoors, outdoors and in industrial workshops. The terminal 20 can be a mobile phone, a tablet computer, a laptop computer, etc., which is not limited here. The cloud server 21 can include one or more servers, or one or more processing nodes, or one or more virtual machines running on the server. The cloud server 21 can also be called a server cluster, a management platform, a data processing center, etc., which is not limited in the embodiments of the present application.

[0036] The technical solutions in the embodiments of the present application will be described in detail below in conjunction with the drawings in the embodiments of the present application.

[0037] The following embodiments of the present application provide a method for predicting punctuation marks in ancient Chinese. For example, the method can be Figure 1A The local processor 12 shown or Figure 1B It is implemented by the cloud server 21 shown. Figure 2 The flowchart of the method for predicting ancient Chinese punctuation marks described in the embodiment of the present application is shown as follows: Figure 2 As shown, the ancient Chinese punctuation prediction method includes steps S11 to S15.

[0038] Step S11, obtaining training data.

[0039] Step S12, splitting the training data, and obtaining a training data set using the split data blocks.

[0040] Step S13, constructing an index library using a minimum hash algorithm to obtain reference text of the training data set.

[0041] Step S14, training an initial language model using the reference text and the original text of the training data set to obtain an ancient Chinese punctuation prediction model.

[0042] Step S15, predicting the text to be predicted using the ancient Chinese punctuation prediction model to obtain a prediction result.

[0043] In some possible implementations, training data is obtained, data enhancement is performed using a sliding window, data splitting preprocessing is performed on the training data, and a training data set is obtained using the split data. An index library is constructed for all sentences in the training data set using a minimum hash algorithm to obtain a reference text for the training data set. The initial language model is trained using the reference text and the original text of the training data set. During training, different but most similar sentences are retrieved from the index library based on sentence similarity to fine-tune the initial language model and obtain an ancient Chinese punctuation prediction model. The ancient Chinese punctuation prediction model is used to predict the text to be predicted, and decoding constraints are added during the prediction process to limit the predicted vector to a set of characters or punctuation marks in the original sentence, and other vectors are filtered to obtain a prediction result.

[0044] In the embodiment of the present application, the training data is split and processed to obtain a data-enhanced training data set, the reference text of the training data set is obtained using the minimum hash algorithm, the initial language model is trained using the reference text and the original text, and the ancient Chinese punctuation prediction model is obtained, and the prediction result of the text to be predicted is obtained using the ancient Chinese punctuation prediction model. This ancient Chinese punctuation prediction method can avoid the problems of missed reports and false reports in continuous punctuation prediction, adapt to complex ancient Chinese scenes, quickly complete the sentence segmentation and punctuation of ancient texts, and improve the accuracy of ancient Chinese punctuation prediction.

[0045] Figure 3A The flowchart of the method for predicting ancient Chinese punctuation marks described in the embodiment of the present application is shown as follows: Figure 3A As shown, the step S12 includes steps S121 to S123.

[0046] Step S121 , performing sentence splitting and text splitting on the training data to obtain a short sentence set and a text set.

[0047] Step S122, obtaining the maximum number of short sentences and the maximum number of texts within the window range according to the short sentence set and the text set.

[0048] Step S123, acquiring the training data set according to the maximum number of short sentences, the maximum number of texts, the short sentence set and the text set.

[0049] In some possible implementations, Figure 3B The diagram is a schematic diagram of the splitting process described in the embodiment of the present application. Figure 3B As shown, a sliding window design is set to enhance the data, and the training data is split into m short sentences, recorded as . Perform text splitting on the training data and split it into n text blocks, recorded as .

[0050] The short sentences after splitting are Start to End, where The maximum number of sentences within the window is The text of each window after splitting is recorded as ,from Start to do this operation continuously, recorded as . is the maximum size of the text window, is the maximum value of the overlapping windows. The training data set is obtained according to the maximum number of short sentences, the maximum number of texts, the short sentence set and the text set.

[0051] In the embodiment of the present application, sentence splitting and text splitting of the training data not only meets the length limit of the training data set, but also increases the number of trainable data samples, provides more diverse learning samples for model training, and helps to improve the generalization ability of the model.

[0052] Figure 4 The flowchart of the method for predicting ancient Chinese punctuation marks described in the embodiment of the present application is shown as follows: Figure 4 As shown, the training data set includes at least one text to be retrieved, and the step S13 includes steps S131 to S134.

[0053] Step S131, obtaining a hash signature vector of the training data set based on a minimum hash algorithm.

[0054] Step S132: construct a signature index library based on a local sensitive hash forest using the hash signature vector.

[0055] Step S133, obtaining a hash signature vector of the text to be retrieved based on the minimum hash algorithm.

[0056] Step S134: Search the signature index library using the hash signature vector of the text to be retrieved to obtain a reference text of the text to be retrieved.

[0057] In some possible implementations, minimum hash is a method that can quickly determine the similarity of sets. The principle is that the probability that the minimum hash values ​​obtained after applying a certain random permutation to two sets are equal is equal to the Jaccard similarity of the two sets. Based on the minimum hash algorithm, 64 groups of permutation functions are used to hash the elements of the set, and the minimum hash value obtained by each operation is retained to obtain the hash signature vector of the training data set. All the hash signature vectors are based on the local sensitive hash forest to build a signature index library. Specifically, after the input text is obtained by the minimum hash operation to obtain the signature vector, it still takes linear time to complete the screening of the associated samples in the signature library. The datasketch development package is used to further pre-build all the signature vectors in the signature library into a local sensitive hash forest (LSH Forest), and the data is dispersed into multiple hash buckets to significantly reduce the amount of data to be compared and improve the search efficiency. The initialization parameters and the number of prefix trees are 64 and 8 respectively. In addition, since the sample contains punctuation and the predicted text does not, in order to maintain consistency, the punctuation marks in the sample text are removed before generating the signature vector, and the signature index library is generated with punctuation-free text.

[0058] Based on the minimum hash algorithm, the hash signature vector of the text to be retrieved is obtained, and the hash signature vector of the text to be retrieved is used to search in the signature index library to obtain the reference text of the text to be retrieved. Given the input text to be retrieved x, the reference text e and the known target text t with punctuation are retrieved and used as the input, reference text and output result of the initial language model respectively, and the initial language model is fine-tuned.

[0059] In one embodiment of the present application, using the hash signature vector of the text to be searched to search in the signature index library to obtain the reference text of the text to be searched includes: using the hash signature vector of the text to be searched to search in the signature index library to obtain at least one search result. Using Jaccard similarity to filter the search results, obtaining the search results with Jaccard similarity greater than a set threshold as the reference text of the text to be searched.

[0060] Specifically, the hash signature vector of the text to be detected is used to search in the signature index library to obtain at least one search result. The search results are screened for character similarity using Jaccard similarity, and sentences with a similarity greater than 0.2 to the original sentence are retained as the reference text to be searched. It should be noted that the threshold is set to be adjusted according to the specific application scenario, and this application is not limited to this.

[0061] In one embodiment of the present application, the classical Chinese punctuation prediction method includes: when predicting the text to be predicted using the classical Chinese punctuation prediction model, limiting the next predicted character to be consistent with the character of the text to be predicted.

[0062] In some possible implementations, the generative large language model learns the structure, rules and semantics of the language by predicting the next character. When facing specific tasks, there will be uncontrollable output results. Taking the automatic punctuation task of ancient texts as an example, even through fine-tuning of instructions, the model will still be uncontrollable, which is manifested in that the characters written out are neither punctuation nor Chinese characters in the original text, or just copy the entire sentence without any modification. When using the ancient punctuation prediction model to predict the text to be predicted, de-coding constraints are imposed. The ancient punctuation task cannot delete the original characters or add new Chinese characters. Based on this feature, the next predicted character is restricted to be consistent with the characters of the text to be predicted to solve the problem of inconsistent output characters.

[0063] In one embodiment of the present application, the ancient Chinese punctuation prediction method includes: when using the ancient Chinese punctuation prediction model to predict the text to be predicted, obtaining at least one candidate prediction result, and taking the candidate prediction result with the largest probability value as the prediction result.

[0064] In some possible implementations, when the ancient Chinese punctuation prediction model is used to predict the text to be predicted, at least one candidate prediction result is obtained, each of the candidate prediction results is a character with a different probability, and the candidate prediction result with the largest probability value is obtained as the prediction result to solve the problem of inconsistent output characters.

[0065] In some other possible implementations, the classical Chinese punctuation prediction model is a prediction model that performs retrieval enhancement after the training data splitting process described in an embodiment of the present application, and performs character restrictions during prediction. The classical Chinese punctuation prediction model also includes a first prediction model and a second prediction model. The first prediction model is the classical Chinese punctuation prediction model for retrieval enhancement after the training data splitting process described in an embodiment of the present application, and the second prediction model is an classical Chinese punctuation prediction model that has not been split and processed by the training data, and does not perform character restrictions during prediction. The first prediction model, the second prediction model and the classical Chinese punctuation prediction model are combined into an expert combination model, and a voting mechanism is used to obtain a decoding result. Only when the output results of two or more models are the same can they be used as the prediction result.

[0066] In one embodiment of the present application, the initial language model is trained using the reference text and the original text of the training data set to obtain an ancient Chinese punctuation prediction model, including: using the original text of the training data set and the reference text to input the initial language model for training to obtain an initial ancient Chinese punctuation prediction model; and using back propagation to adjust the parameters of the initial ancient Chinese punctuation prediction model until the model converges to obtain the ancient Chinese punctuation prediction model.

[0067] In some possible implementations, the original text of the training data set and the reference text are input into the initial language model for training to obtain an initial ancient Chinese punctuation prediction model; the parameters of the initial ancient Chinese punctuation prediction model are adjusted by back propagation, and the cycle is repeated multiple times until the model converges to obtain the ancient Chinese punctuation prediction model. The original text of the training data set is: Buy machines to weave velvet, felt, yarn, feathers, foreign shirts, pants, foreign socks, foreign umbrellas, etc. Refine lake sand to make glassware, refine refined copper to imitate clocks and watches, which are lifelike and strong and cheap. This is the third battle with various items. Shanghai papermaking, Kanto rolling tobacco, Nanyang, planting sugarcane in Zhongzhou, opening grape gardens, brewing wine and making sugar. This is the fourth battle with various foods. The reference text is: Buy new machines widely, weave various cloths by ourselves, one province is well done, and then promote it to other provinces. This is the second battle with foreign cloth. Buy machines to weave velvet, felt, woolen yarn, feathers, foreign shirts, pants, socks, and umbrellas, etc., refine lake sand to make glassware, refine copper to make replica clocks, which are lifelike, strong and cheap. This is the third one to fight with all kinds of things. The output result of the ancient Chinese punctuation prediction model is: Buy machines to weave velvet, felt, woolen yarn, feathers, foreign shirts, pants, socks, and umbrellas. Refine lake sand to make glassware. Refine copper to make replica clocks, which are lifelike, strong and cheap. This is the third one to fight with all kinds of things. Shanghai makes paper, Kanto rolls tobacco, Nanyang expands sugarcane planting, Zhongzhou opens grape gardens, brews wine and makes sugar. This is the fourth one to fight with all kinds of food.

[0068] Figure 5 The structure diagram of the ancient Chinese punctuation prediction system described in the embodiment of the present application is shown as follows: Figure 5 As shown, the ancient Chinese punctuation prediction system 100 includes a data acquisition module 110 , a data processing module 120 , a reference text module 130 , a model acquisition module 140 and a result prediction module 150 .

[0069] The data acquisition module 110 is used to acquire training data.

[0070] The data processing module 120 is used to split the training data and obtain a training data set using the split data blocks.

[0071] The reference text module 130 is used to construct an index library using a minimum hash algorithm to obtain reference text of the training data set.

[0072] The model acquisition module 140 is used to train the initial language model using the reference text and the original text of the training data set to obtain an ancient Chinese punctuation prediction model.

[0073] The result prediction module 150 is used to predict the text to be predicted by using the ancient Chinese punctuation prediction model to obtain a prediction result.

[0074] In the embodiment of the present application, the data processing module 120 is used to split the training data to obtain a data-enhanced training data set, the reference text module 130 is used to obtain the reference text of the training data set using the minimum hash algorithm, the model acquisition module 140 is used to train the initial language model using the reference text and the original text to obtain the ancient Chinese punctuation prediction model, and the result prediction module 150 is used to obtain the prediction result of the text to be predicted using the ancient Chinese punctuation prediction model. This ancient Chinese punctuation prediction system 100 can avoid the problems of missed reports and false reports in continuous punctuation prediction, adapt to complex ancient Chinese scenes, quickly complete the sentence segmentation and punctuation of ancient texts, and improve the accuracy of ancient Chinese punctuation prediction.

[0075] In the several embodiments provided in the present application, it should be understood that the disclosed system, device or method can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of modules / units is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules or units, which can be electrical, mechanical or other forms.

[0076] The modules / units described as separate components may or may not be physically separated, and the components displayed as modules / units may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules / units may be selected according to actual needs to achieve the purpose of the embodiments of the present application. For example, the functional modules / units in the various embodiments of the present application may be integrated into one processing module, or each module / unit may exist physically separately, or two or more modules / units may be integrated into one module / unit.

[0077] Those of ordinary skill in the art should further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0078] An embodiment of the present application also provides an electronic device. Figure 6 The structure diagram of the electronic device 600 according to the embodiment of the present application is shown. Figure 6 As shown, in this embodiment, the electronic device 600 includes a memory 610 and a processor 620 .

[0079] The memory 610 is used to store computer programs; preferably, the memory 610 includes: ROM, RAM, disk, USB flash drive, memory card or CD-ROM and other media that can store program codes.

[0080] Specifically, the memory 610 may include a computer system readable medium in the form of a volatile memory, such as a random access memory (RAM) and / or a cache memory. The electronic device 600 may further include other removable / non-removable, volatile / non-volatile computer system storage media. The memory 610 may include at least one program product having a set (e.g., at least one) of program modules, which are configured to perform the functions of the various embodiments of the present application. It is understood that the memory 610 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), which is used as an external cache. By way of exemplary but not limiting description, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM). The memory described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable categories of memory.

[0081] The processor 620 is connected to the memory 610 and is used to execute the computer program stored in the memory 610 so that the electronic device 600 executes the ancient Chinese punctuation prediction method described in any embodiment of the present application.

[0082] Optionally, the processor 620 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.

[0083] Optionally, the electronic device 600 in this embodiment may further include a display 630. The display 630 is communicatively connected with the memory 610 and the processor 620, and is used to display a graphical user interface (GUI) interaction interface related to the ancient Chinese punctuation prediction method described in the embodiment of the present application.

[0084] The present application also provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method for predicting ancient Chinese punctuation marks described in any embodiment of the present application is implemented.

[0085] The terms "component", "module", "system", etc. used in this specification are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process, a processor, an object, an executable file, an execution thread, a program and / or a computer running on a processor. By way of illustration, both applications and computing devices running on a computing device can be components. One or more components may reside in a process and / or an execution thread, and a component may be located on a computer and / or distributed between two or more computers. In addition, these components may be executed from various computer-readable media having various data structures stored thereon. Components may, for example, communicate through local and / or remote processes according to signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system and / or a network, such as the Internet interacting with other systems through signals).

[0086] The descriptions of the processes or structures corresponding to the above-mentioned figures have different emphases. For parts that are not described in detail in a certain process or structure, please refer to the relevant descriptions of other processes or structures.

[0087] The above embodiments are merely illustrative of the principles and effects of the present application and are not intended to limit the present application. Anyone familiar with the technology may modify or change the above embodiments without violating the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by a person of ordinary skill in the art without departing from the spirit and technical ideas disclosed in the present application shall still be covered by the claims of the present application.

Claims

1. A method for predicting punctuation marks in ancient Chinese, characterized in that: include: Get training data; Splitting the training data, and obtaining a training data set using the split data blocks; The training data is split and processed, and the training data set is obtained by using the split data blocks, including: sentence splitting and text splitting of the training data to obtain a short sentence set and a text set; obtaining a maximum number of short sentences and a maximum number of texts within a window range according to the short sentence set and the text set; and obtaining the training data set according to the maximum number of short sentences, the maximum number of texts, the short sentence set and the text set; Using a minimum hash algorithm to construct an index library to obtain a reference text of the training data set; the training data set includes at least one text to be retrieved, and using the minimum hash algorithm to construct an index library to obtain a reference text of the text to be retrieved includes: obtaining a hash signature vector of the training data set based on the minimum hash algorithm; using the hash signature vector to construct a signature index library based on a local sensitive hash forest; obtaining a hash signature vector of the text to be retrieved based on the minimum hash algorithm; using the hash signature vector of the text to be retrieved to search in the signature index library to obtain a reference text of the text to be retrieved; Using the reference text and the original text of the training data set to train the initial language model to obtain an ancient Chinese punctuation prediction model; The ancient Chinese punctuation prediction model is used to predict the text to be predicted to obtain a prediction result.

2. The ancient Chinese punctuation prediction method according to claim 1, characterized in that: Using the hash signature vector of the text to be retrieved to search in the signature index library, obtaining the reference text of the text to be retrieved includes: Using the hash signature vector of the text to be searched, searching in the signature index library, and obtaining at least one search result; The search results are screened using the Jaccard similarity, and the search results whose Jaccard similarity is greater than a set threshold are obtained as reference texts of the text to be searched.

3. The ancient Chinese punctuation prediction method according to claim 1, characterized in that: include: When the ancient Chinese punctuation prediction model is used to predict the text to be predicted, the next predicted character is restricted to be consistent with the character of the text to be predicted.

4. The ancient Chinese punctuation prediction method according to claim 1, characterized in that: include: When the ancient Chinese punctuation prediction model is used to predict the text to be predicted, at least one candidate prediction result is obtained, and the candidate prediction result with the largest probability value is used as the prediction result.

5. The ancient Chinese punctuation prediction method according to claim 1, characterized in that: Using the reference text and the original text of the training data set to train the initial language model to obtain the ancient Chinese punctuation prediction model includes: Using the original text of the training data set and the reference text to input the initial language model for training, to obtain an initial classical Chinese punctuation prediction model; The parameters of the initial ancient Chinese punctuation prediction model are adjusted by back propagation until the model converges to obtain the ancient Chinese punctuation prediction model.

6. A system for predicting punctuation marks in ancient Chinese, characterized in that: include: A data acquisition module, used to acquire training data; A data processing module is used to split the training data and obtain a training data set using the split data blocks; the data processing module is also used to: split the training data into sentences and texts to obtain a short sentence set and a text set; obtain the maximum number of short sentences and the maximum number of texts within a window range according to the short sentence set and the text set; and obtain the training data set according to the maximum number of short sentences, the maximum number of texts, the short sentence set and the text set; A reference text module, used to construct an index library using a minimum hash algorithm to obtain reference text of the training data set; The training data set includes at least one text to be retrieved, and the reference text module is further used to: obtain a hash signature vector of the training data set based on a minimum hash algorithm; use the hash signature vector to build a signature index library based on a local sensitive hash forest; obtain a hash signature vector of the text to be retrieved based on the minimum hash algorithm; use the hash signature vector of the text to be retrieved to search in the signature index library to obtain a reference text of the text to be retrieved; A model acquisition module, used to train an initial language model using the reference text and the original text of the training data set to obtain an ancient Chinese punctuation prediction model; The result prediction module is used to predict the text to be predicted by using the ancient Chinese punctuation prediction model to obtain a prediction result.

7. An electronic device, characterized in that: The electronic device comprises: Memory for storing computer programs; A processor, wherein the processor is used to execute the computer program stored in the memory so that the electronic device executes the ancient Chinese punctuation prediction method as described in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the ancient Chinese punctuation prediction method described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Ancient book punctuation filling method and device

    CN112199927A

  • Model construction method, device and system based on ALBERT and medium

    CN112906366A

  • System and method for inputting text into electronic devices

    US20130041857A1