Word segmentation model training method, word segmentation processing method and device

CN116151220BActive Publication Date: 2026-10-09MASHANG CONSUMER FINANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210916373.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-01
Publication Date
2026-10-09
Estimated Expiration
2042-08-01

AI Technical Summary

Technical Problem

[0004]本申请实施例的目的是提供一种分词模型训练方法、分词处理方法和装置,有利于提升文本分词的准确率和处理效率,解决常用方法中流程复杂、效率低、限制词长度等问题

Benefits of technology

[0021]In this embodiment, a lexicon associated with the text to be segmented is obtained. The lexicon includes multiple word segments and their corresponding word frequencies. Each word segment includes at least a portion of the words constituting the text to be segmented, and the word frequency indicates the number of times the corresponding word segment appears in the text. A preset number of sample sentences are generated based on the word segments and their corresponding word frequencies in the lexicon, and the sample sentences are labeled with word segments. The word segmentation labels indicate the segmentation results obtained by segmenting the sample sentences. The word segmentation model to be trained is iteratively trained based on the labeled sample sentences to obtain the trained word segmentation model, which is then used to segment the text. Therefore, this application generates sample sentences based on the lexicon associated with the text to be segmented, enabling the trained word segmentation model to adapt to the text and improving the accuracy and speed of text segmentation. By automatically generating sample sentences for training the word segmentation model based on the word segments in the lexicon and automatically labeling the sample sentences, no manual labeling is required, and sample sentences can be generated efficiently to train a high-accuracy word segmentation model. The word segmentation model is iteratively trained using multiple sample sentences after word segmentation annotation. This results in a simple and easy-to-understand model with advantages such as high speed and convenient optimization iteration. This application does not limit the length of the segmented words, enabling accurate segmentation of words of varying lengths. Furthermore, this application can be flexibly applied to various fields. By obtaining a thesaurus associated with the text to be segmented, sample sentences suitable for practical applications can be generated, thereby training a word segmentation model applicable to those applications. This demonstrates good transferability, versatility, and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116151220B_ABST
    Figure CN116151220B_ABST
Patent Text Reader

Abstract

The application discloses a word segmentation model training method, a word segmentation processing method and device. The method comprises: obtaining a word library associated with a text to be segmented, the word library comprising a plurality of word segments and word frequencies corresponding to the word segments, the word segments comprising at least part of words constituting the text to be segmented, and the word frequencies being used to indicate the number of times of occurrence of the corresponding word segments in the text to be segmented; generating a preset number of sample sentences according to the word segments and the corresponding word frequencies in the word library, and performing word segmentation annotation on the sample sentences, the word segmentation annotation being used to indicate a word segmentation processing result obtained by cutting the sample sentences; and iteratively training a word segmentation model to be trained based on the sample sentences after the word segmentation annotation, to obtain a trained word segmentation model, the trained word segmentation model being used to perform word segmentation processing on the text to be segmented. The application can improve the accuracy and speed of text word segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and in particular to a word segmentation model training method, word segmentation processing method and apparatus. Background Technology

[0002] In English, words are clearly separated by punctuation marks. Natural language processing (NLP) can easily divide an English sentence into words using simple operations, without needing to consider whether a word is a new word. Unlike English, the smallest unit of language used to express meaning in Chinese is the word, and there are no clear delimiters between words. Therefore, word segmentation is the primary processing operation for Chinese text in NLP.

[0003] Natural language processing (NLP) requires computers to accurately extract the meaning of sentences, involving syntax, semantic components, semantic structure, and context. All of this is based on accurately segmenting a sentence into combinations of multiple words. In deeper NLP processes, such as personalized recommendation, sentiment analysis, topic classification, and public opinion analysis, high-accuracy word segmentation is a prerequisite. However, the emergence of new words often interferes with existing word segmentation software, leading to unsatisfactory segmentation results and consequently affecting subsequent processing of Chinese text. How to train a word segmentation model to improve segmentation accuracy is the technical problem this application aims to solve. Summary of the Invention

[0004] The purpose of this application is to provide a word segmentation model training method, word segmentation processing method, and apparatus, which are beneficial to improving the accuracy and processing efficiency of text word segmentation and solving problems such as complex processes, low efficiency, and word length limitations in commonly used methods.

[0005] Firstly, a word segmentation model training method is provided, including:

[0006] Obtain a lexicon associated with the text to be segmented, the lexicon including multiple word segments and the word frequency corresponding to each word segment, the word segments including at least some words constituting the text to be segmented, and the word frequency being used to indicate the number of times the corresponding word segment appears in the text to be segmented;

[0007] A preset number of sample sentences are generated based on the word segments and their corresponding word frequencies in the lexicon, and the sample sentences are segmented and labeled. The word segmentation labels are used to indicate the word segmentation results obtained by segmenting the sample sentences.

[0008] The word segmentation model to be trained is iteratively trained based on multiple sample sentences after word segmentation annotation to obtain the trained word segmentation model, which is used to perform word segmentation processing on the text to be segmented.

[0009] Secondly, a word segmentation processing method is provided, including:

[0010] The process involves obtaining a text to be segmented and a trained segmentation model. The trained segmentation model is obtained by iteratively training a segmentation model based on multiple sample sentences with segmentation annotations. The multiple sample sentences are multiple sample sentences from a preset number of sample sentences. The segmentation annotations are used to indicate the segmentation results obtained by segmenting the sample sentences. The preset number of sample sentences are generated based on the segmented words and their corresponding word frequencies in a lexicon associated with the text to be segmented. The lexicon includes multiple segmented words and their corresponding word frequencies. The segmented words include at least some of the words that constitute the text to be segmented. The word frequencies are used to indicate the number of times the corresponding segmented word appears in the text to be segmented.

[0011] The text to be segmented is input into the trained word segmentation model to obtain the word segmentation result of the text to be segmented.

[0012] Thirdly, a word segmentation model training device is provided, including:

[0013] The first acquisition module acquires a lexicon associated with the text to be segmented. The lexicon includes multiple word segments and the word frequency corresponding to each word segment. The word segments include at least a portion of the words that constitute the text to be segmented. The word frequency is used to indicate the number of times the corresponding word segment appears in the text to be segmented.

[0014] The first generation module generates a preset number of sample sentences based on the word segments in the lexicon and their corresponding word frequencies, and performs word segmentation annotation on the sample sentences. The word segmentation annotation is used to indicate the word segmentation processing result obtained by segmenting the sample sentences.

[0015] The first training module iteratively trains the word segmentation model to be trained based on multiple sample sentences after word segmentation and annotation, and obtains the trained word segmentation model, which is used to perform word segmentation processing on the text to be segmented.

[0016] Fourthly, a word segmentation processing apparatus is provided, comprising:

[0017] The second acquisition module acquires the text to be segmented and the trained segmentation model. The trained segmentation model is obtained by iteratively training the segmentation model to be trained based on multiple sample sentences after segmentation annotation. The multiple sample sentences are multiple sample sentences from a preset number of sample sentences. The segmentation annotation is used to indicate the segmentation processing result obtained by segmenting the sample sentences. The preset number of sample sentences are generated based on the segmented words and their corresponding word frequencies in the lexicon associated with the text to be segmented. The lexicon includes multiple segmented words and the word frequencies corresponding to each segmented word. The segmented words include at least some of the words that constitute the text to be segmented. The word frequencies are used to indicate the number of times the corresponding segmented word appears in the text to be segmented.

[0018] The second processing module inputs the text to be segmented into the trained word segmentation model to obtain the word segmentation result of the text to be segmented.

[0019] Fifthly, an electronic device is provided, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method as described in the first or second aspect.

[0020] In a sixth aspect, a computer-readable storage medium is provided on which a computer program is stored, which, when executed by a processor, implements the steps of the method as described in the first or second aspect.

[0021] In this embodiment, a lexicon associated with the text to be segmented is obtained. The lexicon includes multiple word segments and their corresponding word frequencies. Each word segment includes at least a portion of the words constituting the text to be segmented, and the word frequency indicates the number of times the corresponding word segment appears in the text. A preset number of sample sentences are generated based on the word segments and their corresponding word frequencies in the lexicon, and the sample sentences are labeled with word segments. The word segmentation labels indicate the segmentation results obtained by segmenting the sample sentences. The word segmentation model to be trained is iteratively trained based on the labeled sample sentences to obtain the trained word segmentation model, which is then used to segment the text. Therefore, this application generates sample sentences based on the lexicon associated with the text to be segmented, enabling the trained word segmentation model to adapt to the text and improving the accuracy and speed of text segmentation. By automatically generating sample sentences for training the word segmentation model based on the word segments in the lexicon and automatically labeling the sample sentences, no manual labeling is required, and sample sentences can be generated efficiently to train a high-accuracy word segmentation model. The word segmentation model is iteratively trained using multiple sample sentences after word segmentation annotation. This results in a simple and easy-to-understand model with advantages such as high speed and convenient optimization iteration. This application does not limit the length of the segmented words, enabling accurate segmentation of words of varying lengths. Furthermore, this application can be flexibly applied to various fields. By obtaining a thesaurus associated with the text to be segmented, sample sentences suitable for practical applications can be generated, thereby training a word segmentation model applicable to those applications. This demonstrates good transferability, versatility, and scalability. Attached Figure Description

[0022] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0023] Figure 1 This is one of the flowcharts illustrating a word segmentation model training method provided in an embodiment of this application.

[0024] Figure 2 This is a second flowchart illustrating a word segmentation model training method provided in one embodiment of this application.

[0025] Figure 3 This is the third flowchart of a word segmentation model training method provided in one embodiment of this application.

[0026] Figure 4 This is the fourth flowchart illustrating a word segmentation model training method provided in one embodiment of this application.

[0027] Figure 5This is the fifth flowchart illustrating a word segmentation model training method provided in one embodiment of this application.

[0028] Figure 6 This is the sixth flowchart of a word segmentation model training method provided in one embodiment of this application.

[0029] Figure 7 This is the seventh flowchart of a word segmentation model training method provided in one embodiment of this application.

[0030] Figure 8 This is the eighth flowchart of a word segmentation model training method provided in one embodiment of this application.

[0031] Figure 9 This is a flowchart illustrating a word segmentation method provided in one embodiment of this application.

[0032] Figure 10 This is a schematic diagram of the structure of a word segmentation model training device provided in one embodiment of this application.

[0033] Figure 11 This is a schematic diagram of the structure of a word segmentation processing device provided in one embodiment of this application.

[0034] Figure 12 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. Detailed Implementation

[0035] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0036] Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application. The drawing numbers in this application are only used to distinguish the various steps in the scheme and are not used to limit the execution order of the steps; the specific execution order is as described in the specification.

[0037] Deep learning technology can be applied to sentence segmentation. By training a segmentation model, it can perform segmentation on sentences to be segmented, thereby facilitating further analysis of the sentence's grammar, components, and meaning. This can be achieved using rule-based or statistical segmentation methods.

[0038] Rule-based methods typically involve language experts constructing templates based on morphological principles, semantic information, or part-of-speech information, and then matching the text. This method is highly accurate and targeted, but the rules are usually domain-specific, and the complexity of manually writing and maintaining rules is relatively high. Rules are usually not used directly but rather as an additional module combined with other methods, which has limitations.

[0039] Statistical word segmentation methods identify words in text by calculating statistical features such as word frequency, word formation probability, left and right adjacency entropy, and adjacency variation number using a large amount of experimental data. Statistical methods are relatively flexible, not limited by domain, easily scalable, and highly portable, but they suffer from drawbacks such as data sparsity and low accuracy.

[0040] For example, while statistical word segmentation methods have advantages in some aspects, they typically impose limitations on word length, such as setting a maximum word length. Furthermore, the calculation of left and right entropy in statistical methods often leads to inefficiency in the entire word segmentation process. Additionally, they suffer from drawbacks such as data sparsity and low accuracy.

[0041] To address the problems existing in the prior art, embodiments of this application provide a word segmentation model training method, such as... Figure 1 As shown, it includes:

[0042] S11: Obtain a lexicon associated with the text to be segmented. The lexicon includes multiple word segments and the word frequency corresponding to each word segment. The word segments include at least some of the words that constitute the text to be segmented. The word frequency is used to indicate the number of times the corresponding word segment appears in the text to be segmented.

[0043] Different thesaurus stores specialized vocabulary from different fields. For example, there are thesaurus containing commonly used words in daily life, thesaurus containing technical vocabulary from the financial field, and thesaurus containing technical vocabulary from the communications field. In this application example, the corresponding domain thesaurus can be selected based on the domain to which the text to be segmented belongs. The obtained thesaurus can be a Chinese thesaurus. The thesaurus includes multiple word segments and their corresponding word frequencies. The word segments in the thesaurus include at least some of the words that constitute the text to be segmented. In other words, the thesaurus and the text to be segmented contain common words.

[0044] The lexicon contains the word frequency for each word segmentation, which indicates the number of times the word segmentation occurs in the text to be segmented. Generally, the word frequency of a word often differs in different application domains. This word frequency can be expressed in various forms, such as expressing the number of times the word segmentation occurs in the text to be segmented as a positive integer.

[0045] S12: Generate a preset number of sample sentences based on the word segments in the dictionary and their corresponding word frequencies, and perform word segmentation annotation on the sample sentences. The word segmentation annotation is used to indicate the word segmentation processing result obtained by segmenting the sample sentences.

[0046] In this step, word segments from the lexicon are used as components of sentences to construct sample sentences, resulting in a preset number of sample sentences. This preset number can be set according to actual needs.

[0047] Specifically, word segments can be selected from the lexicon based on their corresponding word frequencies to form sample sentences. The probability of a word segment being selected can be positively correlated with its word frequency; that is, the higher the word frequency of a word segment, the higher the probability of it being selected.

[0048] In this step, several word segments can be selected based on their frequency to form components of a sample sentence. Then, these selected word segments are arranged to create a sample sentence. Optionally, punctuation can be added based on the number of word segments to generate a sample sentence containing punctuation.

[0049] Optionally, word segments in the dictionary can be selected repeatedly. For example, multiple generated sample sentences may contain the same word segment. Or, a single generated sample sentence may contain multiple identical word segments.

[0050] In this application, since sample sentences are generated based on word segmentation from a lexicon, the segmentation tags can be retained while combining the sample sentences to achieve word segmentation annotation of the sample sentences. These word segmentation tags are used to identify which words constitute the sentence and can also indicate the "correct answer" of word segmentation for the sample sentence.

[0051] Optionally, the sample sentences can be automatically segmented and tagged in this step. The tagging format can be selected as "BIO" or "BMESO", etc. This automatic tagging step can save manpower and resources, and can also avoid problems such as incorrect tagging and omissions.

[0052] S13: Based on multiple sample sentences after word segmentation annotation, the word segmentation model to be trained is iteratively trained to obtain the trained word segmentation model, which is used to perform word segmentation processing on the text to be segmented.

[0053] Multiple sample sentences with word segmentation annotations are used as training samples and input into the word segmentation model to be trained to obtain the word segmentation model. Through sample sentences with word segmentation annotations, the ability of the word segmentation model to segment words in a sentence can be effectively trained.

[0054] The models used in this application may be, for example, Long Short-Term Memory (LSTM) or the pre-trained language model BERT (Encoder Representation from Transformers).

[0055] The solution provided in this application generates sample sentences based on the word frequencies of word segmentation in the lexicon. This enables the word frequencies of word segmentation in the sample sentences to fit with the word frequencies in the text to be segmented, thereby allowing the trained word segmentation model to perform word segmentation accurately and effectively. This makes the word segmentation results more consistent with real usage habits in the field, thus improving the accuracy of word segmentation.

[0056] The examples provided in this application demonstrate that, based on the characteristics of Chinese text, sample sentences can be generated from a dictionary associated with the text to be segmented, effectively improving the accuracy of text segmentation. The segmentation model is trained on the text to be segmented, enabling accurate segmentation and exhibiting high processing efficiency.

[0057] Compared to the complex model training methods commonly used, the example in this application has the advantage of automatically generating high-quality sample sentences, resulting in high training and application efficiency. Furthermore, the example in this application has no limit on the length of the segmented words, enabling flexible segmentation of words of varying lengths in the text to be segmented.

[0058] This application example generates sample sentences based on a lexicon associated with the text to be segmented, enabling the trained segmentation model to be adapted to the text to be segmented, thus improving the accuracy and speed of text segmentation. Furthermore, by automatically generating sample sentences for training the segmentation model based on the segmented words in the lexicon, and automatically performing segmentation annotation on the sample sentences, no manual annotation is required. This efficient generation of sample sentences allows for the training of a highly accurate segmentation model.

[0059] Moreover, the example in this application iteratively trains the word segmentation model to be trained based on multiple sample sentences after word segmentation and annotation, making the trained word segmentation model simple in structure and easy to understand, with the advantages of fast speed and simple optimization iteration.

[0060] In addition, this application does not limit the length of word segmentation, and can be used to accurately segment words of different lengths.

[0061] Furthermore, the examples in this application can be flexibly applied to various fields. By obtaining the vocabulary associated with the text to be segmented, sample sentences suitable for practical application fields can be generated, and then a segmentation model suitable for practical application fields can be trained. It has good transferability, versatility and scalability.

[0062] Based on the solutions provided in the above embodiments, optionally, such as Figure 2As shown, the text to be segmented includes multiple sentences to be segmented. Step S11 above includes:

[0063] S21: Obtain a dictionary containing common words in each of the multiple sentences to be segmented.

[0064] In this example, the text to be segmented includes multiple sentences, where each sentence is the one that will be segmented using the segmentation model. In practical applications, the text to be segmented can also be an article, several paragraphs, etc. For an article or paragraph, it can first be split into multiple sentences based on punctuation marks to obtain the text to be segmented.

[0065] In addition, multiple sentences to be segmented in the text can be arranged in order, which can be determined by the position of the sentences to be segmented in the original article or paragraph.

[0066] In the text to be segmented, each sentence consists of at least one segmented word. The vocabulary obtained in this step contains common segmented words with each sentence, so that the segmented words in the obtained vocabulary match the text to be segmented. This enables the generation of sample sentences that are closer to the text to be segmented in subsequent steps, thereby improving the segmentation accuracy of the trained segmentation model for each sentence in the text to be segmented.

[0067] S22: Determine the word frequency corresponding to each word based on the number of times each word in the dictionary appears in the text to be segmented.

[0068] Different fields use different words. In this application example, a related thesaurus is obtained for the text to be segmented, so that the words in the thesaurus are similar to the words used in the text to be segmented.

[0069] In practical applications, the relevant thesaurus can be selected based on the domain of the text to be segmented. For example, if the text to be segmented relates to the industrial field, then a thesaurus within that field can be obtained, ensuring that the segmented words in the obtained thesaurus are similar to those used in the text to be segmented.

[0070] Specifically, the frequency of a word segment in the dictionary appearing in the text to be segmented can be determined by searching, and this frequency is then used as the word frequency of the segment.

[0071] In practical applications, the word frequency of a word can also be determined by combining the frequency of its daily use and the number of times it appears in the text to be segmented.

[0072] For example, if all words in the dictionary have appeared in the text to be segmented, then the frequency of the word is taken as the word frequency. For instance, if "complaint" appears 10 times in the text to be segmented, then the word frequency of "complaint" is 10.

[0073] If at least some words in the dictionary do not appear in the text to be segmented, then the frequency of these unseen words is 1, and the frequency of all words that have appeared is the number of times they appeared + 1.

[0074] The solution provided by the embodiments of this application obtains a lexicon and each sentence to be segmented contains common segments. The word frequency is determined based on the number of times each segment appears in the text to be segmented, so that the determined word frequency represents the frequency of the segment in the text to be segmented.

[0075] This application example matches the vocabulary with the text to be segmented from the perspectives of word segmentation and word frequency, thereby making the generated sample sentences closer to the sentences to be segmented, and enabling the word segmentation model trained in subsequent steps to have higher word segmentation accuracy for the text to be segmented.

[0076] Based on the solutions provided in the above embodiments, optionally, such as Figure 3 As shown, step S12 above includes:

[0077] S31: Determine the number M of sentences to be segmented contained in the text to be segmented.

[0078] For text containing punctuation marks, the number M of sentences to be segmented in the text can be determined based on the punctuation marks. For example, the text can be divided according to the punctuation marks to obtain multiple sentences to be segmented, thus determining the number M.

[0079] Some text to be segmented may have special formats, such as using spaces, newlines, or other format specifiers to divide the text into multiple sentences to be segmented to determine the number M.

[0080] S32: Determine the number of sample sentences to be generated N based on the number M of sentences to be segmented, where N and M are positive integers and the ratio of N to M is greater than a preset ratio.

[0081] For example, assuming the preset ratio is 1 / 3, then the total number of sample sentences N generated will be greater than M / 3, meaning the total number of sample sentences N generated is greater than one-third of the number of sentences M to be segmented in the text to be segmented. The preset ratio can be adjusted according to parameters such as the number of words in the dictionary and the number of sentences to be segmented.

[0082] S33: Generate N sample sentences based on the word segments and their corresponding word frequencies in the lexicon, and perform word segmentation and annotation on the N sample sentences.

[0083] Generally speaking, the more sentences and complex the words in the text to be segmented, the more difficult it is to train a segmentation model. To address this, this application's embodiments determine the number of sample sentences to generate based on the number of sentences in the text, thereby training the segmentation model with an appropriate number of sample sentences and optimizing the segmentation performance of the trained model.

[0084] Based on the solutions provided in the above embodiments, optionally, such as Figure 4 As shown, step S12 above includes:

[0085] S41: Determine the set of character counts of the text to be segmented, wherein the character count of each sentence to be segmented in the text to be segmented belongs to the set of character counts.

[0086] In this step, the aforementioned character count set is determined based on each sentence in the text to be segmented. Each item in the character count set is a positive integer, representing the number of characters in the sentence to be segmented. Optionally, the character count set described in this example does not contain duplicate items.

[0087] S42: The ratio of the total number of characters in the word segmentation in the dictionary to the total number of word segments in the dictionary is determined as the average length of the word segmentation.

[0088] In this embodiment, the execution order of steps S41 and S42 can be interchanged, or they can be executed simultaneously. In this step, the average length of the words in the dictionary is determined. A word may consist of multiple characters. For example, "resistance" is a word composed of 2 characters, and "integrated circuit" is a word composed of 4 characters. The total number of characters in these two words is 6 (the sum of 2 and 4), and the total number of words is 2 (the number of words "resistance" and "integrated circuit" mentioned above). Their ratio is 3, meaning the average length of these two words is 3. In this step, the ratio of the total number of characters in the words contained in the dictionary to the total number of words in the dictionary is determined as the average length of the words.

[0089] In practical applications, the calculated ratio may be non-integer, and the non-integer average word segmentation length can be directly applied. Alternatively, rounding, rounding up or down, etc., can be used to determine an integer as the average word segmentation length.

[0090] S43: Generate a preset number of sample sentences based on the character count set, the average length of the word segmentation, the word segments in the word library and their corresponding word frequencies, and perform word segmentation annotation on the sample sentences. The number of word segments contained in the generated sample sentences is the ratio of a random number in the character count set to the average length of the word segmentation.

[0091] Specifically, for the generation of each sentence, the length L of the sample sentence is determined by a random number x, satisfying the following formula: L = x / l.

[0092] Here, the length of the sample sentence is L, indicating that the sample sentence contains L words. l represents the average length of words in the lexicon, and the value of the random number x is randomly selected from the above character count set. For example, if x is selected as 5 from the character count set, the average word count length l is 2, and according to the above formula, x / l is 2.5.

[0093] Optionally, if the calculated ratio is not an integer, it can be rounded up or down to determine the number of words to be segmented for easier subsequent processing. For example, the length L of the sample sentence satisfies the following formula: L = ROUNDUP(x / l).

[0094] ROUNDUP() means rounding up. Based on the example above, after determining that x / l is 2.5, rounding 2.5 up can calculate the length L of the sample sentence as 3.

[0095] The scheme provided in this application can determine the length of the generated sample sentence. Specifically, determining the sentence length based on the average length of the segmented words and a random number from the character set ensures that the determined sentence length is close to the sentence to be segmented, thereby better training the model and improving the segmentation accuracy of the trained model.

[0096] Based on the solutions provided in the above embodiments, optionally, such as Figure 5 As shown, after step S41 above, the following steps are also included:

[0097] S51: Divide the set of word counts into multiple word count subsets, where each word count subset contains the word count of at least one sentence to be segmented.

[0098] For example, suppose the word count set contains integers greater than or equal to 1 and less than or equal to 40. Then, in this step, the word count set can be divided into three subsets: [1, 10], [11, 30], and [31, max]. In this example, max is 40.

[0099] It should be understood that the above intervals can be used to mark and distinguish different word count subsets. For example, [1, 10] means that the word count values ​​in the word count subset belong to the word count subset of this interval. However, it does not mean that the sentences to be segmented in the text to be segmented must necessarily have the 10 lengths from 1 to 10.

[0100] For example, the [1,10] word count subset can specifically contain the five elements 1, 4, 6, 8, and 9, representing that the text to be segmented contains at least one sentence with a length of 1, at least one sentence with a length of 4, at least one sentence with a length of 6, at least one sentence with a length of 8, and at least one sentence with a length of 9. That is, the above [1,10] word count subset can be specifically represented as {1,4,6,8,9}.

[0101] S52: Determine the weights corresponding to the plurality of word count subsets respectively, wherein the weight corresponding to any word count subset represents the probability that the word count of the sentence to be segmented in the text to be segmented belongs to the corresponding word count subset.

[0102] For example, the text to be segmented contains 100 sentences. Among them, 20 sentences have a length range of [1, 10], accounting for 20%. Sentences with a length range of [11, 30] account for 50%, and sentences with a length range of [31, max] account for 30%. Accordingly, the percentage corresponding to each length range can be used as the weight of the corresponding subset of word counts. That is, the weight of the subset with [1, 10] words is 0.2, the weight of the subset with [11, 30] words is 0.5, and the weight of the subset with [31, max] words is 0.3.

[0103] Step S43 above includes:

[0104] S53: Determine the target word count subset based on the weights corresponding to each word count subset in the word count set.

[0105] The target word count subset determined in this step is a subset of words selected from the above word count set based on weights. The probability of a word count subset being selected as the target word count subset is positively correlated with its corresponding weight. In this example, assume that the target word count subset determined based on the corresponding weights is the word count subset of [11, 30].

[0106] S54: Select a random number from the target word count subset, and determine the target number of words contained in the target sample sentence based on the random number and the average length of the word segments.

[0107] This step involves selecting a random number from the chosen target word count subset. For example, assuming the target word count subset is [11, 30], this step randomly selects one element from the elements contained in this subset as the random number. Subsequently, the ratio of the selected random number to the average length of the word segments is determined as the target number of words contained in the target sample sentence.

[0108] S55: Construct a target sample sentence containing the target number of segmented words based on the lexicon, and perform word segmentation and annotation on the target sample sentence.

[0109] Through the above steps, a reasonable subset of target characters is selected, and random numbers are drawn to determine the number of word segments to construct the target sample sentence. The probability of each word segment being selected in the dictionary is related to its corresponding weight, and the weight of a word depends on its frequency. Assuming the sentence length is L, L words are selected based on their weights, and finally, these L words are randomly arranged and combined to form the target sample sentence.

[0110] For example, if there are 100 sentences in the text to be segmented, and the length of each sentence is a positive integer, then the proportion of sentences with lengths in the range [1, 10] is 20%, the proportion of sentences with lengths in the range [11, 30] is 50%, and the proportion of sentences with lengths in the range [31, max] is 30%. Therefore, the probability that the random number x is in the range [1, 10] is 20%, the probability of it being in the range [11, 30] is 50%, and the probability of it being in the range [31, max] is 30%, where max represents the maximum length of a sentence in the text to be segmented.

[0111] Through the examples in this application, the word frequency of word segmentation can be used as a weight to make the probability of word segmentation being selected correspond to the frequency of word segmentation, thereby generating target sample sentences that are closer to the sentence to be segmented, and thus enabling the trained word segmentation model to have greater word segmentation accuracy for the text to be segmented.

[0112] Optionally, the above character count set can also be divided into more subsets. For example, each character value in the character count set can be divided into a separate subset, and each character value can have a corresponding weight.

[0113] Based on the solutions provided in the above embodiments, optionally, such as Figure 6 As shown, step S55 above includes:

[0114] S61: Select a target number of target words from the word library based on word frequency, wherein the probability of a word being selected is positively correlated with the word frequency of the word.

[0115] In this application example, the word frequency of word segmentation is used as the weight, so that the probability of word segmentation being selected corresponds to the frequency of word segmentation, thereby generating target sample sentences that are closer to the sentence to be segmented, and thus enabling the trained word segmentation model to have better word segmentation accuracy for the text to be segmented.

[0116] S62: Randomly arrange and combine the target words to generate a pre-constructed first target sample sentence.

[0117] After selecting the target number of target words, these target words are arranged in a random order to generate the first target sample sentence containing the target number of consecutive target words.

[0118] S63: Add punctuation to the pre-constructed first target sample sentence according to the target number to generate a second target sample sentence, and perform word segmentation and tagging on the second target sample sentence, wherein the number of punctuation added to the pre-constructed first target sample sentence is positively correlated with the target number.

[0119] The scheme provided in this application can generate sample sentences containing punctuation. The number of punctuation marks to be added to each sentence depends on the length of the sentence. For example, for sentences with a length range of [1, 10], one punctuation mark is added to the end of the sentence; for sentences with a length range of [11, 30], one punctuation mark is added to the end of the sentence, and one punctuation mark is randomly added between words in the sentence; for sentences with a length range of [31, max], one punctuation mark is added to the end of the sentence, and two punctuation marks are randomly added between words in the sentence. The type of punctuation mark added can be randomly selected from {, . ? !}.

[0120] The solution provided in this application embodiment generates a second target sample sentence containing punctuation that matches the sentence length, making the generated second target sample sentence closer to sentences containing punctuation in actual applications. This makes the generated second target sample sentence similar to the sentence to be segmented, enabling the trained segmentation model to perform segmentation on the text to be segmented more accurately.

[0121] Based on the solutions provided in the above embodiments, optionally, such as Figure 7 As shown, step S13 above includes:

[0122] S71: For each sample sentence, input the sample sentence into the segmentation model that has been trained in stages to obtain the predicted segmentation result of the sample sentence.

[0123] In this example, the word segmentation model undergoes phased training. For instance, a predetermined number of iterations is set, and the word segmentation model to be trained is subjected to a predetermined number of iterations to obtain a phased-trained word segmentation model. In this step, each sample sentence is input into the aforementioned phased-trained word segmentation model to obtain the predicted word segmentation result for the sample sentence. This predicted word segmentation result can characterize the word segmentation accuracy of the aforementioned phased-trained word segmentation model.

[0124] S72: Based on the predicted word segmentation results and word segmentation annotations of the sample sentences, determine the accuracy of the predicted word segmentation results of the sample sentences.

[0125] The above steps yield the segmentation results predicted by the staged training model for each sample sentence. In this step, the predicted segmentation results for each sample sentence are compared with the segmentation annotations of that sample sentence to determine whether the staged training model has correctly performed the segmentation.

[0126] For example, if the predicted word segmentation result for the same sample sentence is consistent with the word segmentation annotation, then the word segmentation model trained in stages is determined to have correctly segmented that sample sentence. If the predicted word segmentation result for the same sample sentence is inconsistent with the word segmentation annotation (e.g., missegmentation, omissions), then the word segmentation model trained in stages is determined to have incorrectly segmented that sample sentence. In this scheme, the ratio of the number of correctly segmented sample sentences to the total number of sample sentences input into the word segmentation model trained in stages can be defined as the accuracy of the predicted word segmentation result.

[0127] Furthermore, in practical applications, the criteria for correct word segmentation can be adjusted according to the actual situation. For example, if the consistency between the predicted word segmentation result and the word segmentation annotation for the same sample sentence is greater than 80%, then the word segmentation of that sample sentence is determined to be correct.

[0128] Alternatively, if there are fewer than two instances of misclassification or omission, the sentence segmentation process for that sample is considered correct.

[0129] Optionally, for sample sentences where the predicted word segmentation results are inconsistent with the word segmentation annotations, manual review can be used to further determine whether the word segmentation model has made a mistake in segmenting the sample sentences, thereby improving the accuracy of the accuracy rate.

[0130] S73: When the accuracy of the predicted word segmentation result is lower than the preset accuracy, adjust the generation parameters of the sample sentences to regenerate the sample sentences and iteratively train the word segmentation model trained in stages to obtain the trained word segmentation model. The generation parameters include at least one of the word frequency corresponding to the segmentation and the number of sample sentences.

[0131] In this step, if the accuracy rate is lower than the preset accuracy rate, it indicates that the word segmentation ability of the word segmentation model trained in this stage is not up to standard and cannot accurately perform word segmentation processing, requiring further iterative training.

[0132] Subsequently, the generation parameters of the sample sentences are adjusted to regenerate the sample sentences and further iteratively train the segmentation model that has been trained in stages until a segmentation model with an accuracy exceeding the preset accuracy is trained to meet the needs of practical applications.

[0133] In addition to the accuracy rate mentioned above, more segmentation performance metrics can be set according to actual needs to further determine whether the segmentation model meets the requirements of practical applications. For example, the segmentation processing speed can be preset to improve the segmentation efficiency of the model.

[0134] Optionally, the generation parameters for the above sample sentences include at least one of the word frequency corresponding to the word segmentation and the number of sample sentences.

[0135] Because language development and word usage frequency have a certain time sensitivity, previously high-frequency words may become low-frequency words, and new words may emerge to replace older ones. Therefore, sample sentences generated based on a lexicon may deviate somewhat from sentences used in existing applications. Based on this, the word frequencies corresponding to segmented words can be adjusted according to changes in word usage in actual applications, making the adjusted word frequencies closer to the usage frequencies of words in the actual application scenario. Subsequently, new sample sentences can be generated based on the adjusted word frequencies to further iterate and train the staged word segmentation model, optimizing its segmentation performance.

[0136] Furthermore, the number of sample sentences used to train the model also affects the model's word segmentation accuracy. If the number of sample sentences is too small, leading to insufficient model training, it will also result in poor word segmentation performance. In such cases, the number of sample sentences can be increased, for example, by using data augmentation techniques to generate new sample sentences.

[0137] After adjusting the above-mentioned generation parameters and generating new sample sentences, the new sample sentences can be used to perform iterative training on the word segmentation model to optimize the word segmentation effect. The solution provided in this application embodiment enables iterative training based on the word segmentation effect of the model, further improving the word segmentation capability of the model.

[0138] Furthermore, the word segmentation model can be reinforced in the lexicon based on the actual missegmented words to generate new sample sentences and perform reinforcement training. For example, the rule for adding missegmented words could be: if the word already exists in the lexicon, increase its frequency by doubling the original frequency. If the word does not exist in the lexicon, add it and set its frequency according to actual needs.

[0139] The solution provided in the embodiments of this application can be found in [reference]. Figure 8 This method improves the accuracy and speed of text segmentation by acquiring a lexicon, calculating word frequencies, generating sentences, automatically annotating, training models, and iteratively optimizing. Notably, it achieves a relatively accurate segmentation model without requiring manual annotation or domain knowledge.

[0140] Moreover, the word segmentation model has the advantages of simple structure, easy understanding, fast speed, and convenient optimization iteration. In addition, this word segmentation model does not limit the length of words, can effectively solve out-of-vocabulary words, and has good transferability, universality, and scalability.

[0141] The process involves several key components: First, a comprehensive lexicon is used for sentence generation. A large, complete lexicon sourced from texts similar to the text to be segmented significantly improves segmentation accuracy. Second, word frequency analysis, through the design of frequency rules, generates sample sentences that more closely resemble the sentences in the original text, further enhancing segmentation accuracy. Third, sentence generation, through the design of sentence generation rules, produces diverse sentences that more closely approximate the samples in the original text (e.g., sentence length distribution, word frequency), also increasing segmentation accuracy. Fourth, automatic annotation saves time and resources compared to traditional methods and avoids mislabeling and omissions. Finally, model training and iterative optimization further improve the accuracy of the segmentation model.

[0142] This application example can be fine-tuned to adapt to other tasks or application scenarios, and other excellent solutions can be embedded into the implementation logic, thus possessing good portability, versatility, and scalability. For example, after word segmentation, based on this application example, word frequency can be counted, and filtering can be performed by setting a word frequency threshold. This can be used for new word discovery or adding new words to the dictionary, thereby enriching the dictionary. It can effectively solve text segmentation problems such as complex processes, low efficiency, and limitations on word length.

[0143] The following examples illustrate practical applications of this application. For instance, this application can be applied to the identification and classification of online articles. Due to the complexity and wide range of topics covered in online articles, accurate classification using only a single word segmentation model is difficult. The solution provided by the embodiments of this application can effectively improve word segmentation accuracy.

[0144] Specifically, for a web article to be segmented, the word segmentation model training method provided in this application first obtains a lexicon associated with the web article. Assuming the web article relates to the economic field, a lexicon containing economic-specific vocabulary can be obtained. This lexicon includes multiple word segments and their corresponding word frequencies, ensuring that the lexicon contains vocabulary from the web article to be segmented, and that the word frequencies in the lexicon represent the number of times the corresponding word segment appears in the web article.

[0145] Then, a predetermined number of sample sentences are generated based on the word segments and their corresponding word frequencies in the lexicon, and the sample sentences are then segmented and labeled. Since the lexicon contains specialized vocabulary in the economic field, the generated sample sentences are close to sentences in online articles that are also in the economic field and require word segmentation.

[0146] Next, the word segmentation model to be trained is iteratively trained based on multiple sample sentences after word segmentation and annotation to obtain the trained word segmentation model.

[0147] Subsequently, the online article to be segmented is input into the trained segmentation model to perform segmentation on the online article.

[0148] Since the vocabulary obtained in this application example belongs to the same economic field as the online article to be segmented, and contains proper nouns in the economic field, the generated sample sentences are closer to the sentences to be segmented in the online article. This enables the trained segmentation model to more accurately segment online articles in the economic field.

[0149] Optionally, for other text documents belonging to the economic field, the word segmentation model trained above can also be used to perform word segmentation processing without retraining the word segmentation model.

[0150] To address the problems existing in the prior art, embodiments of this application also provide a word segmentation processing method, such as... Figure 9 As shown, it includes:

[0151] S91: Obtain the text to be segmented and the trained segmentation model, wherein the trained segmentation model is obtained by iteratively training the segmentation model to be trained based on multiple sample sentences after segmentation annotation. The multiple sample sentences are multiple sample sentences from a preset number of sample sentences. The segmentation annotation is used to indicate the segmentation processing result obtained by segmenting the sample sentences. The preset number of sample sentences are generated based on the segmented words and their corresponding word frequencies in the lexicon associated with the text to be segmented. The lexicon includes multiple segmented words and the word frequencies corresponding to each segmented word. The segmented words include at least some words that constitute the text to be segmented. The word frequency is used to indicate the number of times the corresponding segmented word appears in the text to be segmented.

[0152] The word segmentation model trained in this step can be obtained by training the word segmentation model according to any of the above embodiments. The text to be segmented can specifically be texts such as online novels, news reports, and comments.

[0153] For text to be segmented that contains special formats, the format of the text to be segmented can be standardized first, so that it can be input into the segmentation model for processing. Standardizing the format may include dividing the text to be segmented into multiple sentences according to punctuation marks or paragraph marks, and then inputting the sentences to be segmented into the segmentation model one by one in the order of the text to perform segmentation processing sentence by sentence.

[0154] S92: Input the text to be segmented into the trained word segmentation model to obtain the word segmentation result of the text to be segmented.

[0155] The sample sentences used to train the word segmentation model are automatically generated based on word segmentation data from the dictionary. Furthermore, the word segmentation of the sample sentences is automatically labeled without manual labeling, which enables the training of a highly accurate word segmentation model.

[0156] Moreover, iteratively training the word segmentation model based on multiple sample sentences after word segmentation annotation can make the trained word segmentation model simple in structure and easy to understand, with the advantages of fast speed and simple optimization iteration.

[0157] Moreover, the word segmentation model does not limit the length of the segments and can be used to accurately segment words of different lengths. By obtaining a lexicon associated with the text to be segmented, sample sentences suitable for practical applications are generated, and then a word segmentation model suitable for practical applications is trained, which has good transferability, versatility and scalability.

[0158] In this step, the text to be segmented is input into the trained word segmentation model. This word segmentation model can efficiently and accurately perform word segmentation on the text to be segmented and output the word segmentation result, which can express the segmentation result of the text to be segmented.

[0159] Furthermore, the word segmentation results output by the word segmentation model can also be used for text proofreading, text parsing, and other tasks. Examples of practical applications are given below.

[0160] For example, the word segmentation results of the text to be segmented obtained in this application can be applied to perform semantic parsing on the text to be segmented.

[0161] Specifically, based on the word segmentation results, the text to be segmented can be divided into multiple words. By identifying the semantics of each word through a pre-trained model, the semantics expressed by each word in the sentence can be obtained, thereby determining the overall meaning of the sentence and the overall meaning of the text.

[0162] For example, the word segmentation results of the text to be segmented obtained in this application can be used to proofread the text. Specifically, based on the word segmentation results, sentences with abnormal word segmentation results containing an excessive number of segments can be filtered out, and then manual proofreading can be used to check for typos caused by typos.

[0163] Alternatively, the word segmentation results of the text to be segmented obtained in this application example can also be used to further analyze the word frequency of each word in the text. Among them, words with too low word frequency may be misspelled due to typos, which facilitates targeted correction.

[0164] In addition, the word segmentation results of the text to be segmented obtained in this application example can also be used to optimize the text to be segmented, search for other related texts, etc., and can be flexibly applied according to actual needs.

[0165] To address the problems existing in the prior art, this application also provides a word segmentation model training device 100, such as... Figure 10 As shown, it includes:

[0166] The first acquisition module 101 acquires a lexicon associated with the text to be segmented. The lexicon includes multiple word segments and the word frequency corresponding to each word segment. The word segments include at least some of the words that constitute the text to be segmented. The word frequency is used to indicate the number of times the corresponding word segment appears in the text to be segmented.

[0167] The first generation module 102 generates a preset number of sample sentences based on the word segments in the lexicon and their corresponding word frequencies, and performs word segmentation annotation on the sample sentences. The word segmentation annotation is used to indicate the word segmentation processing result obtained by segmenting the sample sentences.

[0168] The first training module 103 iteratively trains the word segmentation model to be trained based on multiple sample sentences after word segmentation and annotation, and obtains the trained word segmentation model, which is used to perform word segmentation processing on the text to be segmented.

[0169] The apparatus provided in this application provides a lexicon associated with the text to be segmented. The lexicon includes multiple word segments and their corresponding word frequencies. Each word segment includes at least a portion of the words constituting the text to be segmented. The word frequency indicates the number of times the corresponding word segment appears in the text to be segmented. A preset number of sample sentences are generated based on the word segments in the lexicon and their corresponding word frequencies. The sample sentences are then labeled with word segments. The word segments are labeled with word segments to indicate the segmentation results obtained by segmenting the sample sentences. The word segmentation model to be trained is iteratively trained based on the multiple labeled sample sentences to obtain the trained word segmentation model. The trained word segmentation model is used to perform word segmentation processing on the text to be segmented.

[0170] Among them, by generating sample sentences based on the vocabulary associated with the text to be segmented, the trained word segmentation model can be adapted to the text to be segmented, thereby improving the accuracy and speed of text segmentation.

[0171] This application example automatically generates sample sentences for training a word segmentation model based on word segmentation data from a lexicon, and automatically performs word segmentation annotation on the sample sentences without manual annotation. It can efficiently generate sample sentences to train a word segmentation model with high accuracy.

[0172] This application example uses multiple sample sentences after word segmentation and annotation to iteratively train the word segmentation model to be trained, making the trained word segmentation model simple in structure and easy to understand, with the advantages of fast speed and simple optimization iteration.

[0173] This application does not limit the length of word segmentation, and can be used to accurately segment words of different lengths.

[0174] Furthermore, the examples in this application can be flexibly applied to various fields. By obtaining the vocabulary associated with the text to be segmented, sample sentences suitable for practical application fields can be generated, and then a segmentation model suitable for practical application fields can be trained. It has good transferability, versatility and scalability.

[0175] In this application, the modules in the apparatus provided can also implement the method steps provided in the above-described embodiment of the word segmentation model training method. Alternatively, the apparatus provided in this application may include other modules besides those described above to implement the method steps provided in the above-described embodiment of the word segmentation model training method. Furthermore, the apparatus provided in this application can achieve the technical effects achievable by the above-described embodiment of the word segmentation model training method.

[0176] To address the problems existing in the prior art, this application also provides a word segmentation processing device 110, such as... Figure 11 As shown, it includes:

[0177] The second acquisition module 111 acquires the text to be segmented and the trained segmentation model. The trained segmentation model is obtained by iteratively training the segmentation model to be trained based on multiple sample sentences after segmentation annotation. The multiple sample sentences are multiple sample sentences from a preset number of sample sentences. The segmentation annotation is used to indicate the segmentation processing result obtained by segmenting the sample sentences. The preset number of sample sentences are generated based on the segmented words and their corresponding word frequencies in the lexicon associated with the text to be segmented. The lexicon includes multiple segmented words and the word frequencies corresponding to each segmented word. The segmented words include at least some words that constitute the text to be segmented. The word frequency is used to indicate the number of times the corresponding segmented word appears in the text to be segmented.

[0178] The second processing module 112 inputs the text to be segmented into the trained word segmentation model to obtain the word segmentation result of the text to be segmented.

[0179] The device provided in this application embodiment enables the trained word segmentation model to perform word segmentation processing, and can perform accurate and efficient word segmentation on the text to be segmented.

[0180] The obtained word segmentation result can express the segmentation result of the above-mentioned text to be segmented. This word segmentation result can be used to further process the above-mentioned text to be segmented, such as for text proofreading, text parsing, etc.

[0181] In this application, the modules in the apparatus provided can also implement the method steps provided in the above-described word segmentation method embodiment. Alternatively, the apparatus provided in this application may include other modules besides those described above to implement the method steps provided in the above-described word segmentation method embodiment. Furthermore, the apparatus provided in this application can achieve the technical effects achievable by the above-described word segmentation method embodiment.

[0182] Furthermore, corresponding to the above Figures 1 to 9 Based on the same technical concept, this application also provides an electronic device for executing the above-described word segmentation model training method or word segmentation processing method, such as... Figure 12 As shown.

[0183] Electronic devices can vary considerably due to differences in configuration or performance. They may include one or more processors 1201 and memory 1202, with memory 1202 storing one or more application programs or data. Memory 1202 can be temporary or persistent storage. The application programs stored in memory 1202 may include one or more modules (not shown), each module including a series of computer-executable instructions for the electronic device. Furthermore, processor 1201 may be configured to communicate with memory 1202 and execute the series of computer-executable instructions stored in memory 1202 on the electronic device. The electronic device may also include one or more power supplies 1203, one or more wired or wireless network interfaces 1204, one or more input / output interfaces 1205, one or more keyboards 1206, etc.

[0184] In one specific embodiment, the electronic device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for use in the electronic device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following:

[0185] Obtain a lexicon associated with the text to be segmented, the lexicon including multiple word segments and the word frequency corresponding to each word segment, the word segments including at least some words constituting the text to be segmented, and the word frequency being used to indicate the number of times the corresponding word segment appears in the text to be segmented;

[0186] A preset number of sample sentences are generated based on the word segments and their corresponding word frequencies in the lexicon, and the sample sentences are segmented and labeled. The word segmentation labels are used to indicate the word segmentation results obtained by segmenting the sample sentences.

[0187] The word segmentation model to be trained is iteratively trained based on multiple sample sentences after word segmentation annotation to obtain the trained word segmentation model, which is used to perform word segmentation processing on the text to be segmented.

[0188] The electronic device in this embodiment acquires a lexicon associated with the text to be segmented. The lexicon includes multiple word segments and their corresponding word frequencies. Each word segment includes at least a portion of the words constituting the text to be segmented. The word frequency indicates the number of times the corresponding word segment appears in the text to be segmented. A preset number of sample sentences are generated based on the word segments in the lexicon and their corresponding word frequencies. The sample sentences are then labeled with word segments. The word segments are labeled with word segments to indicate the word segmentation results obtained by segmenting the sample sentences. The word segmentation model to be trained is iteratively trained based on the multiple labeled sample sentences to obtain the trained word segmentation model. The trained word segmentation model is used to perform word segmentation processing on the text to be segmented.

[0189] Among them, by generating sample sentences based on the vocabulary associated with the text to be segmented, the trained word segmentation model can be adapted to the text to be segmented, thereby improving the accuracy and speed of text segmentation.

[0190] This application example automatically generates sample sentences for training a word segmentation model based on word segmentation data from a lexicon, and automatically performs word segmentation annotation on the sample sentences without manual annotation. It can efficiently generate sample sentences to train a word segmentation model with high accuracy.

[0191] This application iteratively trains the word segmentation model based on multiple sample sentences after word segmentation and annotation, resulting in a simple and easy-to-understand word segmentation model with advantages such as fast speed and convenient optimization iteration. This application does not limit the length of the segmented words, enabling accurate segmentation of words of varying lengths.

[0192] In addition, this application can be flexibly applied to a variety of fields. By obtaining the vocabulary associated with the text to be segmented, sample sentences suitable for the actual application field can be generated, and then a segmentation model suitable for the actual application field can be trained. It has good transferability, versatility and scalability.

[0193] In another specific embodiment, the electronic device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for use in the electronic device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following:

[0194] The process involves obtaining a text to be segmented and a trained segmentation model. The trained segmentation model is obtained by iteratively training a segmentation model based on multiple sample sentences with segmentation annotations. The multiple sample sentences are multiple sample sentences from a preset number of sample sentences. The segmentation annotations are used to indicate the segmentation results obtained by segmenting the sample sentences. The preset number of sample sentences are generated based on the segmented words and their corresponding word frequencies in a lexicon associated with the text to be segmented. The lexicon includes multiple segmented words and their corresponding word frequencies. The segmented words include at least some of the words that constitute the text to be segmented. The word frequencies are used to indicate the number of times the corresponding segmented word appears in the text to be segmented.

[0195] The text to be segmented is input into the trained word segmentation model to obtain the word segmentation result of the text to be segmented.

[0196] The electronic device in this embodiment performs word segmentation processing through a trained word segmentation model, enabling accurate and efficient word segmentation of the text to be segmented. The resulting word segmentation can express the segmentation result of the text to be segmented. This word segmentation result can be used for further processing of the text to be segmented, such as for text proofreading and text parsing.

[0197] It should be noted that the embodiments of electronic devices in this specification and the embodiments of word segmentation model training methods and word segmentation processing methods in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding methods mentioned above, and the repeated parts will not be described again.

[0198] Furthermore, corresponding to the above Figures 1 to 9 Based on the same technical concept, this application also provides a storage medium for storing computer-executable instructions. In one specific embodiment, the storage medium can be a USB flash drive, optical disc, hard disk, etc. When the computer-executable instructions stored in the storage medium are executed by a processor, they can achieve the following process:

[0199] Obtain a lexicon associated with the text to be segmented, the lexicon including multiple word segments and the word frequency corresponding to each word segment, the word segments including at least some words constituting the text to be segmented, and the word frequency being used to indicate the number of times the corresponding word segment appears in the text to be segmented;

[0200] A preset number of sample sentences are generated based on the word segments and their corresponding word frequencies in the lexicon, and the sample sentences are segmented and labeled. The word segmentation labels are used to indicate the word segmentation results obtained by segmenting the sample sentences.

[0201] The word segmentation model to be trained is iteratively trained based on multiple sample sentences after word segmentation annotation to obtain the trained word segmentation model, which is used to perform word segmentation processing on the text to be segmented.

[0202] When the computer-executable instructions stored in the storage medium in this embodiment are executed by the processor, the following steps are taken: First, a dictionary associated with the text to be segmented is obtained. The dictionary includes multiple word segments and their corresponding word frequencies. Each word segment includes at least a portion of the words constituting the text to be segmented. The word frequency indicates the number of times the corresponding word segment appears in the text. A preset number of sample sentences are generated based on the word segments in the dictionary and their corresponding word frequencies. The sample sentences are then labeled with word segments, indicating the segmentation result obtained by segmenting the sample sentences. The segmentation model to be trained is iteratively trained based on the labeled sample sentences to obtain the trained segmentation model. The trained segmentation model is then used to segment the text to be segmented.

[0203] Among them, by generating sample sentences based on the vocabulary associated with the text to be segmented, the trained word segmentation model can be adapted to the text to be segmented, thereby improving the accuracy and speed of text segmentation.

[0204] This application automatically generates sample sentences for training a word segmentation model based on word segmentation data from a lexicon, and automatically performs word segmentation annotation on the sample sentences. This eliminates the need for manual annotation and can efficiently generate sample sentences to train a word segmentation model with high accuracy.

[0205] This application iteratively trains the word segmentation model to be trained based on multiple sample sentences after word segmentation and annotation, making the trained word segmentation model simple in structure and easy to understand, with the advantages of fast speed and simple optimization iteration.

[0206] This application does not limit the length of word segments, and can be used to accurately segment words of different lengths.

[0207] In addition, this application can be flexibly applied to a variety of fields. By obtaining the vocabulary associated with the text to be segmented, sample sentences suitable for the actual application field can be generated, and then a segmentation model suitable for the actual application field can be trained. It has good transferability, versatility and scalability.

[0208] In another specific embodiment, the computer-executable instructions stored in the storage medium, when executed by a processor, can achieve the following process:

[0209] The process involves obtaining a text to be segmented and a trained segmentation model. The trained segmentation model is obtained by iteratively training a segmentation model based on multiple sample sentences with segmentation annotations. The multiple sample sentences are multiple sample sentences from a preset number of sample sentences. The segmentation annotations are used to indicate the segmentation results obtained by segmenting the sample sentences. The preset number of sample sentences are generated based on the segmented words and their corresponding word frequencies in a lexicon associated with the text to be segmented. The lexicon includes multiple segmented words and their corresponding word frequencies. The segmented words include at least some of the words that constitute the text to be segmented. The word frequencies are used to indicate the number of times the corresponding segmented word appears in the text to be segmented.

[0210] The text to be segmented is input into the trained word segmentation model to obtain the word segmentation result of the text to be segmented.

[0211] The electronic device in this embodiment performs word segmentation processing through a trained word segmentation model, enabling accurate and efficient word segmentation of the text to be segmented. The resulting word segmentation can express the segmentation result of the text to be segmented. This word segmentation result can be used for further processing of the text to be segmented, such as for text proofreading and text parsing.

[0212] It should be noted that the embodiments of electronic devices in this specification and the embodiments of word segmentation model training methods and word segmentation processing methods in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding methods mentioned above, and the repeated parts will not be described again.

[0213] It should be noted that the embodiments concerning storage media in this specification are based on the same inventive concept as the embodiments concerning word segmentation model training methods and word segmentation processing methods in this specification. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding word segmentation model training methods and word segmentation processing methods mentioned above, and the repeated parts will not be described again.

[0214] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0215] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0216] For ease of description, the above devices are described in terms of function, divided into various units. Of course, when implementing one or more of these specifications, the functions of each unit can be implemented in one or more software and / or hardware.

[0217] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0218] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, one or more of this specification can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0219] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0220] The above description is merely an embodiment of one or more embodiments of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.

Claims

1. A word segmentation model training method, characterized in that, include: Obtain a lexicon associated with the text to be segmented, the lexicon including multiple word segments and the word frequency corresponding to each word segment, the word segments including at least some words constituting the text to be segmented, and the word frequency being used to indicate the number of times the corresponding word segment appears in the text to be segmented; A preset number of sample sentences are generated based on the word segments and their corresponding word frequencies in the lexicon, and the sample sentences are then labeled with word segments. The word segmentation labels are used to indicate the word segmentation results obtained by segmenting the sample sentences. The sample sentences include a second target sample sentence, which is obtained by randomly arranging and combining a target number of target word segments selected from the lexicon, and then adding punctuation to the first target sample sentence obtained by randomly arranging and combining the random combinations. The number of punctuation marks is positively correlated with the target number, and the probability of a word segment being selected is positively correlated with the word frequency of the word segment. The target number is the ratio of a random number within the character count set to the average length of the word segments. The character count set is used to represent the number of characters in each sentence to be segmented in the text to be segmented. The average length of the word segments is the ratio of the total number of characters in the word segments in the lexicon to the total number of word segments in the lexicon. The word segmentation model to be trained is iteratively trained based on multiple sample sentences after word segmentation annotation to obtain the trained word segmentation model, which is used to perform word segmentation processing on the text to be segmented.

2. The method as described in claim 1, characterized in that, The text to be segmented includes multiple sentences to be segmented, and the acquisition of the dictionary associated with the text to be segmented includes: Obtain a dictionary containing common words in each of the multiple sentences to be segmented; The frequency of each word is determined based on the number of times each word in the lexicon appears in the text to be segmented.

3. The method as described in claim 2, characterized in that, The step of generating a preset number of sample sentences based on word segmentation and their corresponding word frequencies in the lexicon includes: Determine the number M of sentences to be segmented contained in the text to be segmented; The number of sample sentences to be generated, N, is determined based on the number M of sentences to be segmented, where N and M are positive integers and the ratio of N to M is greater than a preset ratio. N sample sentences are generated based on the word segments and their corresponding word frequencies in the lexicon, and the N sample sentences are then segmented and labeled.

4. The method as described in claim 2, characterized in that, The step of generating a preset number of sample sentences based on word segments and their corresponding word frequencies in the lexicon, and then performing word segmentation and annotation on the sample sentences, includes: Determine the set of character counts for the text to be segmented, wherein the character count of each sentence to be segmented in the text belongs to the set of character counts; The ratio of the total number of characters in the word segmentation in the dictionary to the total number of word segments in the dictionary is determined as the average length of the word segmentation; A preset number of sample sentences are generated based on the character count set, the average length of the word segments, the word segments in the lexicon and their corresponding word frequencies, and the sample sentences are segmented and labeled. The number of word segments in the generated sample sentences is the ratio of a random number in the character count set to the average length of the word segments.

5. The method as described in claim 4, characterized in that, After determining the set of character counts for the text to be segmented, the method further includes: The set of word counts is divided into multiple word count subsets, and each word count subset contains the word count of at least one sentence to be segmented; The weights corresponding to the plurality of word count subsets are determined respectively, wherein the weight corresponding to any word count subset represents the probability that the word count of the sentence to be segmented in the text to be segmented belongs to the corresponding word count subset; The step of generating a preset number of sample sentences based on the character count set, the average length of the word segments, the word segments in the lexicon and their corresponding word frequencies, and then performing word segmentation and annotation on the sample sentences includes: The target word count subset is determined based on the weight corresponding to each word count subset in the word count set; Random numbers are randomly selected from the target word count subset, and the target number of words contained in the target sample sentence is determined based on the random numbers and the average word segmentation length. Construct a target sample sentence containing a target number of segmented words based on the lexicon, and perform word segmentation and annotation on the target sample sentence.

6. The method as described in claim 5, characterized in that, The step of constructing a target sample sentence containing a target number of segmented words based on the lexicon, and performing word segmentation and annotation on the target sample sentence, includes: Select a target number of target words from the vocabulary based on word frequency; The target words are randomly permuted and combined to generate a pre-constructed first target sample sentence; Punctuation marks are added to the pre-constructed first target sample sentence according to the target number to generate a second target sample sentence, and word segmentation and tagging are performed on the second target sample sentence.

7. The method as described in claim 1, characterized in that, The word segmentation model to be trained is iteratively trained using multiple sample sentences after word segmentation annotation to obtain the trained word segmentation model, including: For each sample sentence, the sample sentence is input into a segmentation model that is trained in stages to obtain the predicted segmentation result of the sample sentence; Based on the predicted word segmentation results and word segmentation annotations of the sample sentences, the accuracy of the predicted word segmentation results of the sample sentences is determined; When the accuracy of the predicted word segmentation result is lower than the preset accuracy, the generation parameters of the sample sentences are adjusted to regenerate the sample sentences and iteratively train the word segmentation model trained in stages to obtain the trained word segmentation model. The generation parameters include at least one of the word frequency corresponding to the segmentation and the number of sample sentences.

8. A word segmentation processing method, characterized in that, include: The process involves obtaining a text to be segmented and a trained segmentation model. The trained segmentation model is obtained by iteratively training a segmentation model based on multiple sample sentences with segmentation annotations. The multiple sample sentences are multiple sample sentences from a preset number of sample sentences. The segmentation annotations are used to indicate the segmentation results obtained by segmenting the sample sentences. The preset number of sample sentences are generated based on the segmented words and their corresponding word frequencies in a lexicon associated with the text to be segmented. The lexicon includes multiple segmented words and their corresponding word frequencies. The segmented words include at least some of the words that constitute the text to be segmented. The word frequencies are used to indicate the number of times the corresponding segmented word appears in the text to be segmented. The text to be segmented is input into the trained word segmentation model to obtain the word segmentation result of the text to be segmented; The sample sentences include a second target sample sentence, which is obtained by randomly arranging and combining a target number of target words selected from the lexicon, and then adding punctuation to the first target sample sentence obtained by the random arrangement and combination. The number of punctuation marks is positively correlated with the target number, and the probability of a word being selected is positively correlated with the word frequency of the word. The target number is the ratio of a random number in the character set to the average length of the word segment. The character set is used to represent the number of characters in each sentence to be segmented in the text to be segmented. The average length of the word segment is the ratio of the total number of characters in the word segment in the lexicon to the total number of words in the lexicon.

9. A word segmentation model training device, characterized in that, include: The first acquisition module acquires a lexicon associated with the text to be segmented. The lexicon includes multiple word segments and the word frequency corresponding to each word segment. The word segments include at least a portion of the words that constitute the text to be segmented. The word frequency is used to indicate the number of times the corresponding word segment appears in the text to be segmented. The first generation module generates a preset number of sample sentences based on the word segments and their corresponding word frequencies in the lexicon, and performs word segmentation annotation on the sample sentences. The word segmentation annotation is used to indicate the word segmentation processing result obtained by segmenting the sample sentences. The sample sentences include a second target sample sentence. The second target sample sentence is obtained by randomly arranging and combining a target number of target word segments selected from the lexicon, and then adding punctuation to the first target sample sentence obtained by randomly arranging and combining them. The number of punctuation marks is positively correlated with the target number, and the probability of word segmentation being selected is positively correlated with the word frequency of the word segmentation. The target number is the ratio of a random number in the character count set to the average length of the word segmentation. The character count set is used to represent the number of characters in each sentence to be segmented in the text to be segmented. The average length of the word segmentation is the ratio of the total number of characters in the word segmentation in the lexicon to the total number of word segments in the lexicon. The first training module iteratively trains the word segmentation model to be trained based on multiple sample sentences after word segmentation and annotation, and obtains the trained word segmentation model, which is used to perform word segmentation processing on the text to be segmented.

10. A word segmentation processing device, characterized in that, include: The second acquisition module acquires the text to be segmented and the trained segmentation model. The trained segmentation model is obtained by iteratively training the segmentation model to be trained based on multiple sample sentences after segmentation annotation. The multiple sample sentences are multiple sample sentences from a preset number of sample sentences. The segmentation annotation is used to indicate the segmentation processing result obtained by segmenting the sample sentences. The preset number of sample sentences are generated based on the segmented words and their corresponding word frequencies in the lexicon associated with the text to be segmented. The lexicon includes multiple segmented words and the word frequencies corresponding to each segmented word. The segmented words include at least some of the words that constitute the text to be segmented. The word frequencies are used to indicate the number of times the corresponding segmented word appears in the text to be segmented. The second processing module inputs the text to be segmented into the trained word segmentation model to obtain the word segmentation result of the text to be segmented. The sample sentences include a second target sample sentence, which is obtained by randomly arranging and combining a target number of target words selected from the lexicon, and then adding punctuation to the first target sample sentence obtained by the random arrangement and combination. The number of punctuation marks is positively correlated with the target number, and the probability of a word being selected is positively correlated with the word frequency of the word. The target number is the ratio of a random number in the character set to the average length of the word segment. The character set is used to represent the number of characters in each sentence to be segmented in the text to be segmented. The average length of the word segment is the ratio of the total number of characters in the word segment in the lexicon to the total number of words in the lexicon.

11. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method as described in any one of claims 1-7 or claim 8.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1-7 or claim 8.

Citation Information

Patent Citations

  • Word segmentation method, device, apparatus and storage medium

    CN109271631A

  • Input method model generation method and device

    CN109710087A