Method, apparatus, electronic device and readable storage medium for determining new words

By performing word sequence mining and supersequence judgment on the sample text set, the problem of existing word segmentation tools being mistakenly split when processing new words is solved, more accurate word segmentation and new word discovery is achieved, and training costs are reduced.

CN111680146BActive Publication Date: 2025-06-20TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010525541.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-10
Publication Date
2025-06-20
Estimated Expiration
2040-06-10

AI Technical Summary

Technical Problem

Existing word segmentation tools are prone to miss splitting when dealing with new online words, personal names, place names, professional terms, etc., resulting in inaccurate word segmentation results, mainly due to errors in the recognition of new words.

Method used

By obtaining the sample text set, word sequence mining is performed to obtain frequent word sequences, determine the supersequence, and determine whether the supersequence is included in the participle of the sample text set, and if not, it is determined as a new word.

Benefits of technology

This method can better identify frequently updated words, words or phrases, improve the accuracy of word segmentation and new word discovery, reduce the need to train complex neural network models and manually label training samples, and reduce training costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111680146B_ABST
    Figure CN111680146B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method, apparatus, electronic device, and readable storage medium for determining new words. The method includes: obtaining a sample text set; performing character sequence mining on the sample text set to obtain frequent character sequences corresponding to each length; determining each supersequence in the frequent character sequences corresponding to each length; for each supersequence, if the supersequence is not included in each word segmentation included in the sample text set, then determining the supersequence as a new word. In the embodiment of the present application, the method of character sequence mining can better screen out frequently updated characters, words, or phrases, which will have important reference value and practical significance in applications such as word segmentation and new word discovery; and in the process of determining new words, there is no need to train a complex neural network model, nor is it necessary to manually label training samples, thereby effectively reducing the training cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing. Specifically, this application relates to a method, apparatus, electronic device, and readable storage medium for determining new words. Background Art

[0002] With the development of language and the continuous emergence and evolution of Internet terms, new words and new professional terms emerge in an endless stream. For many tasks of natural language processing, the quality of word segmentation plays a crucial role in the accuracy of subsequent task flows and is also an important influencing factor for the task effects of other basic tools (such as syntactic analysis, keyword extraction, etc.). In fact, in current word segmentation tools, there are often incorrect segmentations for network new words, personal names, place names, professional terms, etc., that is, the results of word segmentation are inaccurate. The root cause is also the problem caused by incorrect recognition of new words. Therefore, how to mine new words is an important problem that needs to be solved. Summary of the Invention

[0003] The purpose of this application aims to solve at least one of the above technical defects.

[0004] On the one hand, an embodiment of this application provides a method for determining new words. The method includes:

[0005] Obtain a sample text set;

[0006] Perform character sequence mining on the sample text set to obtain frequent character sequences corresponding to each length;

[0007] Determine each super-sequence in the frequent character sequences corresponding to each length;

[0008] For each super-sequence, if the super-sequence is not included in each word segmentation included in the sample text set, then determine the super-sequence as a new word.

[0009] On the other hand, an embodiment of this application provides a text processing method. The method includes:

[0010] Obtain the text to be processed;

[0011] Perform word segmentation processing on the text to be processed based on the word segmentation database to obtain the word segmentations included in the text to be processed. The word segmentation database contains new words determined by the method in the first aspect.

[0012] On the other hand, an embodiment of this application provides a device for determining new words. The device includes:

[0013] A text acquisition module, configured to obtain a sample text set;

[0014] A sequence mining module, configured to perform character sequence mining on the sample text set to obtain frequent character sequences corresponding to each length;

[0015] A super-sequence determination module, configured to determine each super-sequence in the frequent word sequences corresponding to each length;

[0016] A new word determination module, for each super-sequence, if the super-sequence is not included in each word segmentation in the sample text set, then determine the super-sequence as a new word.

[0017] On the other hand, an embodiment of the present application provides a text processing device, including:

[0018] A text acquisition module, configured to acquire the text to be processed;

[0019] A word segmentation processing module, configured to perform word segmentation processing on the text to be processed based on a word segmentation database, to obtain the word segments included in the text to be processed, and the word segmentation database contains new words determined by the method in the first aspect.

[0020] In yet another aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory:

[0021] The memory is configured to store a computer program, and when the computer program is executed by the processor, the processor is caused to execute the method provided in any aspect of the present application.

[0022] In still another aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program runs on a computer, the computer can execute the method provided in any aspect of the present application.

[0023] The beneficial effects brought by the technical solution provided by the embodiment of the present application are:

[0024] In the embodiment of the present application, since the sample text set can be mined for word sequences to obtain frequent word sequences corresponding to each length, and then the super-sequences in the frequent word sequences are screened based on the word segments included in the sample text set, and the screened super-sequences are used as new words. At this time, the method of word sequence mining can better screen out frequently updated words, phrases or expressions, which will have important reference value and practical significance in applications such as word segmentation and new word discovery; and in the process of determining new words, there is no need to train a complex neural network model, nor to manually label training samples, thereby effectively reducing the training cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present application.

[0026] Figure 1 It is a schematic flowchart of a method for determining new words provided by an embodiment of the present application;

[0027] Figure 2 A flowchart showing the process of a text processing method provided by an embodiment of the present application;

[0028] Figure 3 A structural diagram of a device for determining new words provided by an embodiment of the present application;

[0029] Figure 4 A structural diagram of a text processing device provided by an embodiment of the present application;

[0030] Figure 5 A structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0031] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present application.

[0032] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0033] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that combines linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language used by people in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies.

[0034] The emergence of new words and the incorrect splitting of word segmentation are both common phenomena in natural language understanding and difficult problems that existing language processing technologies must overcome. Currently, the main methods for discovering and identifying new words include the new word discovery method based on language models, the new word discovery algorithm based on segmentation, or the new word discovery method based on deep learning. The following briefly introduces these methods.

[0035] 1. New word discovery method based on language models: This method is a method that uses conditional probability for calculation. It converts the probability of words into the combined probability of characters through Bayes' formula and determines new words based on the combined probability of characters.

[0036] 2. New word discovery algorithm based on segmentation: The basis for segmentation of this algorithm lies in the cohesion between elements. By calculating the cohesion between character elements, it selects when to perform word segmentation to identify new words.

[0037] 3. New word discovery method based on deep learning: This method first performs word segmentation, then calculates the amount of information and adds rules for recognition, and discovers new words through a deep learning model.

[0038] However, the existing methods above have the following problems that need to be improved:

[0039] 1. When using the new word discovery method based on language models, its effect depends on the quality of the language model. The premise of probability conversion through Bayes' formula is the independence assumption of features. However, in fact, there is a certain degree of correlation in the appearance of characters, and they are not completely independent.

[0040] 2. In order to avoid cutting out too many invalid words in the new word discovery method based on segmentation, a relatively large cohesion needs to be selected. However, in fact, the cohesion of two characters forming a word does not necessarily have to be very large, that is, there is a problem of not strict enough segmentation criteria.

[0041] 3. The new word discovery method based on deep learning models needs to train neural networks based on a large number of labeled training samples. For industrial applications, there is a problem that the efficiency is difficult to meet the online requirements.

[0042] Based on this, the present application provides a method, device, electronic device, and computer-readable storage medium for determining new words, aiming to solve at least one technical problem in the existing technology.

[0043] The following uses specific embodiments to elaborate in detail on the technical solution of the present application and how the technical solution of the present application solves the above-mentioned at least one technical problem. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of the present application in conjunction with the accompanying drawings.

[0044] The method for determining new words provided by the embodiments of the present application can be used in any electronic device, such as products like smart phones and tablet computers. Of course, this method can also be applied to a server (including but not limited to physical servers and cloud servers), and the server can determine new words from a sample text set based on the method provided by the embodiments of the present application.

[0045] Figure 1 FIG. 4 shows a schematic flowchart of a method for determining new words provided by the embodiments of the present application. As Figure 1 shown, the method includes:

[0046] Step S101, obtain a sample text set.

[0047] Among them, the sample text set includes multiple sample texts, and new words may be included in the multiple sample texts. For example, some sample texts contain newly generated Internet terms, and some sample texts contain newly generated professional terms in a certain field, etc. Among them, in an optional embodiment of the present application, the specific form of the sample text can be a single sentence or a single word, and whether the specific form of the sample text is a single sentence or a single word can be pre-configured according to actual applications.

[0048] In an optional embodiment of the present application, obtaining the sample text set may include:

[0049] Obtain an initial text set, where the initial text set includes each initial text;

[0050] Perform text preprocessing on each initial text respectively to obtain a preprocessing result corresponding to each initial text;

[0051] Based on the preprocessing results corresponding to each initial text, obtain the sample text set;

[0052] Among them, the text preprocessing includes at least one of clause splitting processing and specific character deletion processing.

[0053] Specifically, the clause splitting processing refers to splitting an article or text fragment with multiple clauses into multiple independent sentences; the specific character deletion processing refers to deleting specific characters in the text, and the specific type of the specific characters can be pre-configured, such as being configured according to the actual application scenario and / or experience, and the embodiments of the present application do not limit this. For example, punctuation marks can be set as specific characters. At this time, when there is a comma in the initial text, the comma can be deleted from the initial text.

[0054] In practical applications, the initial text set includes various initial texts. The specific form of the initial text is not limited in the embodiments of the present application. For example, the initial text can be a passage with multiple clauses, or it can be a single sentence. That is to say, the granularity of the initial text is not limited in the embodiments of the present application and can be configured according to actual application needs. As an optional method, since the sample text can be a single sentence or a single sentence, and when the initial text is an article or a text segment, at this time, the article or text segment can be clause-separated, and each clause obtained after the processing (i.e., the preprocessing result) is used as a sample text.

[0055] In an example, assume that an initial text in the initial text set is "I love natural language processing. Natural language processing is the core tool for text analysis and mining. Fine-grained sentiment analysis is a type of sentiment analysis. This article conducts research on the key technologies in fine-grained sentiment analysis". The text preprocessing includes clause-separating and specific character deletion processing, and the specific character is punctuation. At this time, the initial text can be clause-separated and the specific character deleted, and the processed initial text obtained includes four independent clauses: "I love natural language processing", "Natural language processing is the core tool for text analysis and mining", "Fine-grained sentiment analysis is a type of sentiment analysis", and "This article conducts research on the key technologies in fine-grained sentiment analysis", and these four independent clauses are used as 4 sample texts in the sample text set, as shown in Table 1 specifically.

[0056] Table 1

[0057] I love natural language processing Natural language processing is a core tool for text analysis and mining Fine-grained sentiment analysis is a type of sentiment analysis This paper conducts research on the key technologies in fine-grained sentiment analysis

[0058] It can be understood that in the embodiments of the present application, when the text preprocessing includes clause-separating and specific character deletion processing, and the specific character deletion processing is punctuation deletion processing, the initial text can be clause-separated first and then the punctuation can be deleted.

[0059] Step S102, perform word sequence mining on the sample text set to obtain frequent word sequences corresponding to each length.

[0060] Among them, character elements refer to the characters included in the frequent character sequence, and characters of different forms are different character elements, such as the character "字" and the character "与" in the frequent character sequence are different character elements; frequent character sequences refer to character sequences that frequently appear in the sample text set, and the length refers to the number of character elements included in the frequent character sequence. For example, when a frequent character sequence is "natural language", the frequent character sequence includes four character elements, namely, the character element "自", the character element "然", the character element "语" and the character element "言", and the length of the frequent character sequence "自然语言" is 4. Optionally, each length in the embodiment of the present application may refer to the length of a character element to the length of the character elements included in the longest frequent character sequence, or may refer to the length from a set starting length (such as the length of 1 character element or the length of two character elements) to the length of the character elements included in the longest frequent character sequence, which is not limited in the embodiment of the present application.

[0061] In practical applications, after obtaining a sample text set, word sequence mining can be performed on each sample text included in the sample text set to obtain frequent word sequences corresponding to various lengths.

[0062] Step S103, determining each supersequence in the frequent word sequences corresponding to each length.

[0063] In practical applications, if all elements of a frequent word sequence A can be found in the item set of a frequent word sequence B, then the frequent word sequence A is a subsequence of the frequent word sequence B. According to this definition, suppose that for a frequent word sequence A = {a1, a2, ...a n} and frequent word sequence B = {b1, b2, ... b m},n≤m, if there exists a digital sequence 1≤j1≤j2≤...≤j m ≤m,satisfying Then the frequent word sequence A is called a subsequence of the frequent word sequence B, and conversely, the frequent word sequence B is a supersequence of the frequent word sequence A.

[0064] In practical applications, since the lengths of the frequent word sequences mined are different, there may be situations where the elements in some frequent word sequences are included in other frequent word sequences. At this time, it is possible to determine which frequent word sequences in the frequent word sequences belong to supersequences and which frequent word sequences are subsequences. Because the supersequence itself contains more reference information, in order to ensure the integrity of the information and reduce the amount of subsequent data processing, the subsequence can be deleted and the supersequence can be retained.

[0065] Among them, the implementation manner of determining each supersequence in the frequent word sequence can be pre-configured, and the embodiments of the present application do not limit this. For example, the supersequence in the frequent word sequence can be determined by setting an n-gram window, and the number of word elements included in the n-gram window (i.e., n) can be determined according to the length of each frequent word sequence.

[0066] In one example, assume that the length of the frequent word sequence with the largest length among frequent word sequences of each length is 8. At this time, n can be set to 8, and any word element in each frequent word sequence can be selected as the center (i.e., w). For example, the word element "language" is selected as w, and then from each frequent word sequence, find each frequent word sequence that includes the word element "language" and the number of word elements before and after the word element "language" (which can be called the context auxiliary information Context(w) of w) is less than 8. At this time, each found frequent word sequence includes the word element "language", and then determine which of the found frequent word sequences are supersequences and which are subsequences. For example, it is determined that both the frequent word sequence "natural language processing" and the frequent word sequence "natural language" include the word element "language". Since "natural language processing" also includes the context auxiliary information "processing" on the basis of the subsequence "natural language", the supersequence pattern "natural language processing" can be retained and the subsequence "natural language" can be deleted.

[0067] Step S104, for each supersequence, if the supersequence is not included in each word segmentation included in the sample text set, then determine the supersequence as a new word.

[0068] In practical applications, each sample text included in the sample text set can be segmented to obtain the word segmentations included in the sample text set. Among them, the specific implementation manner of segmenting each sample text can be pre-configured, and the embodiments of the present application do not limit this. For example, existing word segmentation tools (such as jieba and other word segmentation tools) can be used to segment the sample text; then for each supersequence, determine whether the supersequence is included in the word segmentations included in the sample text set. If the supersequence is not included, at this time, the supersequence can be determined as a new word. On the contrary, if the supersequence is included in the word segmentations included in the sample text set, it means that the supersequence is an existing word. In addition, in practical applications, it is also possible to determine whether each supersequence is a new word based on the word segmentation result obtained by segmenting the initial text in the initial text set. The embodiments of the present application do not limit this.

[0069] Continuing with the previous example, assuming that based on the word segmentation results of the initial texts in the initial text set, it is determined whether each supersequence is a new word. At this time, the word segmentation results of the initial texts in the initial text set are specifically shown in Table 2. Then, for each supersequence, it is determined whether the word segmentation shown in Table 2 includes this supersequence. If it does not, this supersequence is determined as a new word. The new word results obtained at this time are shown in Table 3. For example, for the supersequence "natural language processing", it is not reflected in Table 2. At this time, the supersequence "natural language processing" is a new word.

[0070] Table 2

[0071] Word segmentation result I love natural language processing Natural language processing is a core tool for text analysis and mining Fine-grained sentiment analysis is a type of sentiment analysis This paper conducts research on the key technologies in fine-grained sentiment analysis

[0072] Table 3

[0073] New word result Fine-grained sentiment analysis Natural language processing

[0074] In the embodiments of the present application, since the character sequence mining can be performed on the sample text set to obtain frequent character sequences corresponding to each length, and then the supersequences in the frequent character sequences are screened based on the word segmentation included in the sample text set, and the screened supersequences are used as new words. At this time, it is possible to better screen out the characters, words or phrases that are frequently updated. Therefore, it has important reference value and practical significance in applications such as word segmentation and new word discovery; and in the process of determining new words, neither a complex neural network model needs to be trained nor manual annotation of training samples is required, thereby effectively reducing the training cost.

[0075] In an alternative embodiment of the present application, performing character sequence mining on the sample text set to obtain frequent character sequences corresponding to each length includes:

[0076] Determining the number of samples corresponding to each character element included in the sample text set. For a character element, the number of samples corresponding to this character element refers to the number of sample texts in the sample text set that contain this character sequence element;

[0077] Filtering the character elements included in the sample text set based on the number of samples corresponding to each character element to obtain a processed sample text set;

[0078] Performing character sequence mining on the processed sample text set to obtain frequent character sequences corresponding to each length.

[0079] In practical applications, the number of samples corresponding to each character element included in the sample text set can be counted. It should be noted that for a character element, when there are multiple such character elements in the same sample, this sample is still counted as one sample, that is, the number of samples is incremented by 1.

[0080] In one example, assume that the sample text set includes sample texts "I love natural language processing", "Natural language processing is the core tool for text analysis and mining", "Fine-grained sentiment analysis is a type of sentiment analysis", and "This paper conducts research on the key technologies in fine-grained sentiment analysis", which are four independent clauses. At this time, the number of samples corresponding to each character element can be counted. For example, the character element "fen" appears in "Natural language processing is the core tool for text analysis and mining", "Fine-grained sentiment analysis is a type of sentiment analysis", and "This paper conducts research on the key technologies in fine-grained sentiment analysis". At this time, the number of samples corresponding to the character element "fen" is 3, while the character element "wo" only appears in "I love natural language processing", so the number of samples corresponding to the character element "wo" is 1. Similarly, the number of samples corresponding to the character element "xi" is 3, the number of samples corresponding to the character element "de" is 3, the number of samples corresponding to the character element "gan" is 3, and the number of samples corresponding to other character elements can be obtained in the same way, which will not be elaborated here.

[0081] Furthermore, based on the number of samples corresponding to each character element, the character elements included in the sample text set can be filtered to obtain the processed sample text set, and then word sequence mining can be performed on the processed sample text set to obtain frequent word sequences corresponding to each length.

[0082] In an alternative embodiment of the present application, filtering the character elements included in the sample text set based on the number of samples corresponding to each character element to obtain the processed sample text set includes:

[0083] For the number of samples of a character element, if the number of samples meets the set conditions, the character element is deleted from the sample text set;

[0084] The number of samples meeting the set conditions includes at least one of the following:

[0085] The number of samples is less than the set value or the proportion of the number of samples is less than the preset value;

[0086] Among them, the proportion of the number of samples refers to the ratio of the number of samples corresponding to the character element to the number of sample texts included in the sample text set.

[0087] In practical applications, for a character element, it can be determined whether the number of samples corresponding to the character element meets the set conditions. If the preset conditions are met and when the character element exists in the sample text, the character element can be deleted from the sample text at this time. Among them, the number of samples meeting the set conditions can include at least one of the number of samples being less than the set value and the proportion of the number of samples being less than the preset value. The proportion of the number of samples refers to the ratio of the number of samples corresponding to the character element to the number of sample texts included in the sample text set. For example, if the number of samples corresponding to a certain character element is 4 and there are 4 sample texts in the sample text set, the proportion of the number of samples of this character element is 4 / 4 = 1 at this time.

[0088] Continuing the previous example, assume that the condition for the number of samples to meet the set condition is that the number of samples is less than the set value, and the set value is set to 2; further, since the number of samples corresponding to the character element "fen" is 3, which is greater than the preset threshold of 2, the character element "fen" can be retained at this time, while the number of samples corresponding to the character element "wo" is 1, which is less than the preset threshold of 2, so the character element "wo" can be deleted. Similarly, it can be determined whether other character elements such as the character element "xi", the character element "de" with the corresponding number of samples of 3, and the character element "gan" need to be deleted. The results of the filtered character elements are shown in Table 4 at this time.

[0089] Table 4

[0090] Character element Number of samples Minute 3 Analyze 3 Of 3 Feeling 2 Emotion 2 This 2 Place 2 Degree 2 Handle 2 Grain 2 However 2 Is 2 Article 2 Fine 2 Speech 2 Language 2 Nature 2

[0091] Furthermore, each sample text in this example can be retained only with the character elements included in the "Character Element" column in Table 4 to obtain the processed sample text set, which can be specifically shown in Table 5. For example, since the element "wo" meets the set conditions, the character element "wo" can be deleted from the sample text "I love natural language processing" at this time to obtain the processed sample (i.e., "love natural language processing").

[0092] Table 5

[0093] Processed sample text Natural language processing Natural language processing is for text analysis Fine-grained sentiment analysis is for sentiment analysis This paper's fine-grained sentiment analysis

[0094] It can be understood that when the condition for the number of samples to meet the set condition is that the proportion of the number of samples is less than the preset value, the proportion of the number of samples corresponding to each character element can be determined, and then the character elements with a proportion less than the preset value can be deleted from the sample text set. For example, in the previous example, the number of samples corresponding to the character element "wo" is 1, and its ratio to the total number of samples is 1 / 4, which is less than the preset value of 1 / 3, so the character element "wo" can be deleted from each sample text, while the number of samples corresponding to the character element "gan" is 3, and its ratio to the total number of samples is 3 / 4, which is not less than the preset value of 1 / 3, so the character element "gan" can be retained at this time.

[0095] In an alternative embodiment of the present application, word sequence mining is performed on the sample text set to obtain frequent word sequences corresponding to each length, including:

[0096] Based on the PrefixSpan (Prefix-Projected Pattern Growth) algorithm, word sequence mining is performed on the sample text set to obtain frequent word sequences corresponding to each length.

[0097] In practical applications, a minimum support threshold can be preset, and then the PrefixSpan algorithm is used to perform word sequence mining on the sample text set to obtain frequent word sequences corresponding to each length. The calculation method of the minimum support is as follows.

[0098] min_sup = a × n

[0099] Where n is the number of samples, a is the minimum support ratio, the minimum support ratio can be adjusted according to the magnitude of the sample text set, and min_sup is the minimum support.

[0100] The specific operation steps for performing word sequence mining based on the PrefixSpan algorithm are as follows:

[0101] 1. Find the word sequence prefixes with a unit length of 1 and the corresponding projected data sets;

[0102] 2. Count the occurrence frequencies of the word sequence prefixes and add the prefixes with a support higher than the minimum support threshold to the data set to obtain the frequent word sequences of the one-item set;

[0103] 3. Recursively mine all prefixes with a length of i that meet the minimum support threshold requirements:

[0104] 4. Mine the projected data sets of the prefixes. If the projected data is an empty set, return the recursion;

[0105] 5. Count the minimum support of each item in the corresponding projected data set, merge each single item that meets the support with the current prefix to obtain a new prefix, and recursively return if the support requirements are not met;

[0106] 6. Let i = i + 1, and the prefixes are the new prefixes after merging the single items. Recursively execute step 3 respectively until the projected data sets of the prefixes are all less than the minimum support;

[0107] 7. Return all the frequent word sequences in the word sequence data set.

[0108] In one example, assuming that the sample text set includes 4 sample texts as shown in Table 5, when the PrefixSpan algorithm is used to mine the word sequence of the sample text set to obtain the frequent word sequences corresponding to each length, the prefix of length 1 (i.e., a prefix) can be mined first, and then each prefix that meets the minimum support threshold and its corresponding adjacent suffix (i.e., the word element included in the subsequent part adjacent to the prefix in the sample text) can be determined. For example, for a prefix "分", its adjacent suffix in the sample texts "自然语处理是文字分析的", "细粒度情感分析的" and "本文精粒情感分析的" are all "析的", and for a prefix "的", its adjacent suffix that does not exist in each sample text (all indicated by "none" in the table), the corresponding adjacent suffixes of other prefixes can be obtained in the same way, as shown in Table 6:

[0109] Table 6

[0110]

[0111] Furthermore, prefixes of length 2 (i.e., binomial prefixes) are mined. At this time, each binomial prefix that meets the minimum support threshold and its corresponding adjacent suffix (i.e., the character elements included subsequently) can be determined. For example, for a prefix "分", it can be determined whether the ratio of the number of samples corresponding to each character element in each of its corresponding adjacent suffixes to the total number of samples (the total number of adjacent suffixes of a prefix) is greater than the minimum support threshold of 1 / 3. For example, for the character element "分析", the ratio of its corresponding number of samples (3) to the total number of samples (3) is 1, and is greater than the minimum support threshold of 1 / 3. At this time, the character element and the prefix can be merged into a binomial prefix "分析", and the adjacent suffix corresponding to the binomial prefix "分析" can be determined. Similarly, other binomial prefixes and corresponding adjacent suffixes can be obtained, as shown in Table 7:

[0112] Table 7

[0113]

[0114] Further, prefixes of length 3 (i.e., three-item prefixes) are mined. At this time, each three-item prefix that meets the minimum support threshold and its corresponding adjacent suffixes can be determined. For example, for the one-item prefix "sentiment score", it can be determined whether the ratio of the number of samples corresponding to each character element in its corresponding adjacent suffixes to the total number of samples is greater than the minimum support threshold of 1 / 3. For example, for the character element "analysis", the ratio of its corresponding number of samples (2) to the total number of samples (2) is 1, and it is greater than the minimum support threshold of 1 / 3. At this time, this character element can be combined with the two-item prefix to form a three-item prefix "sentiment analysis", and the adjacent suffixes corresponding to the three-item prefix "sentiment analysis" can be determined. Similarly, other three-item prefixes and their corresponding adjacent suffixes can be obtained, as shown in Table 8 specifically:

[0115] Table 8

[0116]

[0117] Further, based on the same principle, four-item prefixes and their corresponding adjacent suffixes are mined. For example, for the three-item prefix "language processing", it can be determined whether the ratio of the number of samples corresponding to each character element in its corresponding adjacent suffixes to the total number of samples is greater than the minimum support threshold of 1 / 3. For example, for the character element "management", the ratio of its corresponding number of samples (2) to the total number of samples (2) is 1, and it is greater than the minimum support threshold of 1 / 3. At this time, this character element can be combined with the three-item prefix to form a four-item prefix "language processing", and the adjacent suffixes corresponding to the four-item prefix "language processing" can be determined. Similarly, other four-item prefixes and their corresponding adjacent suffixes can be obtained, and the results are shown in Table 9:

[0118] Table 9

[0119]

[0120] Further, based on the same principle, five-item prefixes and their corresponding adjacent suffixes are mined, and the results are shown in Table 10:

[0121] Table 10

[0122]

[0123] Further, based on the same principle, six-item prefixes and their corresponding adjacent suffixes are mined, and the results are shown in Table 11::

[0124] Table 11

[0125]

[0126] Further, based on the same principle, seven-item prefixes and their corresponding adjacent suffixes are mined, and the results are shown in Table 12:

[0127] Table 12

[0128]

[0129] Further, since the word elements in the adjacent suffixes corresponding to the seven-item prefixes are all less than the minimum support threshold, there will be no eight-item prefixes at this time. Thus, the mining iteration ends, and the above-mentioned prefixes are used as frequent word sequences of each length.

[0130] In an alternative embodiment of the present application, determining each supersequence in the frequent word sequences corresponding to each length includes:

[0131] Performing particle filtering on each frequent word sequence respectively to obtain each filtered frequent word sequence;

[0132] Determining each supersequence in each filtered frequent word sequence.

[0133] In practical applications, after obtaining each frequent word sequence, some frequent word sequences may include some particles. To reduce the data processing volume, at this time, each frequent word sequence can be particle-filtered to obtain each filtered frequent word sequence, and then each supersequence in each filtered frequent word sequence is determined. Among them, common particles include "de", "le", "shi", "deng", "zai", "zhe", etc. The method of particle filtering for frequent word sequences can be carried out by constructing a particle dictionary or based on syntactic analysis, etc.

[0134] Figure 2 The flowchart of a text processing method provided in an embodiment of the present application is shown. As Figure 2 shown, the method includes:

[0135] Step S201, obtaining the text to be processed;

[0136] Step S202, performing word segmentation processing on the text to be processed based on the word segmentation database to obtain the word segments included in the text to be processed; wherein, the word segmentation database contains the new words determined by the method described above

[0137] Wherein, the text to be processed refers to the text that needs to be word-segmented, which can be a single sentence, an article with multiple clauses, or a text fragment. The embodiments of the present application do not limit this.

[0138] In practical applications, after each new word is recognized, the recognized new words can be added to the word library of the existing word segmentation tool. For example, a new word library can be established separately in the existing word segmentation tool, or the new words can be directly added to the original word library of the existing word segmentation tool. Further, when the text to be processed is obtained, when the word segmentation tool performs word segmentation on the text to be processed, if the word segmentation tool includes a new word library and an original word library, at this time, the text to be processed can be segmented based on the new word library first to obtain the new words included in the text to be processed, and then the text to be processed after removing the new words can be segmented based on the original word library to obtain the final word segmentation result; correspondingly, if the original word library in the word segmentation tool includes new words, at this time, it can be first determined whether the text to be processed includes the word with the largest length (referring to the number of characters) in the original word library. If it does, the word in the text to be processed is used as a word segment. Further, it is determined whether the text to be processed after removing the word segment includes the word with the second largest length in the original word library. If so, the word with the second largest length in the text to be processed is used as a word segment, and so on, until all the word segments included in the text to be processed are obtained.

[0139] It can be understood that in this example, only the application scenario of word segmentation is used to illustrate the method provided by the embodiments of the present application. However, the method provided by the embodiments of the present application includes but is not limited to the application scenario of word segmentation, and can also be applied to other application scenarios such as new word discovery, syntactic analysis, and keyword extraction.

[0140] In summary, it can be found that the method provided by the embodiments of the present application discovers new words based on the method of sequence pattern mining, and at the same time sets the matching rules of the word segmentation word library. At this time, it can better identify names such as professional terms and structure names, reduce the problems caused by incorrect splitting of the existing word segmentation tool, and thus can greatly improve the accuracy of other subsequent tasks.

[0141] The embodiments of the present application provide a device for determining new words, such as Figure 3 As shown, the device 60 for determining new words may include: a text acquisition module 601, a sequence mining module 602, a super-sequence determination module 603, and a new word determination module 604, where

[0142] The text acquisition module 601 is configured to acquire a sample text set;

[0143] The sequence mining module 602 is configured to perform character sequence mining on the sample text set to obtain frequent character sequences corresponding to each length;

[0144] The super-sequence determination module 603 is configured to determine each super-sequence in the frequent character sequences corresponding to each length;

[0145] A new word determination module 604, which is used to determine, for each supersequence, that if the supersequence is not included in each word segmentation included in the sample text set, the supersequence is determined as a new word.

[0146] Optionally, when the sequence mining module performs word sequence mining on the sample text set to obtain frequent word sequences corresponding to each length, it is specifically used for:

[0147] Determine the number of samples corresponding to each word element included in the sample text set. For a word element, the number of samples corresponding to the word element refers to the number of sample texts in the sample text set that contain the word sequence element;

[0148] Based on the number of samples corresponding to each word element, filter the word elements included in the sample text set to obtain a processed sample text set;

[0149] Perform word sequence mining on the processed sample text set to obtain frequent word sequences corresponding to each length.

[0150] Optionally, when the sequence mining module filters the word elements included in the sample text set based on the number of samples corresponding to each word element to obtain a processed sample text set, it is specifically used for:

[0151] For the number of samples of a word element, if the number of samples meets the set conditions, delete the word element from the sample text set;

[0152] The number of samples meeting the set conditions includes at least one of the following:

[0153] The number of samples is less than the set value or the proportion of the number of samples is less than the preset value;

[0154] Wherein, the proportion of the number of samples refers to the ratio of the number of samples corresponding to the word element to the number of sample texts included in the sample text set.

[0155] Optionally, when the sequence mining module performs word sequence mining on the sample text set to obtain frequent word sequences corresponding to each length, it is specifically used for:

[0156] Based on the PrefixSpan algorithm, perform word sequence mining on the sample text set to obtain frequent word sequences corresponding to each length.

[0157] Optionally, when the supersequence determination module determines each supersequence in the frequent word sequences corresponding to each length, it is specifically used for:

[0158] Perform particle word filtering on each frequent word sequence respectively to obtain each filtered frequent word sequence;

[0159] Determine each supersequence in each filtered frequent word sequence.

[0160] Optionally, when the text acquisition module acquires the sample text set, it is specifically used for:

[0161] Acquire an initial text set, where each initial text is included in the initial text set;

[0162] Perform text preprocessing on each initial text respectively to obtain a preprocessing result corresponding to each initial text;

[0163] Based on the preprocessing results corresponding to each initial text, obtain the sample text set;

[0164] Among them, the text preprocessing includes at least one of sentence splitting processing and specific character deletion processing.

[0165] The device for determining new words in the embodiments of the present application can execute a method for determining new words provided by the embodiments of the present application, and its implementation principle is similar, which will not be elaborated here.

[0166] The embodiments of the present application provide a text processing device, as Figure 4 shown, the text processing device 70 may include: a text acquisition module 701 and a word segmentation processing module 702, where

[0167] The text acquisition module 701 is used to acquire the text to be processed;

[0168] The word segmentation processing module 702 is used to perform word segmentation processing on the text to be processed based on the word segmentation database to obtain the word segments included in the text to be processed; among them, the new words determined by the method in the above embodiments are included in the word segmentation database.

[0169] The text processing device in the embodiments of the present application can execute a text processing method provided by the embodiments of the present application, and its implementation principle is similar, which will not be elaborated here.

[0170] The embodiments of the present application provide an electronic device, as Figure 5 shown, Figure 5 The electronic device 2000 shown includes: a processor 2001 and a memory 2003. Among them, the processor 2001 and the memory 2003 are connected, such as connected through a bus 2002. Optionally, the electronic device 2000 may further include a transceiver 2004. It should be noted that in practical applications, the transceiver 2004 is not limited to one, and the structure of the electronic device 2000 does not constitute a limitation to the embodiments of the present application.

[0171] Among them, the processor 2001 is applied in the embodiments of the present application and is used to implement Figure 3 and Figure 4 the functions of each module shown.

[0172] The processor 2001 can be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 2001 can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0173] The bus 2002 can include a path for transmitting information between the above components. The bus 2002 can be a PCI bus or an EISA bus, etc. The bus 2002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 5 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0174] The memory 2003 can be a ROM or other types of static storage devices that can store static information and computer programs, a RAM, or other types of dynamic storage devices that can store information and computer programs. It can also be an EEPROM, a CD-ROM, or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store or in the form of a data structure the desired computer program and can be accessed by a computer, but is not limited thereto.

[0175] The memory 2003 is used to store the computer program of the application program for executing the solution of this application, and is controlled by the processor 2001 for execution. The processor 2001 is used to execute the computer program of the application program stored in the memory 2003 to implement Figure 3 or Figure 4 the actions of the device in the illustrated embodiments.

[0176] An embodiment of this application provides an electronic device, including a processor and a memory: The memory is configured to store a computer program, and when the computer program is executed by the processor, the processor implements any one of the methods in the above embodiments.

[0177] An embodiment of this application provides a computer-readable storage medium, which is used to store a computer program. When the computer program runs on a computer, the computer can execute any one of the methods in the above embodiments.

[0178] The nouns and implementation principles related to a computer-readable storage medium in this application can specifically refer to the methods in the embodiments of this application, and will not be elaborated here.

[0179] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise explicitly stated herein, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0180] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for determining new words, characterized in that, including: obtaining a sample text set; performing word segmentation on each sample text in the sample text set respectively to obtain each word segment included in the sample text set; performing word sequence mining on the sample text set to obtain frequent word sequences corresponding to each length; determining each supersequence in the frequent word sequences corresponding to each length; for each of the supersequences, if the supersequence is not included in each word segment included in the sample text set, determining the supersequence as a new word; wherein, the performing word sequence mining on the sample text set to obtain frequent word sequences corresponding to each length includes: determining the ratio corresponding to each word element of each sample text in the sample text set, and respectively adding each word element with the ratio greater than the minimum support threshold as a prefix to the prefix data set; wherein, for each word element, the ratio is the ratio of the number of sample texts containing the word element to the total number of sample texts, and the number of sample texts corresponding to the word element refers to the number of sample texts containing the word element in the sample text set; continuously performing the following update operation on the prefix data set until there is no adjacent suffix word element corresponding to each prefix in the prefix data set with a ratio greater than the minimum support threshold, or the length of each prefix reaches the maximum degree, and stopping the update operation, at this time, the prefixes of each length in the prefix data set are frequent word sequences: for each prefix in the prefix data set, determining the adjacent suffix word element of the prefix; if the ratio corresponding to the adjacent suffix word element of the prefix is greater than the minimum support threshold, merging the prefix and the adjacent suffix word element of the prefix, and adding the merging result as a new prefix to the prefix data set.

2. The method according to claim 1, characterized in that, the performing word sequence mining on the sample text set to obtain frequent word sequences corresponding to each length includes: determining the number of sample texts corresponding to each word element included in the sample text set; filtering the word elements included in the sample text set based on the number of sample texts corresponding to each word element to obtain a processed sample text set; performing word sequence mining on the processed sample text set to obtain frequent word sequences corresponding to each length.

3. The method according to claim 2, characterized in that, the filtering the word elements included in the sample text set based on the number of sample texts corresponding to each word element to obtain a processed sample text set includes: for the number of sample texts of a word element, if the number of sample texts meets the set condition, deleting the word element from the sample text set; the number of sample texts meeting the set condition includes at least one of the following: the number of sample texts is less than the set value or the proportion of the number of sample texts is less than the preset value; wherein, the proportion of the number of sample texts is the ratio of the number of sample texts corresponding to the word element to the number of sample texts included in the sample text set.

4. The method according to claim 1, characterized in that, the performing word sequence mining on the sample text set to obtain frequent word sequences corresponding to each length includes: performing word sequence mining on the sample text set based on the PrefixSpan algorithm of pattern mining based on prefix projection to obtain frequent word sequences corresponding to each length.

5. The method according to claim 1, characterized in that, the determining each supersequence in the frequent word sequences corresponding to each length includes: Perform particle word filtering on each of the frequent word sequences to obtain the filtered frequent word sequences; Determine each supersequence in each of the filtered frequent word sequences.

6. The method according to claim 1, characterized in that, The obtaining of the sample text set includes: Obtain an initial text set, where the initial text set includes each initial text; Perform text preprocessing on each of the initial texts respectively to obtain the preprocessing result corresponding to each initial text; Based on the preprocessing results corresponding to each of the initial texts, obtain the sample text set; Wherein, the text preprocessing includes at least one of clause splitting processing and specific character deletion processing.

7. A text processing method, characterized in that, The method includes: Obtain a text to be processed; Perform word segmentation processing on the text to be processed based on a word segmentation database to obtain the word segments included in the text to be processed, where the word segmentation database contains new words determined by using any one of the methods described in claims 1 to 6.

8. A device for determining new words, characterized in that, Includes: A text acquisition module, configured to acquire a sample text set; A sequence mining module, configured to perform word sequence mining on the sample text set to obtain frequent word sequences corresponding to each length; A supersequence determination module, configured to determine each supersequence in the frequent word sequences corresponding to each length; A new word determination module, configured to perform word segmentation on each sample text in the sample text set respectively to obtain each word segment included in the sample text set; for each of the supersequences, if the supersequence is not included in each word segment included in the sample text set, then determine the supersequence as a new word; Wherein, when the sequence mining module performs word sequence mining on the sample text set, it is specifically configured to: Determine the ratio corresponding to each word element of each sample text in the sample text set, and respectively use each word element whose ratio is greater than the minimum support threshold as a prefix and add it to the prefix data set; wherein, for each word element, the ratio is the ratio of the number of sample texts corresponding to the word element to the total number of sample texts, and the number of sample texts corresponding to the word element refers to the number of sample texts in the sample text set that contain the word element; Continuously perform the following update operations on the prefix data set until there is no adjacent suffix word element corresponding to each prefix in the prefix data set whose ratio is greater than the minimum support threshold, or the length of each prefix reaches the maximum degree, and stop the update operation. At this time, the prefixes of each length in the prefix data set are frequent word sequences: For each prefix in the prefix data set, determine the adjacent suffix word element of the prefix; if the ratio corresponding to the adjacent suffix word element of the prefix is greater than the minimum support threshold, then merge the prefix and the adjacent suffix word element of the prefix, and use the merged result as a new prefix and add it to the prefix data set.

9. The device according to claim 8, wherein, The sequence mining module is configured to: Determine the number of sample texts corresponding to each word element included in the sample text set; Based on the number of sample texts corresponding to each word element, filter the word elements included in the sample text set to obtain a processed sample text set; Perform word sequence mining on the processed sample text set to obtain frequent word sequences corresponding to each length.

10. The device according to claim 9, wherein, The sequence mining module is configured to: For the number of samples of a character element, if the number of samples meets the set conditions, delete the character element from the sample text set; The number of samples meeting the set conditions includes at least one of the following: The number of samples is less than the set value or the proportion of the number of samples is less than the preset value; Wherein, the proportion of the number of samples refers to the ratio of the number of samples corresponding to the character element to the number of sample texts included in the sample text set.

11. The device according to claim 8, wherein, The sequence mining module is used for: Based on the PrefixSpan algorithm for pattern mining of prefix projection, perform character sequence mining on the sample text set to obtain frequent character sequences corresponding to each length.

12. The device according to claim 8, wherein, The supersequence determination module is used for: Perform particle word filtering on each of the frequent character sequences respectively to obtain each filtered frequent character sequence; Determine each supersequence in each of the filtered frequent character sequences.

13. The device according to claim 8, wherein, The text acquisition module is used for: Acquire an initial text set, where the initial text set includes each initial text; Perform text preprocessing on each of the initial texts respectively to obtain a preprocessing result corresponding to each initial text; Based on the preprocessing results corresponding to each of the initial texts, obtain the sample text set; Wherein, the text preprocessing includes at least one of sentence splitting processing and specific character deletion processing.

14. A text processing device, wherein, Includes: A text acquisition module, used for acquiring a text to be processed; A word segmentation processing module, used for performing word segmentation processing on the text to be processed based on a word segmentation database to obtain the word segments included in the text to be processed, wherein the word segmentation database contains new words determined by any one of the methods described in claims 1 to 6.

15. An electronic device, wherein, Includes a processor and a memory: The memory is configured to store a computer program, and when the computer program is executed by the processor, the processor executes the method described in any one of claims 1 - 7.

Citation Information

Patent Citations

  • New word recognition method, device, computer device and storage medium

    CN109408818A

  • Named entity recognition method and device and computer equipment

    CN109858040A