A method, apparatus, and readable storage medium for processing text
By introducing subject matching relationship and context feature analysis, the problem of inaccurate keyword extraction in the prior art is solved, and the accuracy and representativeness of text keywords are improved.
Patent Information
- Application Number
- CN202110797668.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-07-14
AI Technical Summary
The existing keyword extraction methods are mainly based on word granularity characteristics, resulting in inaccurate keywords and inability to effectively express the core content of the text.
By introducing the first sample matching relationship and the second sample matching relationship of the subject, comprehensively consider the keywords and context characteristics, analyze the degree of correlation between the words and the subject, and determine the keywords of the target text.
It improves the accuracy of keyword extraction, avoids the inaccuracy problem caused by extraction from the perspective of word granularity, and enhances the representativeness of text keywords.
Smart Images

Figure CN113821528B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing, and in particular, to a method, device, and readable storage medium for processing text. Background Art
[0002] Category key pattern extraction is one of the important technologies in the development of natural language understanding technology, and there are many scenarios in practical applications. For example, for the processing and analysis of massive text data, a key step is to extract the most important information in the text, and the important information can often be characterized by several pattern features. Therefore, category key pattern extraction plays a very effective role. Another example is that in retrieval systems such as Baidu Library, by extracting article keywords and matching them with retrieval terms or calculating similarity, etc., the accuracy of the recalled results can be improved.
[0003] Currently, keyword extraction methods mainly adopt unsupervised methods, including: keyword extraction based on statistical features, keyword extraction based on graph network models, and keyword extraction based on topic models. Keyword extraction based on statistical features uses word weights, word position information, and word association information as statistical features for keyword extraction; keyword extraction based on graph network models first constructs a network graph of the document, uses the preprocessed words as nodes, the relationships between words as edges, and the weights between edges are generally represented by the association degree between words. Then, based on the network graph, the importance of each node is evaluated, and the nodes are sorted according to the importance. The words represented by the top several nodes in the sorting are selected as keywords; keyword extraction based on topic models first tokenizes the text and selects candidate keywords, obtains a topic model according to corpus learning, calculates the topic distribution of the article and the candidate keyword distribution based on the obtained latent topic model. Next, calculates the topic similarity between the document and the candidate keywords and sorts them, and selects the top n words as keywords.
[0004] However, when performing keyword extraction, keyword extraction based on statistical features, keyword extraction based on graph network models, and keyword extraction based on topic models only consider the feature extraction at the word granularity, and thus the extracted keywords cannot well express the core content of the text, resulting in inaccurate extracted keywords. Summary of the Invention
[0005] This application provides a method, device, and readable storage medium for processing text, which can improve the accuracy of text keyword extraction.
[0006] On the one hand, an embodiment of this application provides a method for processing text, including:
[0007] Determine the first keyword set and the second keyword set corresponding to each word in the target text. The first keyword set is the keyword set that matches each word in the target text in the first sample matching relationship, and the second keyword set is the keyword set that matches each word in the target text in the second sample matching relationship. The first sample matching relationship includes the matching relationship between the keywords of the first text category and the support degree, and the second sample matching relationship includes the matching relationship between the keywords of the second text category and the support degree;
[0008] Determine the number of first samples hit by each keyword in the first keyword set in the sample set associated with the first sample matching relationship, and determine the number of second samples hit by each keyword in the second keyword set in the sample set associated with the second sample matching relationship;
[0009] Determine the comprehensive support degree of each word according to the number of first samples, the number of second samples, the number of samples in the sample set associated with the first sample matching relationship, and the number of samples in the sample set associated with the second sample matching relationship. The comprehensive support degree of each word represents the degree of association between each word and any subject in the first text category and the second text category;
[0010] Determine the target subject of the target text from the first text category and the second text category according to the comprehensive support degree of each word;
[0011] Match each word with the keyword corresponding to the target subject to obtain the keyword of the target text.
[0012] The second aspect of the embodiments of the present application provides a text processing device, including:
[0013] A first determination unit, configured to determine the first keyword set and the second keyword set corresponding to each word in the target text. The first keyword set is the keyword set that matches each word in the target text in the first sample matching relationship, and the second keyword set is the keyword set that matches each word in the target text in the second sample matching relationship. The first sample matching relationship includes the matching relationship between the keywords of the first text category and the support degree, and the second sample matching relationship includes the matching relationship between the keywords of the second text category and the support degree;
[0014] A second determination unit, configured to determine the number of first samples hit by each keyword in the first keyword set in the sample set associated with the first sample matching relationship, and determine the number of second samples hit by each keyword in the second keyword set in the sample set associated with the second sample matching relationship;
[0015] A third determination unit, configured to determine the comprehensive support degree of each word according to the number of first samples, the number of second samples, the number of samples in the sample set associated with the first sample matching relationship, and the number of samples in the sample set associated with the second sample matching relationship, where the comprehensive support degree of each word represents the degree of association between each word and any subject in the first text category and the second text category;
[0016] A fourth determination unit, configured to determine the target subject of the target text from the first text category and the second text category according to the comprehensive support degree of each word;
[0017] A matching unit, configured to match each word with the keyword corresponding to the target subject to obtain the keyword of the target text.
[0018] In a possible design, the third determination unit is specifically configured to:
[0019] Determine the first support degree of each word according to the number of first samples and the number of samples in the sample set associated with the first sample matching relationship;
[0020] Determine the second support degree of each word according to the number of second samples and the number of samples in the sample set associated with the second sample matching relationship;
[0021] Determine the comprehensive support degree of each word according to the first support degree and the second support degree.
[0022] In a possible design, the fourth determination unit is specifically configured to:
[0023] Determine the first target words in each word whose comprehensive support degree is greater than the comprehensive support degree threshold;
[0024] Determine the subject corresponding to the first target word in the first text category and the second text category as the target subject;
[0025] Or,
[0026] Determine the second target word with the largest comprehensive support degree in each word;
[0027] Determine the subject corresponding to the second target word in the first text category and the second text category as the target subject.
[0028] In a possible design, the apparatus further includes:
[0029] A construction unit, and the construction unit is configured to:
[0030] Obtain a training text set, where the training text set includes the training texts associated with the first text category and the training texts associated with the second text category;
[0031] For each text in the training text set, perform sentence splitting to obtain the sentence set corresponding to each text;
[0032] Process the sentence set corresponding to each text to obtain the first character sequence corresponding to each text;
[0033] Remove the keywords in the first character sequence that are less than the support threshold to obtain the second character sequence corresponding to each text;
[0034] Determine the keywords in the second character sequence and the corresponding support degrees of the keywords;
[0035] Determine the keywords in the second character sequence corresponding to the first text category and the corresponding support degrees of the keywords as the first sample matching relationship of the first text category;
[0036] Determine the keywords in the second character sequence corresponding to the second text category and the corresponding support degrees of the keywords as the second sample matching relationship of the second text category.
[0037] In a possible design, the building unit processes the sentence set corresponding to each text, and the steps to obtain the first character sequence corresponding to each text include:
[0038] Filter at least one of punctuation marks, letters, and numbers in the sentence set corresponding to each text;
[0039] Split the filtered sentence set corresponding to each text into word units to obtain the first character sequence.
[0040] In a possible design, the steps for the building unit to determine the keywords in the second character sequence include:
[0041] Determine the target keywords with the character number i in the second character sequence and the keyword set associated with the target keywords in the target character sequence set. The target character sequence set is the sample sentence set corresponding to at least one of the first text category and the second text category, and the value of i is the numerical value of the character number in the second character sequence;
[0042] Remove the keywords in the keyword set that are less than the support threshold;
[0043] Recursively based on the character number of the keywords in the second character sequence until the keywords with the most character numbers in the second character sequence and the keyword set corresponding to the keywords with the most character numbers in the target character sequence set are determined, so as to obtain the keywords with different character numbers in the second character sequence.
[0044] In a possible design, the building unit is further used for:
[0045] Perform sentence splitting on the target text to obtain the sentence set corresponding to the target text;
[0046] Process the set of clauses corresponding to the target text to obtain a third word sequence;
[0047] Update the keywords and support degrees in the first sample matching relationship of the target subject according to the third word sequence.
[0048] Another aspect of the embodiments of the present application provides a computer device, which includes at least one connected processor, a memory, and a transceiver. Among them, the memory is used to store program codes, and the processor is used to call the program codes in the memory to execute the steps of the text processing method described in the above aspects.
[0049] Another aspect of the embodiments of the present application provides a computer storage medium, which includes instructions that, when running on a computer, cause the computer to execute the steps of the text processing method described in the above aspects.
[0050] In summary, it can be seen that in the embodiments of the present application, determine the first keyword set and the second keyword set corresponding to each word in the target text; determine the number of first samples hit by each keyword in the first keyword set in the sample set associated with the first sample matching relationship, and determine the number of second samples hit by each keyword in the second keyword set in the sample set associated with the second sample matching relationship; determine the comprehensive support degree of each word according to the number of first samples, the number of second samples, the number of samples in the sample set associated with the first sample matching relationship, and the number of samples in the sample set associated with the second sample matching relationship. The comprehensive support degree of each word represents the degree of association between each word and at least one subject in the first text category and the second text category; determine the target subject of the target text from the first text category and the second text category according to the comprehensive support degree of each word; match each word with the keyword corresponding to the target subject to obtain the keyword of the target text. It can be seen from this that in the present application, when extracting keywords in the text, the first sample matching relationship and the second sample matching relationship of the subject are introduced, the keywords and the context features that appear together with the keywords are comprehensively considered, the subject to which the target text belongs is determined by analyzing the degree of association between each word and the subject, and each word is matched with the keyword of the subject to which the target text belongs to determine the keyword of the target text. In this way, it is possible to avoid extracting keywords only from the perspective of word granularity when extracting keywords in the text, thereby causing inaccurate extraction of keywords. Description of the Drawings
[0051] Figure 1 It is a network architecture diagram of the text processing system provided by the embodiments of the present application:
[0052] Figure 2 It is a schematic diagram of an embodiment of the text processing method provided by the embodiments of the present application;
[0053] Figure 3 Another schematic diagram of the text processing method provided by the embodiment of the present application;
[0054] Figure 4 Schematic diagram of the virtual structure of the text processing device provided by the embodiment of the present application;
[0055] Figure 5 Schematic diagram of the hardware structure of the server provided by the embodiment of the present application. Detailed implementation manners
[0056] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.
[0057] Terms such as "first" and "second" in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that shown or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or modules does not have to be limited to those steps or modules clearly listed, but may include other steps or modules not clearly listed or inherent to these processes, methods, products or devices. The division of modules in the present application is only a logical division, and there may be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another system, or some feature vectors can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between each other may be through some interfaces. The indirect coupling or communication connection between modules may be electrical or other similar forms, which are not limited in the present application. And the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed to multiple circuit modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present application.
[0058] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an architecture of the text processing system in the embodiment of the present application. As Figure 1 shown, it includes a client 101, a network 102, a server 103, and K databases 104. Each of the K databases stores a sample matching relationship corresponding to a subject;
[0059] The client 101 obtains the target text. The client 101 uploads the target text to the server 103 through the network 102. After the server 103 obtains the target text, it can process the target text to obtain each word segment of the target text, and match each word segment of the processed target text with the first sample matching relationship and the second sample matching relationship respectively, to obtain the first keyword set that matches each word in the first sample matching relationship and the second keyword set that matches each word in the second sample matching relationship. Then, determine the first text books hit by each keyword in the first keyword set in the sample set associated with the first sample matching relationship, and determine the second sample numbers hit by each keyword in the second keyword set in the sample set associated with the second sample matching relationship, and determine the comprehensive support degree of each word according to the first sample number, the second sample number, the number of samples in the sample set associated with the first sample matching relationship, and the number of samples in the sample set associated with the second sample matching relationship. Among them, the comprehensive support degree of each word represents the degree of association between each word and any subject in the first text category and the second text category, and determine the target subject of the target text from the first text category and the second text category according to the comprehensive support degree of each word; Finally, match each word with the keyword corresponding to the target subject to obtain the keyword of the target text. It can be seen from this that in this application, when extracting keywords in the text, by introducing the first sample matching relationship and the second sample matching relationship of the subject, comprehensively considering the keywords and the context features that appear together with the keywords, and determining the subject to which the target text belongs by analyzing the degree of association between each word and the subject, and matching each word with the keyword of the subject to which the target text belongs to determine the keyword of the target text. In this way, it is possible to avoid extracting keywords only from the word granularity perspective when extracting keywords from the text, which may lead to inaccurate extraction of keywords.
[0060] The following is combined with Figure 2 the text processing method of this application for illustration. Please refer to Figure 2 , Figure 2 which is a schematic diagram of an embodiment of the text processing method provided by an embodiment of this application, including:
[0061] 201. Determine the mining category and divide the training samples of each category.
[0062] In this embodiment, a database can be constructed and training sets of multiple categories can be obtained. These categories can be, for example, news categories, entertainment categories, sports categories, etc. First, the category to be mined needs to be determined. Taking the news category as an example, the news events in this news category can be divided into company events, market events, macro policy events, etc. At the same time, each category is further divided into secondary categories and tertiary categories. For example, company events can be divided into multiple secondary categories and tertiary categories, as shown in Table 1 below:
[0063] Table 1
[0064]
[0065]
[0066] It can be understood that the above Table 1 is only for illustration. Of course, there are also other categories and category contents, which are not specifically limited. Then, the training sample sets of each event type are obtained. Through the historical text data of various categories that have been obtained or manual marking, the sample training sets of event types are obtained to complete the data preparation work.
[0067] 202. Mine the pattern features of each category through sequence pattern.
[0068] In this embodiment, the pattern features of each category can be mined through sequence pattern. Based on the training samples of each category text, the frequent sequence pattern features of each category are mined based on the frequent word sequence pattern, that is, the keywords of each category and their corresponding support degrees.
[0069] 203. Calculate the comprehensive support degree of each word in the text of the keyword to be extracted in each category.
[0070] In this embodiment, after the pattern features of each category are mined, each word in the text of the keyword to be extracted can be respectively matched with the pattern features of each category to obtain the comprehensive support degree of each word in each category.
[0071] 204. Determine the keyword of the text of the keyword to be extracted according to the comprehensive support degree of each word in each category.
[0072] In this embodiment, the target category with the highest comprehensive support degree can be used as the category of the text for keyword extraction. Then, each word in the text for keyword extraction is matched with the keyword corresponding to the target category to obtain the keyword of the text for keyword extraction.
[0073] In summary, it can be seen that in the embodiments provided by the present application, by mining the frequent pattern features of various categories, comprehensively considering the keywords and the context features that appear together with the keywords, and determining the category to which the text of the keyword to be extracted belongs by analyzing the degree of association between each word in the text of the keyword to be extracted and the category, and matching each word with the keywords of the category to which the text of the keyword to be extracted belongs to determine the keywords of the text of the keyword to be extracted. Thus, when extracting keywords from the text, it is possible to avoid extracting text features only from the word granularity perspective, which may lead to inaccurate extraction of keywords.
[0074] Combined with the above introduction, the processing method of the text in the present application will be introduced from the perspective of a text processing device below. The text processing device can be a server or a service unit in the server.
[0075] Please refer to Figure 3 , Figure 3 which is a schematic diagram of an embodiment of the processing method of the text provided by the embodiment of the present application, including:
[0076] 301. Determine the first keyword set and the second keyword set corresponding to each word in the target text.
[0077] In this embodiment, the text processing device can obtain a target text, which is a text for which keywords are to be extracted, and process the target text. Here, the processing is to split sentences according to punctuation marks and filter at least one of the tag symbols, letters, and data in the split sentences. Finally, a set of words corresponding to the target text is obtained (in addition, to simplify the process, a threshold value, such as 3, can also be set for the set of words corresponding to the target text, and words that appear less than 3 times in the words corresponding to the target text are removed to obtain a set of words, which is not specifically limited). After that, a first keyword set and a second keyword set corresponding to each word in the target text are determined. The first keyword set is a keyword set that matches each word in the target text in the first sample matching relationship. The second keyword set is a keyword set that matches the degree of each word in the target text in the second sample matching relationship. The first sample matching relationship includes the matching relationship between the keywords of the first text category and the support degree. The second sample matching relationship includes the matching relationship between the keywords of the second text category and the support degree. The first text category can be, for example, the entertainment category, and the second text category is other categories except the entertainment category, such as the sports category, the news category, the current affairs category, etc., which is not specifically limited (it can be understood that when each word in the target text is matched with the keywords in the first sample matching relationship and the second sample matching relationship, if each word in the target text fails to match the first sample matching relationship and the second sample matching relationship, the support degree threshold can be reduced, the first sample matching relationship and the second sample matching relationship are adjusted, and based on the adjusted first sample matching relationship and the second sample matching relationship, each word in the target text is matched again).
[0078] The construction of the first sample matching relationship and the second sample matching relationship will be described below:
[0079] Step 1: Obtain a training text set, which includes training samples associated with the first text category and training samples associated with the second text category.
[0080] In this step, the text processing device can obtain training samples associated with the first text category and training samples associated with the second text category. The first text category can be, for example, the "news category", and the second text category is all other categories except the "news category", such as the "entertainment category", the "sports category", and the "current affairs category", etc. That is, the text processing device can obtain the historical acquired category text data or manually marked text data corresponding to each subject as the training data samples for the first sample matching relationship and the second sample matching relationship of the first text category. The training data samples will be described below with actual samples. Please refer to Table 2, which is the training data sample associated with the event type of "company event_company operation_performance growth" in the news category:
[0081] Table 2
[0082]
[0083] Step 2: Sentence each text in the training text set to obtain a sentence set corresponding to each text.
[0084] In this embodiment, after obtaining the training text set, the text processing device can divide each text in the training text set into sentences to obtain a sentence set for each text. The sentence division here divides each text in the training text set by regular matching punctuation marks to obtain a sentence set corresponding to each text.
[0085] Step 3: Process the sentence set corresponding to each text to obtain the first word sequence corresponding to each text.
[0086] In this embodiment, after obtaining the sentence set corresponding to each text, the text processing device can process the sentence corresponding to each text to obtain the first word sequence corresponding to each text. Specifically, the text processing device can filter at least one of punctuation marks, letters and numbers in the sentence set corresponding to each text by regular filtering. For ease of understanding, the text processing is described below with reference to a specific example. Table 3 is a data sample after processing the training data sample in Table 2:
[0087] Table 3
[0088] Processed data sample In this year, Rizhao Iron & Steel's monthly performance increase ranked first in the province year-on-year Great Wall Motor's monthly sales increased significantly month-on-month, breaking the market ice with actions Li-Ning expects its mid-term earnings to increase by more than 100 million yuan year-on-year In the first half of the year, China Shenhua's net profit of Shenhua Finance increased year-on-year Shuangjian Co., Ltd. expects its profit in the first half of the year to increase both year-on-year and month-on-month, reaching a record high, exceeding previous years The sales volume of electric vehicles of BMW Group exceeded the 10,000 mark, and the monthly delivery increased month-on-month In this year, Tencent Video's monthly business revenue increased month-on-month The annual net profit of Beauty Net increased to HK$100 million year-on-year After the listing of Tencent Music, its performance showed an obvious upward trend month-on-month
[0089] Afterwards, the sentence set corresponding to each filtered text is split according to character units to obtain the first character sequence (if the text is a Chinese text, the character unit can be a single character, if the text is an English text, the character unit can be a single word, and there is no specific limitation). For example, the data sample in Table 3 "This year, Rizhao Steel's year-on-year growth rate ranks first in the province" is split according to character units to obtain: "this, year, month, part, day, Zhao, steel, iron, industry, performance, same, ratio, increase, amplitude, position, column, all, province, first" is the first character sequence corresponding to the data sample "This year, Rizhao Steel's year-on-year growth rate ranks first in the province", thereby obtaining the first character sequence in the sentence set corresponding to each text.
[0090] Step 4: Remove keywords whose support value is less than a threshold value from the first word sequence to obtain a second word sequence corresponding to each text.
[0091] In this embodiment, the text processing device can count the number of sample clauses in which all word sequences appear in the clauses of each sample object, and filter out keywords with a support threshold less than the minimum. Assume that the support threshold within the smallest subject is 1 / 3, that is, it must appear at least 4 times in the 9 data samples in Table 3 to meet the support threshold, otherwise the keyword is filtered out. Taking the 9 data samples in Table 3 as an example, the keywords with a support threshold less than the minimum in the first word sequence are removed, and the second word sequence corresponding to each text is obtained. The results are shown in Table 4, including each keyword and the number of samples in which each keyword appears:
[0092] Table 4
[0093] Keywords Ratio Increase Year Month-on-month Year-on-year Month The number of samples that appear in the data sample 9 8 6 5 5 4
[0094] The results obtained after filtering the data samples in Table 3 through the support threshold are shown in Table 5:
[0095] Table 5
[0096] Support threshold filtering result Year-on-year increase in [Month] of [Year] Month-on-month increase in [Month] Year-on-year increase Year-on-year increase in [Year] Year-on-year and month-on-month increase in [Year] Month-on-month increase in [Month] Month-on-month increase in [Month] of [Year] Year-on-year increase in [Year] Month-on-month
[0097] It should be noted that in this application, word sequences are used as the objects for keyword mining, and keywords with various character counts that meet the support threshold in the text are mined based on the Prefixspan algorithm (the Prefixspan algorithm is an association algorithm used to mine frequent sequence patterns in text). The calculation of the support threshold can be carried out through the following formula:
[0098] min_sup = a × n
[0099] Among them, min_sup is the support threshold, n is the number of training texts in a certain subject, and a is the minimum support rate. The minimum support rate is adjusted according to the magnitude of the training data set.
[0100] Step 5: Determine the keywords in the second word sequence and the support corresponding to the keywords.
[0101] In this step, the text processing device can determine the keywords in the second word sequence and the support corresponding to the keywords. The following is a specific description:
[0102] Step 51: Determine the target keywords with the number of characters i in the second word sequence and the set of keywords associated with the target keywords in the target word sequence set. The target word sequence set is the set of word sequences after support filtering corresponding to the first text category or the second text category. Taking the above example for illustration, it is the result after support threshold filtering of Table 5. Here, the value of i is the numerical value of the number of characters in the second word sequence. Taking Table 5 as an example for illustration, with the value of the number of characters i in the second word sequence starting from 1 and recursively increasing by 1 each time as the benchmark, determine the target keywords with the number of characters i and the set of keywords associated with the target keywords in the target word sequence set. Table 6 illustrates with i being 1:
[0103] Table 6
[0104]
[0105]
[0106] Step 52: Eliminate the keywords in the keyword set that are less than the support threshold.
[0107] In this step, the text processing device can eliminate the keywords in the keyword set that are less than the support threshold. The support threshold is 1 / 3 (that is, the number of occurrences of this word sequence must be greater than 3 times). Then, the keywords in the keyword set that are less than the support threshold can be eliminated. The following is an illustration in combination with Table 7:
[0108] Table 7
[0109]
[0110]
[0111] In Table 6, among the keywords associated with "year" in the target sequence set, namely "ring", "year", and "month", none of them reach the support threshold. Therefore, they need to be eliminated.
[0112] Step 53: Recursively proceed based on the number of characters of the keywords in the second word sequence until the keyword with the most characters in the second word sequence and the set of keywords corresponding to the keyword with the most characters in the target word sequence set are determined, thereby obtaining the keywords of the second word sequence.
[0113] In this step, the text processing device recursively proceeds based on the number of characters of the keywords in the second word sequence until the keyword with the most characters in the second word sequence and the set of keywords corresponding to the keyword with the most characters in the target word sequence set are determined, thereby obtaining the keywords of the second word sequence.
[0114] The following takes the example of the target keyword with a character count of 2 obtained after secondary recursion with "year" and explains it in combination with Table 6. After excluding "ring", "year", and "month" that do not reach the support threshold, the target keywords with a character count of 2 that meet the support threshold are recursively obtained respectively (which are "year-on-year", "year-on-year increase", "year-on-year same") and the keyword set associated with the target keyword in the target word sequence set, as shown in Table 8 specifically:
[0115] Table 8
[0116]
[0117] Continuing with the example of the target keyword with a character count of 3 obtained after three recursions with "year" (in the keyword set associated with the target keyword in the target word sequence set in Table 7, "year" and "ring" do not reach the support threshold, and after excluding them, the target keywords with a character count of 3 are recursively obtained, which are "year-on-year increase", "year-on-year same", "year-on-year increase in the same period"), as shown in Table 9 specifically:
[0118] Table 9
[0119]
[0120] Continuing with the example of the target keyword with a character count of 4 obtained after four recursions with "year" (in the keyword set associated with the target keyword in the target word sequence set in Table 8, "year" does not reach the support threshold, and after excluding it, the target keyword with a character count of 4 is recursively obtained, which is "year-on-year increase in the same period", and by matching it with the keywords in the keyword set in Table 5, "year-on-year increase in the same period" can be obtained), as shown in Table 10 specifically:
[0121] Table 10
[0122]
[0123] So far, the recursion of the keywords corresponding to "year" ends. Based on the above principle, other target keywords and the keyword set associated with the target keyword in the target word sequence set can be obtained.
[0124] Step 54: Determine the first sample matching relationship of the first text category by taking the keywords of the second word sequence corresponding to the first text category and the support degrees corresponding to the keywords.
[0125] In this embodiment, after obtaining the keywords of the second character sequence corresponding to each text in the training text set, the text processing device may determine the support degree corresponding to the keywords, and determine the keywords of the second character sequence corresponding to the first text category and the support degree corresponding to the keywords as the first sample matching relationship of the first text category. The following takes the keyword "year" as an example in combination with Table 11. Table 10 is the first sample matching relationship of the keyword "year" corresponding to the first text category:
[0126] Table 11
[0127]
[0128] Taking "year" as an example, when the number of characters is 1, the support degree in the corresponding main body is 5 / 9, that is, "year" appears 5 times in 9 sample texts. Thus, the keywords in the second character sequence corresponding to the first text category and the support degree corresponding to the keywords can be obtained.
[0129] Step 55: Determine the keywords of each number of characters in the second character sequence corresponding to the second text category and the support degree corresponding to the keywords of each number of characters as the second sample matching relationship of the second text category.
[0130] In this step, based on Steps 51 to 55, the second sample matching relationship of the second text category can be mined, that is, the sample matching relationship of the second text category. After obtaining the keywords corresponding to each text in the training text set and the support degree corresponding to the keywords, that is, obtaining the keywords of each number of characters in the second character sequence corresponding to the second text category and the support degree corresponding to the keywords of each number of characters, the second text category is a category other than the first text category.
[0131] 302. Determine the number of first samples hit by each keyword in the first keyword set in the sample set associated with the first sample matching relationship, and determine the number of second samples hit by each keyword in the second keyword set in the sample set associated with the second sample matching relationship.
[0132] In this embodiment, after the text processing device obtains the first keyword set and the second keyword set, it can determine the number of first samples hit by each keyword in the first keyword set in the sample set associated with the first sample matching relationship, and determine the number of second samples hit by each keyword in the second keyword set in the sample set associated with the second sample matching relationship. That is, for each keyword in the first keyword set, it is necessary to find the number of training texts hit by each keyword in the training text set corresponding to the first text category. For example, the training data samples in Table 2 are the training text set corresponding to the first text category, and "year" is a keyword in the first keyword set. Thus, "year" can be matched with the training data samples in Table 2, and the number of samples hit by the keyword "year" is 5. Similarly, the number of first samples hit by each keyword in the first keyword set in the sample set associated with the first sample matching relationship and the number of second samples hit by each keyword in the second keyword set in the sample set associated with the second sample matching relationship can be obtained.
[0133] 303. Determine the comprehensive support degree of each word according to the number of first samples, the number of second samples, the number of samples in the sample set associated with the first sample matching relationship, and the number of samples in the sample set associated with the second sample matching relationship.
[0134] In this embodiment, after the text processing device obtains the number of first samples and the number of second samples, it can determine the number of samples in the sample set associated with the first sample matching relationship and the number of samples in the sample set associated with the second sample matching relationship, that is, determine the number of samples in the training data samples corresponding to the first text category. For example, the training data samples in Table 5 are the training data samples corresponding to the first text category, then the number of samples in the sample set associated with the first sample matching relationship is 9. After that, the first support degree of each word can be determined according to the number of first samples and the number of samples in the sample set associated with the first sample matching relationship. Specifically, it can be calculated by the following formula:
[0135]
[0136] And determine the second support degree of each word according to the number of second samples and the number of samples in the sample set associated with the second sample matching relationship. Specifically, it can be calculated by the following formula:
[0137]
[0138] It can be understood that the second text category can be 1 or multiple, and there is no specific limitation. Correspondingly, the second sample matching relationship corresponds to the second text category.
[0139] Finally, determine the comprehensive support degree of each word according to the first support degree and the second support degree. Specifically, it can be calculated by the following formula:
[0140]
[0141] Among them, since the second support degree of each word in the target text may be 0, it is necessary to add a constant to ensure that the denominator is not 0, and at the same time, this constant cannot affect the comprehensive support degree of each word calculated. Therefore, δ≥0, and δ is a very small number, such as 0.01 or 0.001, and there is no specific limitation.
[0142] 304. Determine the target subject of the target text from the first text category and the second text category according to the comprehensive support degree of each word.
[0143] In this embodiment, after the comprehensive support degree of each word, the text processing device can determine the target subject of the target text from the first text category and the second text category according to the comprehensive support degree of each word. That is, when determining the comprehensive support degree of each word, the comprehensive support degree of each word represents the degree of association between each word and any one subject in the first text category and the second text category. Each word will have a comprehensive support degree relative to any one of the subjects. Thus, the first target words with a comprehensive support degree greater than the comprehensive support degree threshold in the comprehensive support degree of each word can be determined, and the subjects corresponding to the first target words in the first text category and the second text category are determined as the target subjects; or, the second target words with the largest comprehensive support degree of each word are determined; the subjects corresponding to the second target words in the first text category and the second text category are determined as the target subjects. For example, there are three subjects S1, S2, and S3, and the target text has three words, A, B, and C. Among them, the comprehensive support degree of word A relative to S1 is 0.8, relative to S2 is 0.1, and relative to S3 is 0.2. The comprehensive support degree of B relative to S1 is 0.1, relative to S2 is 0.75, and relative to S3 is 0.2. The comprehensive support degree of C relative to S1 is 0.65, relative to S2 is 0.01, and relative to S3 is 0.03. Then, the subject S1 can be determined as the subject corresponding to the target text. Or, if the comprehensive support degree threshold is 0.7, then the subjects S1 and S2 can be used as the subjects corresponding to the target text.
[0144] 305. Match each word with the keyword corresponding to the target subject to obtain the keyword of the target text.
[0145] In this embodiment, after determining the target subject corresponding to the target text, each word in the target text can be matched with the keyword corresponding to the target subject to obtain the keyword of the target text, that is, the keyword identical to each word in the target subject is used as the keyword of the target text.
[0146] It should be noted that the text processing device can also perform sentence segmentation on the target text to obtain the sentence combination corresponding to the target text, and process the sentence set corresponding to the target text to obtain the third word sequence, and update the keywords and support in the first sample matching relationship of the target subject according to the third word sequence. That is, the text processing device can use the target text as a training sample in the target subject to update the first sample matching relationship of the target subject. For details, please refer to the above steps 51 to 55, which have been described in detail above and will not be repeated here.
[0147] It should also be noted that, for different categories of time texts, common sequence patterns in different categories of event texts and sequence pattern features of each category (that is, keywords of each category can be extracted through the above steps) can also be found. An example is shown in Table 12:
[0148] Table 12
[0149]
[0150] The support of the sequence pattern in category 1 and category 2 is shown in Table 13:
[0151] Table 13
[0152] Sequential pattern Within-class support of Category 1 Within-class support of Category 2 Revenue & Growth 4 / 9 1 / 3 [Year] & Revenue & Year-on-year & Growth 2 / 9 1 / 99 Revenue & Growth & Blocked 1 / 90 1 / 3
[0153] Here, the keywords with high support in this category and low support in other categories are used as the keywords of the event text relative to this category. The sequence patterns "income & growth" and "year & income & year-on-year & growth" are the category key sequences of the event text "Tencent's total revenue in 2018 was 312.7 billion yuan, a year-on-year increase of 32%" in category 1. The sequence pattern "income & growth & obstruction" is the category key pattern of the event text "The growth of revenue from games and other industries is obstructed, Tencent plummeted, can the B-end rise again?" in category 2.
[0154] In summary, it can be seen that in the embodiments provided in the present application, when extracting keywords of text, the first sample matching relationship and the second sample matching relationship of the subject are introduced, and the keywords and the context features that appear together with the keywords are comprehensively considered. By analyzing the degree of association between each word and the subject, the subject to which the target text belongs is determined, and each word is matched with the keywords of the subject to which the target text belongs to determine the keywords of the target text. This can avoid extracting keywords of text only from the perspective of word granularity, thereby avoiding the problem of inaccurate extraction of keywords.
[0155] The embodiments of the present application have been described above from the perspective of the text processing method. Next, the embodiments of the present application will be described from the perspective of the text processing device.
[0156] Please refer to Figure 4 , in the embodiments of the present application, a text processing device is provided. The text processing device 400 includes:
[0157] A first determination unit 401, configured to determine a first keyword set and a second keyword set corresponding to each word in the target text. The first keyword set is a keyword set that matches each word in the target text in the first sample matching relationship, and the second keyword set is a keyword set that matches each word in the target text in the second sample matching relationship. The first sample matching relationship includes the matching relationship between the keywords of the first text category and the support degree, and the second sample matching relationship includes the matching relationship between the keywords of the second text category and the support degree;
[0158] A second determination unit 402, configured to determine the number of first samples hit by each keyword in the first keyword set in the sample set associated with the first sample matching relationship, and determine the number of second samples hit by each keyword in the second keyword set in the sample set associated with the second sample matching relationship;
[0159] A third determination unit 403, configured to determine the comprehensive support degree of each word according to the number of first samples, the number of second samples, the number of samples in the sample set associated with the first sample matching relationship, and the number of samples in the sample set associated with the second sample matching relationship. The comprehensive support degree of each word represents the degree of association between each word and any subject in the first text category and the second text category;
[0160] A fourth determination unit 404, configured to determine the target subject of the target text from the first text category and the second text category according to the comprehensive support degree of each word;
[0161] A matching unit 405, configured to match each word with the keywords corresponding to the target subject to obtain the keywords of the target text.
[0162] In a possible design, the third determination unit 403 is specifically configured to:
[0163] Determine the first support degree of each word according to the number of first samples and the number of samples in the sample set associated with the first sample matching relationship;
[0164] Determine the second support degree of each word according to the number of second samples and the number of samples in the sample set associated with the second sample matching relationship;
[0165] Determine the comprehensive support degree of each word according to the first support degree and the second support degree.
[0166] In a possible design, the fourth determination unit 404 is specifically configured to:
[0167] Determine the first target words in each word whose comprehensive support degree is greater than the comprehensive support degree threshold;
[0168] Determine the subjects corresponding to the first target words in the first text category and the second text category as the target subjects;
[0169] Or,
[0170] Determine the second target words with the largest comprehensive support degree in each word;
[0171] Determine the subjects corresponding to the second target words in the first text category and the second text category as the target subjects.
[0172] In a possible design, the text processing device further includes:
[0173] A construction unit 406, and the construction unit 406 is configured to:
[0174] Obtain a training text set, where the training text set includes training texts associated with the first text category and training texts associated with the second text category;
[0175] Perform sentence splitting on each text in the training text set to obtain a sentence splitting set corresponding to each text;
[0176] Process the sentence splitting set corresponding to each text to obtain a first character sequence corresponding to each text;
[0177] Eliminate keywords with a support degree less than the support degree threshold in the first character sequence to obtain a second character sequence corresponding to each text;
[0178] Determine the keywords in the second character sequence and the support degree corresponding to the keywords;
[0179] Determine the keywords in the second character sequence corresponding to the first text category and the support degree corresponding to the keywords as the first sample matching relationship of the first text category;
[0180] Determine the keywords in the second word sequence corresponding to the second text category and the support degree corresponding to the keywords as the second sample matching relationship of the second text category.
[0181] In a possible design, the construction unit 406 processes the set of clauses corresponding to each text, and the obtained first word sequence corresponding to each text includes:
[0182] Filter at least one of punctuation marks, letters, and numbers in the set of clauses corresponding to each text;
[0183] Split the set of clauses corresponding to each text after filtering by word units to obtain the first word sequence.
[0184] In a possible design, the construction unit 406 determines the keywords of the second word sequence, including:
[0185] Determine the target keywords with the character number i in the second word sequence and the set of keywords associated with the target keywords in the set of target word sequences. The set of target word sequences is the set of sample clauses corresponding to at least one subject in the first text category and the second text category. The value of i is the numerical value of the character number in the second word sequence;
[0186] Eliminate the keywords in the set of keywords that are less than the support degree threshold;
[0187] Recursively based on the character number of the keywords in the second word sequence until the keywords with the most character numbers in the second word sequence and the set of keywords corresponding to the keywords with the most character numbers in the set of target word sequences are determined, so as to obtain the keywords with each character number in the second word sequence.
[0188] In a possible design, the construction unit 406 is further used for:
[0189] Clause the target text to obtain the set of clauses corresponding to the target text;
[0190] Process the set of clauses corresponding to the target text to obtain the third word sequence;
[0191] Update the keywords and support degrees in the first sample matching relationship of the target subject according to the third word sequence.
[0192] In summary, it can be seen that in the embodiments provided in the present application, when extracting keywords of a text, the first sample matching relationship and the second sample matching relationship of the subject are introduced, and the keywords and the context features that appear together with the keywords are comprehensively considered. The subject to which the target text belongs is determined by analyzing the degree of association between each word and the subject, and the keywords of the target text are determined by matching each word with the keywords of the subject to which the target text belongs. Therefore, it is possible to avoid extracting keywords of a text only from the perspective of word granularity, thereby avoiding the problem of inaccurate extraction of keywords.
[0193] The embodiments of the present application further provide another text processing device, which is deployed on a server. Please refer to Figure 5 , Figure 5 FIG. is a schematic structural diagram of a server provided by an embodiment of the present invention. The server 500 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 522 (for example, one or more processors) and a memory 532, and one or more storage media 530 (for example, one or more mass storage devices) for storing application programs 542 or data 544. Among them, the memory 532 and the storage media 530 may be transient storage or persistent storage. The program stored in the storage media 530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 522 may be configured to communicate with the storage media 530 and execute a series of instruction operations in the storage media 530 on the server 500.
[0194] The server 500 may further include one or more power supplies 526, one or more wired or wireless network interfaces 550, one or more input / output interfaces 558, and / or one or more operating systems 541, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.
[0195] The steps performed by the text processing device in the above embodiments may be based on the Figure 5 shown server structure.
[0196] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computer, it implements the method flow related to the text processing device in any of the above method embodiments. Correspondingly, the computer may be the above text processing device.
[0197] The embodiments of the present application also provide a computer program or a computer program product including a computer program. When the computer program is executed on a certain computer, the computer will implement the method flow related to the text processing device in any of the above method embodiments. Correspondingly, the computer may be the above-mentioned text processing device.
[0198] In the above Figure 2 or the corresponding embodiment of 3, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0199] As for the text processing device and the server disclosed in the present application, multiple servers can form a blockchain, and the server is a node on the blockchain.
[0200] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).
[0201] It should be understood that the processor mentioned in this application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0202] It should also be understood that the number of processors in this application may be one or multiple, and can be specifically adjusted according to the actual application scenario. This is only an exemplary illustration here and is not limited. The number of memories in the embodiments of this application may be one or multiple, and can be specifically adjusted according to the actual application scenario. This is only an exemplary illustration here and is not limited.
[0203] It should also be noted that when the text processing device includes a processor (or processing unit) and a memory, the processor in this application may be integrated with the memory, or the processor and the memory may be connected through an interface, which can be specifically adjusted according to the actual application scenario and is not limited.
[0204] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0205] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other may be through some interfaces, and the indirect couplings or communication connections of the devices or units may be in electrical, mechanical, or other forms.
[0206] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0207] In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0208] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or other devices, etc.) to execute all or part of the steps of the method described in the embodiments of the present application Figure 2 or the method described in Embodiment 3.
[0209] It should be understood that the storage medium or memory mentioned in this application may include volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synch link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0210] It should be noted that the memories described herein are intended to include, but are not limited to, these and any other suitable types of memories.
[0211] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for processing text, characterized in that, Including: Determine a first keyword set and a second keyword set corresponding to each word in the target text. The first keyword set is the keyword set that matches each word in the target text in the first sample matching relationship. The second keyword set is the keyword set that matches each word in the target text in the second sample matching relationship. The first sample matching relationship includes the matching relationship between the keywords of the first text category and the support degree. The second sample matching relationship includes the matching relationship between the keywords of the second text category and the support degree; Determine the number of first samples hit by each keyword in the first keyword set in the sample set associated with the first sample matching relationship, and determine the number of second samples hit by each keyword in the second keyword set in the sample set associated with the second sample matching relationship; Determine the comprehensive support degree of each word according to the first sample number, the second sample number, the number of samples in the sample set associated with the first sample matching relationship, and the number of samples in the sample set associated with the second sample matching relationship. The comprehensive support degree of each word represents the degree of association between each word and any one of the first text category and the second text category; Determine the target subject of the target text from the first text category and the second text category according to the comprehensive support degree of each word; Match each word with the keyword corresponding to the target subject to obtain the keyword of the target text.
2. The method according to claim 1, wherein The determining the comprehensive support degree of each word according to the first sample number, the second sample number, the number of samples in the sample set associated with the first sample matching relationship, and the number of samples in the sample set associated with the second sample matching relationship includes: Determine the first support degree of each word according to the first sample number and the number of samples in the sample set associated with the first sample matching relationship; Determine the second support degree of each word according to the second sample number and the number of samples in the sample set associated with the second sample matching relationship; Determine the comprehensive support degree of each word according to the first support degree and the second support degree.
3. The method according to claim 1, characterized in that The determining the target subject of the target text from the first text category and the second text category according to the comprehensive support degree of each word includes: Determine the first target words in each word whose comprehensive support degree is greater than the comprehensive support degree threshold; Determine the subject corresponding to the first target word in the first text category and the second text category as the target subject; Or, Determine the second target word with the largest comprehensive support degree in each word; Determine the subject corresponding to the second target word in the first text category and the second text category as the target subject.
4. The method according to any one of claims 1 to 3, characterized in that, The construction of the first sample matching relationship and the second sample matching relationship includes: Obtain a training text set, which includes the training texts associated with the first text category and the training texts associated with the second text category; Perform sentence segmentation on each text in the training text set to obtain the sentence segmentation set corresponding to each text; Process the sentence segmentation set corresponding to each text to obtain the first character sequence corresponding to each text; Remove the keywords with a support threshold less than the support threshold in the first character sequence to obtain the second character sequence corresponding to each text; Determine the keywords in the second character sequence and the support corresponding to the keywords; Determine the keywords in the second character sequence corresponding to the first text category and the support corresponding to the keywords as the first sample matching relationship of the first text category; Determine the keywords in the second character sequence corresponding to the second text category and the support corresponding to the keywords as the second sample matching relationship of the second text category.
5. The method according to claim 4, wherein The processing the sentence segmentation set corresponding to each text to obtain the first character sequence corresponding to each text includes: Filter at least one of punctuation marks, letters, and numbers in the sentence segmentation set corresponding to each text; Split the sentence segmentation set corresponding to each text after filtering by character units to obtain the first character sequence.
6. The method according to claim 4, characterized in that, The determining the keywords in the second character sequence includes: Determine the target keyword with the character number i in the second character sequence and the keyword set associated with the target keyword in the target character sequence set, where the target character sequence set is the sample sentence segmentation set corresponding to at least one subject in the first text category and the second text category, and the value of i is the numerical value of the character number in the second character sequence; Remove the keywords with a support threshold less than the support threshold in the keyword set; Perform recursion based on the character number of the keywords in the second character sequence until the keyword with the most character numbers in the second character sequence and the keyword set corresponding to the keyword with the most character numbers in the target character sequence set are determined, so as to obtain the keywords with different character numbers in the second character sequence.
7. The method according to any one of claims 1-3 and 5-6, characterized in that The method further includes: Perform sentence segmentation on the target text to obtain the sentence segmentation set corresponding to the target text; Process the sentence segmentation set corresponding to the target text to obtain a third character sequence; Update the keywords and supports in the first sample matching relationship of the target subject according to the third character sequence.
8. A text processing device, characterized in that, including: A first determination unit, configured to determine a first keyword set and a second keyword set corresponding to each word in the target text, where the first keyword set is the keyword set that matches each word in the target text in the first sample matching relationship, and the second keyword set is the keyword set that matches each word in the target text in the second sample matching relationship, the first sample matching relationship includes the matching relationship between the keywords and supports of the first text category, and the second sample matching relationship includes the matching relationship between the keywords and supports of the second text category; A second determination unit, configured to determine the number of first samples hit by each keyword in the first keyword set in the sample set associated with the first sample matching relationship, and determine the number of second samples hit by each keyword in the second keyword set in the sample set associated with the second sample matching relationship; A third determination unit, configured to determine the comprehensive support degree of each word according to the first sample number, the second sample number, the number of samples in the sample set associated with the first sample matching relationship, and the number of samples in the sample set associated with the second sample matching relationship, where the comprehensive support degree of each word represents the degree of association between each word and any one of the main bodies in the first text category and the second text category; A fourth determination unit, configured to determine the target main body of the target text from the first text category and the second text category according to the comprehensive support degree of each word; A matching unit, configured to match each word with the keyword corresponding to the target main body to obtain the keyword of the target text.
9. The device according to claim 8, characterized in that, The third determination unit is specifically configured to: Determine the first support degree of each word according to the first sample number and the number of samples in the sample set associated with the first sample matching relationship; Determine the second support degree of each word according to the second sample number and the number of samples in the sample set associated with the second sample matching relationship; Determine the comprehensive support degree of each word according to the first support degree and the second support degree.
10. The device according to claim 8, wherein The fourth determination unit is specifically configured to: Determine a first target word whose comprehensive support degree is greater than a comprehensive support degree threshold among each word; Determine the main body corresponding to the first target word in the first text category and the second text category as the target main body; Or, Determine a second target word with the largest comprehensive support degree among each word; Determine the main body corresponding to the second target word in the first text category and the second text category as the target main body.
11. The device according to any one of claims 8 to 10, characterized in that The apparatus further includes a construction unit; the construction unit is configured to: Obtain a training text set, where the training text set includes training texts associated with the first text category and training texts associated with the second text category; Perform sentence splitting on each text in the training text set to obtain a sentence splitting set corresponding to each text; Process the sentence splitting set corresponding to each text to obtain a first character sequence corresponding to each text; Remove keywords with a support degree less than a support degree threshold in the first character sequence to obtain a second character sequence corresponding to each text; Determine the keywords in the second character sequence and the support degree corresponding to the keywords; Determine the keywords in the second character sequence corresponding to the first text category and the support degree corresponding to the keywords as the first sample matching relationship of the first text category; Determine the keywords in the second character sequence corresponding to the second text category and the support degree corresponding to the keywords as the second sample matching relationship of the second text category.
12. The device according to claim 11, characterized in that, The building unit processes the set of clauses corresponding to each text to obtain the first character sequence corresponding to each text, including: Filtering at least one of punctuation marks, letters, and numbers in the set of clauses corresponding to each text; Splitting the filtered set of clauses corresponding to each text into word units to obtain the first character sequence.
13. The device according to claim 11, characterized in that, The building unit determines the keywords of the second character sequence, including: Determining a target keyword with the number of characters i in the second character sequence and a set of keywords associated with the target keyword in the target character sequence set, where the target character sequence set is a set of sample clauses corresponding to at least one subject in the first text category and the second text category, and the value of i is the numerical value of the number of characters in the second character sequence; Removing keywords in the set of keywords that are less than the support threshold; Recursively based on the number of characters of the keywords in the second character sequence until determining the keyword with the most characters in the second character sequence and the set of keywords corresponding to the keyword with the most characters in the target character sequence set, to obtain the keywords with different numbers of characters in the second character sequence.
14. The device according to any one of claims 12-13, characterized in that, The building unit is further configured to: Clause the target text to obtain a set of clauses corresponding to the target text; Process the set of clauses corresponding to the target text to obtain a third character sequence; Update the keywords and support degrees in the first sample matching relationship of the target subject according to the third character sequence.
15. A computer device, characterized in that, Including: A memory, a processor, and a bus system; Wherein, the memory is used to store programs, and the bus system is used to connect the memory and the processor to enable communication between the memory and the processor; The processor is used to execute the programs in the memory, and the processor is used to execute the text processing method according to any one of claims 1 to 7 according to the instructions in the program code.
16. A computer storage medium, characterized in that, It includes instructions that, when running on a computer, cause the computer to execute the text processing method according to any one of claims 1 - 7.
17. A computer program product, characterized in that, Including a computer program that, when executed on a computer, causes the computer to execute the text processing method according to any one of claims 1 - 7.
Citation Information
Patent Citations
Text classification method, apparatus and device and storage medium
CN110597988A
Event type identification method and device
CN111767730A