A method, apparatus, and readable storage medium for processing text

By using the positive sample matching relationship and negative sample matching relationship to calculate text association scores in electronic document classification, the problem of inaccurate official document distribution in the prior art is solved, and higher distribution accuracy is achieved.

CN113821594BActive Publication Date: 2025-05-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110796094.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-14
Publication Date
2025-05-27
Estimated Expiration
2041-07-14

AI Technical Summary

Technical Problem

The existing template-based electronic document classification method relies on manual formulation of rules and templates, resulting in high time and labor costs, poor generalization ability of rules, and inability to accurately distribute official documents.

Method used

Each word in the target text is matched based on the positive sample matching relationship and the negative sample matching relationship of the first subject. If the match fails, the mutual information of each word and the first keyword and the mutual information of the second keyword are calculated, and the correlation score of the target text is determined based on this information, and then whether to distribute the text to the first subject is determined.

Benefits of technology

There is no need to manually formulate rules and templates, which improves the accuracy of official document text distribution and avoids the problem of inaccurate distribution caused by insufficient generalization ability of manual rules and templates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113821594B_ABST
    Figure CN113821594B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method for processing text and related devices, which can improve the accuracy of text distribution. The method includes: matching each word in the target text based on the positive sample matching relationship and negative sample matching relationship of the first entity, where the positive sample matching relationship includes the matching relationship between the keyword and the support degree of the first entity, and the negative sample matching relationship includes the matching relationship between the keyword and the support degree of the second entity; if the matching fails, determining the mutual information between each word and the first keyword, and determining the mutual information between each word and the second keyword; determining the association score of the target text according to the mutual information between each word and the first keyword, the support degree of the first keyword, the mutual information between each word and the second keyword, and the support degree of the second keyword; if the association score of the target text meets the text distribution condition, distributing the target text to the first entity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital government affairs, and particularly to a method, device, and readable storage medium for processing text. Background Art

[0002] In the field of digital government affairs, automatic document distribution is an essential way to achieve the digital transformation of government affairs and online handling of people's livelihood services. The proposal of new governance and service concepts such as "one network for all people's livelihood services" and "one network for all government services" has accelerated the digital development of the government. Among them, a large amount of government affairs data generated by people's livelihood services and social governance, such as data on the handling of people's livelihood matters, official document texts, and digital services, all need to be better mined and analyzed to truly achieve and accelerate the intelligence of the government affairs industry and improve the convenience of people and government staff in handling matters.

[0003] Currently, the method for establishing an electronic official document classification and grading system is mainly based on a template-based electronic official document classification method. This method constructs a corresponding sensitive word library and matching rules for document assignment tags, learns based on the input sensitive words and imported source files, and generates a source file learning module for the template. The text is subjected to sensitive word matching and rule recognition according to the exported template to obtain document classification and thus achieve automatic text distribution.

[0004] However, the template-based electronic official document classification method overly relies on manually given rules and templates, consuming a large amount of time and labor costs in constructing the sensitive word library and matching rules. At the same time, due to the limitations of the rules and the free format of official document texts, the constructed rules often have a reduced generalization ability and insufficient generality after a certain period of time, resulting in many official documents being unable to be accurately distributed. Summary of the Invention

[0005] This application provides a method, device, and readable storage medium for processing text, which improves the accuracy of official document text distribution.

[0006] One aspect of the embodiments of this application provides a method for processing text, including:

[0007] Matching each word in the target text based on the positive sample matching relationship and negative sample matching relationship of the first subject, where the positive sample matching relationship includes the matching relationship between the keywords of the first subject and the support degree, and the negative sample matching relationship includes the matching relationship between the keywords of the second subject and the support degree;

[0008] If the matching fails, determine the mutual information between each word and the first keyword, and determine the mutual information between each word and the second keyword, where the first keyword is the keyword with the most characters among the keywords of the first subject, and the second keyword is the keyword with the most characters among the keywords of the second subject;

[0009] Determine the association score of the target text according to the mutual information between each word and the first keyword, the support degree of the first keyword, the mutual information between each word and the second keyword, and the support degree of the second keyword, where the association score represents the degree of association between the target text and the first entity;

[0010] If the association score of the target text meets the text distribution condition, then distribute the target text to the first entity.

[0011] In a second aspect of the embodiments of the present application, a text processing device is provided, including:

[0012] A matching unit, configured to match each word in the target text based on the positive sample matching relationship and the negative sample matching relationship of the first entity, where the positive sample matching relationship includes the matching relationship between the keyword and the support degree of the first entity, and the negative sample matching relationship includes the matching relationship between the keyword and the support degree of the second entity;

[0013] A first determination unit, configured to determine the mutual information between each word and the first keyword and the mutual information between each word and the second keyword if the matching fails, where the first keyword is the keyword with the most characters among the keywords of the first entity, and the second keyword is the keyword with the most characters among the keywords of the second entity;

[0014] A second determination unit, configured to determine the association score of the target text according to the mutual information between each word and the first keyword, the support degree of the first keyword, the mutual information between each word and the second keyword, and the support degree of the second keyword, where the association score represents the degree of association between the target text and the first entity;

[0015] A distribution unit, configured to distribute the target text to the first entity if the association score of the target text meets the text distribution condition.

[0016] In a possible design, the second determination unit is specifically configured to:

[0017] Determine the first association score of each word according to the mutual information between each word and the first keyword and the support degree of the first keyword;

[0018] Determine the second association score of each word according to the mutual information between each word and the second keyword and the support degree of the second keyword;

[0019] Determine the association score of the target text according to the first association score of each word and the second association score of each word.

[0020] In a possible design, the first determination unit is further configured to:

[0021] If the match is successful, determine the first keyword set that matches each word in the positive sample matching relationship, and determine the second keyword set that matches each word in the negative sample matching relationship;

[0022] Determine that each first keyword in the first keyword set hits the first target keyword with the most characters in the positive sample matching relationship, and determine that each second keyword in the second keyword set hits the second target keyword with the most characters in the negative sample matching relationship;

[0023] Determine the number of first sample clauses hit by the first target keyword in the sample clause set associated with the positive sample matching relationship, determine the number of second sample clauses hit by the second target keyword in the sample clause set associated with the negative sample matching relationship, and the target number of all sample clauses in the sample clause set associated with the positive sample matching relationship;

[0024] Determine the support weight of the target text according to the number of first sample clauses, the number of second sample clauses, and the target number;

[0025] Distribute the target text according to the support weight of the target text.

[0026] In a possible design, the first determination unit determines the support weight of the target text according to the number of first sample clauses, the number of second sample clauses, and the target number, including:

[0027] Determine the positive support weight of the target text according to the number of first sample clauses and the target number;

[0028] Determine the negative support weight of the target text according to the number of second sample clauses and the target number;

[0029] Determine the support weight of the target text according to the positive support weight of the target text and the negative support weight of the target text.

[0030] In a possible design, the device further includes:

[0031] A third determination unit, and the third determination unit is used for:

[0032] Obtain a training text set, where the training text set includes the training texts associated with the first subject and the training texts associated with the second subject;

[0033] Segment each text in the training text set to obtain a clause set corresponding to each text;

[0034] Process the clause set corresponding to each text to obtain a first character sequence corresponding to each text;

[0035] Remove the keywords in the first word sequence that are less than the support threshold to obtain the second word sequence corresponding to each text;

[0036] Determine the keywords in the second word sequence and the support corresponding to the keywords;

[0037] Determine the keywords in the second word sequence corresponding to the first subject and the support corresponding to the keywords as the positive sample matching relationship of the first subject;

[0038] Determine the keywords in the second word sequence corresponding to the second subject and the support corresponding to the keywords as the negative sample matching relationship of the first subject.

[0039] In a possible design, the third determination unit processes the set of clauses corresponding to each text to obtain the first word sequence corresponding to each text, including:

[0040] Perform stop word filtering on the set of clauses corresponding to each text based on a preset stop word library;

[0041] Perform named entity recognition on the set of clauses corresponding to each text after stop word filtering;

[0042] Split the set of clauses corresponding to each text after named entity recognition into word units to obtain the first word sequence.

[0043] In a possible design, the third determination unit determines the keywords in the second word sequence, including:

[0044] Determine the third keyword with the character count of i in the second word sequence and the set of keywords associated with the third keyword in the target word unit set, where the target word unit set is the set of word units corresponding to at least one of the first subject and the second subject, and the value of i is the numerical value of the character count in the second word sequence;

[0045] Remove the keywords in the set of keywords that are less than the support threshold;

[0046] Recursively based on the character count of the keywords in the second word sequence until the keyword with the most character counts in the second word sequence and the set of keywords corresponding to the keyword with the most character counts in the target word unit set are determined, so as to obtain the keywords with various character counts in the second word sequence.

[0047] Another aspect of the embodiments of the present application provides a computer device, which includes at least one connected processor, a memory, and a transceiver. Among them, the memory is used to store program code, and the processor is used to call the program code in the memory to execute the steps of the text processing method described in the above aspects.

[0048] Another aspect of the embodiments of the present application provides a computer storage medium, which includes instructions that, when running on a computer, cause the computer to execute the steps of the text processing method described in the above aspects.

[0049] In summary, it can be seen that in the embodiments of the present application, the positive sample matching relationship and the negative sample matching relationship of the first subject are matched with each word of the target text to be processed. If the match fails, the mutual information between each word in the target text and the keyword with the largest number of characters in the keywords of the first subject, and the mutual information between each word in the target text and the keyword with the largest number of characters in the keywords of the second subject are calculated. And based on the mutual information, the support degree of the keyword with the largest number of characters in the keywords of the first subject, and the support degree of the keyword with the largest number of characters in the keywords of the second subject, the association score of the target text is determined. This association score indicates the degree of association between the target text and the first subject. If the association score of the target text meets the text distribution condition, the target text is distributed to the first subject. Thus, it can be seen that in the present application, when performing text distribution, there is no need to manually formulate rules and templates. Instead, by introducing the positive sample matching relationship and the negative sample matching relationship of the subject, the association score of the target text is calculated through the mutual information and support degree between each word and the keyword with the largest number of characters, and the target text is distributed based on the association score. This can avoid the problem that the official document text cannot be accurately distributed due to the insufficient generalization ability of manually formulating rules and templates. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 The network architecture diagram of the text distribution system provided by the embodiments of the present application:

[0051] Figure 2 The schematic diagram of an embodiment of the text processing method provided by the embodiments of the present application;

[0052] Figure 3 The schematic diagram of another embodiment of the text processing method provided by the embodiments of the present application;

[0053] Figure 4 The schematic diagram of the virtual structure of the text processing device provided by the embodiments of the present application;

[0054] Figure 5 The schematic diagram of the hardware structure of the server provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0056] In the description, claims and the above-mentioned drawings of this application, terms such as "first" and "second" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or modules does not necessarily have to be limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products or devices. The division of modules in this application is only a logical division, and there may be other division methods when implemented in actual applications. For example, multiple modules can be combined or integrated into another system, or some feature vectors can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, and the indirect couplings or communication connections between modules can be electrical or other similar forms, which are not limited in this application. And the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed to multiple circuit modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application.

[0057] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the architecture of the text distribution system in the embodiment of this application. As Figure 1 shown, it includes user 101, client 102, network 103, server 104 and K databases 105. Among them, each of the K databases stores a positive sample matching relationship corresponding to a subject. The positive sample matching relationship of a certain subject is the negative sample matching relationship of other subjects. The positive sample matching relationship of a certain subject includes the matching relationship between the keywords of the subject and the support degree.

[0058] User 101 inputs the target text through the client 102. The client 102 uploads the target text input by User 101 to the server 104 through the network 103. After obtaining the target text, the server 104 can process the target text to obtain each word segment of the target text, and match each word segment of the processed target text with the positive sample matching relationship and the negative sample matching relationship of the first subject. If the match fails, the mutual information between each word in the target text and the keyword with the most characters in the keywords of the first subject, and the mutual information between the keyword with the most characters in the keywords of the second subject are calculated, and the association score of the target text is determined based on the two mutual informations, the support degree of the keyword with the most characters in the keywords of the first subject, and the support degree of the keyword with the most characters in the keywords of the second subject. This association score identifies the degree of association between the target text and the first subject. If the association score of the target text meets the text distribution condition, the target text is distributed to the database 105 corresponding to the first subject. When the match is successful, the first keyword set that matches each word in the positive sample matching relationship is determined, and the second keyword set that matches each word in the negative sample matching relationship is determined; and it is determined that each first keyword in the first keyword set hits the first target keyword with the most characters in the positive sample matching relationship, and each second keyword in the second keyword set hits the second target keyword with the most characters in the negative sample matching relationship; the number of first sample clauses hit by the first target keyword in the sample clause set associated with the positive sample matching relationship is determined, and the number of second sample clauses hit by the second target keyword in the sample clause set associated with the negative sample matching relationship and the target number of all sample clauses in the sample clause set associated with the positive sample matching relationship are determined; the support weight of the target text is determined according to the number of first sample clauses, the number of second sample clauses, and the target number; finally, the target text is distributed to the database 105 of the corresponding subject according to the support weight of the target text.

[0059] In summary, it can be seen that in this application, when distributing text, there is no need to manually formulate rules and templates. Instead, the positive sample matching relationship and negative sample matching relationship of the subject are introduced. The association score of the target text is calculated through the mutual information between each word and the keyword with the most characters in the positive sample matching relationship, the mutual information between each word and the keyword with the most characters in the negative sample matching relationship, and the support degree. Based on the association score, the target text is distributed. At the same time, when the keyword matching in the positive sample matching relationship and negative sample matching relationship of each word is successful, the positive support degree weight of the keyword in the positive sample matching relationship of each word and the negative support degree weight of the keyword in the negative sample matching relationship of each word are calculated. The support degree weight of the target text is determined according to the positive support degree weight and the negative support degree weight. Finally, the target text is distributed according to the support degree weight, which can avoid the problem that the official document text cannot be accurately distributed due to the insufficient generalization ability of manually formulating rules and templates.

[0060] The following combines Figure 2 the text processing method of this application for illustration. Please refer to Figure 2 , Figure 2 which is a schematic diagram of an embodiment of the text processing method provided by the embodiment of this application, including:

[0061] 201. Build a document distribution database to obtain a training set of distribution departments.

[0062] In this embodiment, a document distribution database can be built and multiple training sets of distribution departments can be obtained. The distribution departments can be, for example, departments such as "City Exclusive Institution", "City Education Bureau", and "City Urban Management Committee". The training set is the official document text in the distribution department. For example, the official document text in the "City Exclusive Institution" can be "I reported the first non-normal means case to the exclusive institution in a certain district and hope it can be filed as soon as possible."

[0063] 202. Mine the positive pattern features and negative pattern features of the distribution department through sequence pattern mining.

[0064] In this embodiment, the positive pattern features and negative pattern features of the distribution department can be mined through text sequence pattern mining. The negative pattern of a certain distribution department is the positive pattern of other distribution departments. Based on the training positive samples and training negative samples of the official document texts of each distribution department, the positive sequence pattern features and negative sequence pattern features of the official documents of the distribution department are respectively mined based on the frequent word sequence pattern.

[0065] 203. Match the positive pattern features and negative pattern features of the document to be predicted to obtain the support degree weight.

[0066] In this embodiment, after mining the positive pattern features and negative pattern features of each distribution department, the official document to be predicted can be matched with the positive pattern features and negative pattern features of each distribution department. If the match is successful, the support weight corresponding to the official document to be predicted is determined.

[0067] 204. Calculate the mutual information between the official document to be predicted and the positive pattern features, and the mutual information between the official document to be predicted and the negative pattern features.

[0068] In this embodiment, if no feature corresponding to the official document to be predicted is matched in the positive pattern features and negative pattern features of each distribution department, calculate the mutual information between each word in the official document to be predicted and the positive pattern features of each distribution department, and calculate the mutual information between each word in the official document to be predicted and the negative pattern features of each distribution department.

[0069] 205. Distribute the official document to be predicted according to the support degree and mutual information.

[0070] In this embodiment, after obtaining the mutual information between each word in the official document to be predicted and the positive pattern features of each distribution department, and the mutual information between each word in the official document to be predicted and the negative pattern features of each distribution department, the official document to be predicted can be distributed according to the mutual information between each word in the official document to be predicted and the positive pattern features of each distribution department, the mutual information between each word in the official document to be predicted and the negative pattern features of each distribution department, the support degree of the positive pattern features, and the support degree of the negative pattern features.

[0071] In summary, it can be seen that in the embodiment provided by the present application, by constructing the positive pattern features and negative pattern features of each distribution department and matching them with the text of the official document to be predicted, if the match is successful, the support weight is obtained, and the text of the official document to be predicted is distributed according to the support weight. Further, when the match fails, the mutual information weight between each word in the official document to be predicted and the positive pattern features of each distribution department, and the mutual information weight between each word in the official document to be predicted and the negative pattern features of each distribution department are calculated, and the official document to be predicted is distributed according to the mutual information weight between each word in the official document to be predicted and the positive pattern features of each distribution department and the mutual information weight between each word in the official document to be predicted and the negative pattern features of each distribution department. Therefore, in the embodiment provided by the present application, when performing text distribution, there is no need to manually formulate rules and templates, thereby avoiding the inaccurate distribution of official document text caused by the insufficient generalization ability of manually formulating rules and templates.

[0072] Combined with the above introduction, the method for processing text in the present application will be introduced from the perspective of a text processing device below. The text processing device can be a server or a service unit in the server.

[0073] Please refer to Figure 3 , Figure 3 , which is a schematic diagram of an embodiment of the text processing method provided by the embodiment of the present application, including:

[0074] 301. Match each word in the target text based on the positive sample matching relationship and negative sample matching relationship of the first subject.

[0075] In this embodiment, the text processing device can obtain the target text, which is an official document text to be distributed. For example, the target text can be "Construction will be carried out at the intersection of XX Avenue and XX Road at 8 o'clock in the morning on holidays, disturbing the residents with noise". After that, the text processing device can process the official document text to be distributed. The processing here includes, but is not limited to, clause segmentation of the target text according to punctuation, stop word filtering, and named entity recognition. Finally, a set of words corresponding to the target text is obtained (it can be understood that, in order to simplify the matching, a threshold can be set, for example, 3, and words that appear less than 3 times in the target text are removed. Of course, it can also be set according to the actual situation). Then, each word in the target text is matched based on the positive sample matching relationship and negative sample matching relationship of the first subject constructed in advance. Among them, the positive sample matching relationship includes the matching relationship between the keywords of the first subject and the support degree, and the negative sample matching relationship includes the matching relationship between the keywords of the second subject and the support degree.

[0076] The construction of the positive sample matching relationship and negative sample matching relationship of the first subject will be described below:

[0077] Step 1. Obtain a training text set, which includes training samples associated with the first subject and training samples associated with the second subject.

[0078] In this step, the text processing device can obtain the training samples associated with the first subject and the training samples associated with the second subject. The first subject can be, for example, "municipal exclusive agency", and the second subject is all other subjects except "municipal exclusive agency", such as "municipal education bureau", "municipal science and technology bureau", and "municipal urban management committee" and other subjects. That is, the text processing device can obtain the historical approved official document text data corresponding to each subject as the training text set for the positive sample matching relationship and negative sample matching relationship of the first subject. Among them, N relevant case texts (i.e., official document texts) are collected for each subject as training positive samples, and the subjects are marked with category identifiers (Identity document, ID), such as: 0, 1, 2...N, to construct the official document text data as shown in Table 1 below:

[0079] Table 1

[0080]

[0081]

[0082] The following is an example of official document text and the main body. Please refer to Table 2:

[0083] Table 2

[0084]

[0085]

[0086] After determining the positive training samples of the official document text for each main body, the official document texts of other main bodies are negative samples for this main body. For example, the official document text of the main body "City Exclusive Institution" is the positive sample of "City Exclusive Institution", and the official document text of "City Education Bureau" is the negative sample of the main body "City Exclusive Institution".

[0087] Step 2: Segment each text in the training text set to obtain the sentence set corresponding to each text.

[0088] In this embodiment, after obtaining the training text set, the text processing device can segment each text in the training text set. The same official document text is an identification object. However, to extract the keyword features in the official document text, the identification object needs to be segmented into sentences. The keywords extracted from each sentence constitute the pattern feature set of the entire official document text. Here, the sentences are segmented by regular expression matching punctuation marks in the official document text to obtain the sentence set corresponding to each official document text. For example, for the official document text "I reported the first non-normal means case to the exclusive institution in a certain district and hoped that it could be filed as soon as possible", the sentences obtained after segmentation are "I reported the first non-normal means case to the exclusive institution in a certain district" and "Hoped that it could be filed as soon as possible", these two sentences.

[0089] Step 3: Process the sentence set corresponding to each text to obtain the first character sequence corresponding to each text.

[0090] In this embodiment, after the text processing device obtains the clause set corresponding to each text, it can process the clauses corresponding to each text to obtain the first character sequence elements corresponding to each text. Specifically, the text processing device can perform stop word filtering on the clause set corresponding to each text based on a preset stop word library. This stop word filtering includes, but is not limited to, filtering out useless information such as "date and time, name, email, mobile phone number", etc.; performing named entity recognition on the clause set corresponding to each text after stop word filtering; for the distribution of official document texts, the entity names of organizations that appear in official document texts, such as exclusive organizations and exclusive organizations, etc., often play a great role in the automatic classification and distribution of official documents. Therefore, here it is necessary to perform named entity recognition on the organization names based on a named entity recognition (NER) tool to obtain the organization name entities in the official document text, and identify the same organization name entity with a unified code. For example, the exclusive organization is identified as #. The purpose of doing this is to ensure that the organization name entity words will not be split when constructing the positive sample matching relationship and the negative sample matching relationship. For the convenience of understanding, the stop word filtering and named entity recognition of official document texts will be described below with specific examples. Tables 3 and 4 are the positive samples of the official document text of the main body "City Exclusive Organization" after stop word filtering and named entity recognition:

[0091] Table 3

[0092]

[0093] Table 4

[0094]

[0095] After performing stop word filtering and named entity recognition on the official document text, the clause set corresponding to each text can be split by character units to obtain the first character sequence (it can be understood that here, if it is a Chinese text, the character unit corresponds to a single character, and if it is an English text, it corresponds to a single word, which is not specifically limited). For example, the clauses "I reported to a certain district # about the first case of extortion by non-normal means" and "Hope to file a case as soon as possible" in Sample 1 of Table 4 are split by character units to obtain: "I, to, a, certain, district, #, report, to, the, police, reflect, extortion, case, hope, can, as, soon, as, possible, file, a, case". This "I, to, a, certain, district, #, report, to, the, police, reflect, extortion, case, hope, can, as, soon, as, possible, file, a, case" is the first character sequence element corresponding to Sample 1.

[0096] Step 4: Remove the keywords in the first character sequence that are less than the support threshold to obtain the second character sequence corresponding to each text.

[0097] In this embodiment, the text processing device can count the number of sample clauses in which all word sequences appear in the clauses of each sample object, and filter out keywords with a support threshold less than the support threshold. Assume that the support threshold is 1 / 3, that is, at least 2 sample clauses must appear in these 6 sample clauses to meet the support threshold, otherwise the keyword is filtered out. Taking the positive sample of the training official document "municipal exclusive agency" in Table 3 as an example, the keywords with a support threshold less than the support threshold in the first word sequence are removed, and the second word sequence corresponding to each text is obtained. The results are shown in Table 5, including each keyword and the number of sample clauses hit by each keyword:

[0098] Table 5

[0099] word case file report check verify document police fraud # number of sample clauses 5 3 2 2 2 2 2 2 2

[0100] It should be noted that in this application, word sequences are used as the objects for keyword mining, and keywords with various character counts that meet the support threshold are mined from official document texts based on the Prefixspan algorithm. The calculation of this support threshold can be carried out through the following formula:

[0101] min_sup = a × n

[0102] Among them, min_sup is the support threshold, n is the number of official document texts in a certain subject, a is the minimum support rate, and the minimum support rate is adjusted according to the magnitude of the training data set. At the same time, this application uses a "snowballing" method, that is, a relatively high support threshold is set for each round of mining to ensure the accuracy of keyword mining, and the recall rate is improved through multiple rounds of iterative mining.

[0103] Step 5: Determine the keywords in the second word sequence and the support corresponding to the keywords.

[0104] In this step, the text processing device can determine the keywords in the second word sequence elements and the support corresponding to the keywords. The following is a specific description:

[0105] Step 51: Determine the third keyword with the character count of i in the second word sequence elements and the keyword set associated with the third keyword in the target word unit set. The target word unit set is the word unit set corresponding to the first subject or the second subject. Among them, the value of i is the numerical value of the character count in the second word sequence. Taking Table 5 as an example, taking the value of the character count i of the second word sequence starting from 1 and recursively adding 1 each time as the benchmark, determine the third keyword with the character count of i and the keyword set associated with the third keyword in the target word unit set. Table 6 shows the case where i is 1:

[0106] Table 6

[0107]

[0108]

[0109] Among them, the set of target word units is shown in Table 7 as follows:

[0110] Table 7

[0111] Sample 1 # Report fraud cases file a case Sample 2 case file and verify a case Sample 3 fraud # Report, verify, file and check

[0112] Taking the "#" with a character count of 1 as an example for illustration, since this "#" appears in both Sample 1 and Sample 3, the set of associated keywords in the set of target word units is "alarm fraud cases" and "alarm nuclear case filing and investigation".

[0113] Step 52: Eliminate the keywords in the set of keywords that are less than the support threshold.

[0114] In this step, the text processing device can eliminate the keywords in the set of keywords that are less than the support threshold. Suppose the support threshold is 1 / 3 (that is, the keyword must appear in 1 out of 3 sentences). Taking the third keyword as "#" for example, the set of keywords associated with this keyword "#" in the set of target word units, and the keywords in this set of keywords meet the support threshold. As shown in Table 8, the words "fraud", "case", "nuclear", "filing", and "investigation" in Table 6 do not reach the support threshold and are eliminated, obtaining the keywords and the number of sample clauses as shown in Table 8:

[0115] Table 8

[0116] Keywords in the keyword set report police case number of sample clauses 2 2 2

[0117] Step 53: Recursively process based on the character count of the keywords in the second word sequence until the keyword with the most characters in the second word sequence and the set of keywords corresponding to the keyword with the most characters in the set of target word units are determined, obtaining the keywords of the second word sequence.

[0118] In this step, the text processing device recursively processes based on the character count of the keywords in the second word sequence until the keyword with the most characters in the second word sequence and the set of keywords corresponding to the keyword with the most characters in the set of target word units are determined, obtaining the keywords of the second word sequence.

[0119] Taking "#" as an example for secondary recursion to obtain the third keyword with 2 characters, the following illustrates how to determine the third keyword with 2 characters and the set of keywords associated with the third keyword in the set of target word units. Specifically, as shown in Table 9:

[0120] Table 9

[0121]

[0122] Among them, the keywords associated with "# Report" in the target word unit set are "Police Fraud Cases" and "Police Verification and Filing for Investigation" respectively. That is, by matching "# Report" with the target word unit set in Table 7, the corresponding keyword set can be obtained.

[0123] Continuing with the example of the third keyword with a character count of 3 obtained after three recursions with "#", this paper illustrates how to determine the third keyword with a character count of 3 and the keyword set associated with the third keyword in the target word unit set, as shown in Table 10:

[0124] Table 10

[0125]

[0126] Taking the first row in Table 10 as an example, the keywords associated with "# Alarm" in the target word unit set are "Fraud Cases" and "Verification and Filing for Investigation" respectively. That is, by matching "# Report" with the target word unit set in Table 7, the corresponding keyword set can be obtained.

[0127] Continuing with the example of the third keyword with a character count of 4 obtained after three recursions with "#", this paper illustrates how to determine the third keyword with a character count of 4 and the keyword set associated with the third keyword in the target word unit set, as shown in Table 11:

[0128] Table 11

[0129]

[0130] Among them, the keywords associated with "# Alarm Case" in the target word unit set are "Case" and "Investigation" respectively. That is, by matching "# Report" with the target word unit set in Table 7, the corresponding keyword set can be obtained. Thus, the recursion of the keywords corresponding to "#" ends. From this, the third keyword corresponding to each character count and the keyword set associated with the third keyword in the target word unit set can be obtained.

[0131] Step 54: Determine the positive sample matching relationship of the first subject as the keywords of the second word sequence corresponding to the first subject and the support degree corresponding to the keywords.

[0132] In this embodiment, after the text processing device obtains the keywords of the second word sequence elements corresponding to each text in the training text set, it can determine the support degree corresponding to the keywords, and determine the positive sample matching relationship of the first subject as the keywords of the second word sequence corresponding to the first subject and the support degree corresponding to the keywords. The following takes the keyword "#" as an example in combination with Table 12. Table 12 shows the positive sample matching relationship of the keyword "#" corresponding to the first subject:

[0133] Table 12

[0134]

[0135]

[0136] Taking "#" as an example, when the number of characters is 1, its corresponding support degree is 1 / 3, that is, "#" appears twice in 6 clauses, and the same is true for "# report". Thus, the keywords in the second character sequence element corresponding to the first subject and the support degree corresponding to the keyword can be obtained.

[0137] Step 52: Determine the negative sample matching relationship of the first subject by the keywords with each number of characters in the second character sequence corresponding to the second subject and the support degree corresponding to the keywords with each number of characters.

[0138] In this step, since the first subject and the second subject are positive and negative sample matching relationships with each other, that is, the positive sample matching relationship of the first subject is the negative sample matching relationship of the second subject, and the positive sample matching relationship of the second subject is the negative sample matching relationship of the first subject. After obtaining the keywords corresponding to each text in the training text set and the support degree corresponding to the keyword, that is, the keywords with each number of characters in the second character sequence element corresponding to the second subject and the support degree corresponding to the keywords with each number of characters are obtained.

[0139] 302. If the matching fails, then determine the mutual information between each word and the first keyword, and determine the mutual information between each word and the second keyword.

[0140] In this embodiment, when the text processing device matches each word in the target text based on the positive sample matching relationship and the negative sample matching relationship of the first subject, if each word in the target text fails to match successfully (that is, when each word in the target text is separately matched with the keywords in the positive sample matching relationship and the negative sample matching relationship, if each word in the target text is not the same as the keywords in the positive sample matching relationship and the negative sample matching relationship, it is considered a matching failure; if there is a word in the target text that is the same as the keywords in the positive sample matching relationship and the negative sample matching relationship, it is considered a matching success. Here, taking the target text as "Recently, mobile phone theft cases have frequently occurred. Please ask the relevant departments to file a case for investigation" and the positive sample matching relationship as the keywords and support degrees in Table 12 as an example. After matching each word in the target text with the positive sample matching relationship, there are two keywords in the positive sample matching relationship that are the same as the words "relevant departments" and "case" in the target text, that is, "# case" in Table 12, then it is determined that the matching is successful. If there are no keywords in the positive sample matching relationship in Table 12 that are the same as each word in the target text, then it is determined that the matching fails), the text processing device can calculate the mutual information between each word segment in the target text and the first keyword, and calculate the mutual information between each word and the second keyword, where the first keyword is the keyword with the most characters among the keywords of the first subject, and the second keyword is the keyword with the most characters among the keywords of the second subject. For example, in Table 12, the first keyword is "# alarm case". The mutual information is described as follows:

[0141] If keyword x and keyword y often appear together, then the mutual information between keyword x and keyword y is relatively large. The formula for the mutual information between keyword x and keyword y can be defined as:

[0142]

[0143] At the same time, when determining the mutual information, the consideration of word frequency is introduced in the calculation, and more attention is paid to high-frequency words.

[0144]

[0145] Among them, x is the word to be mined, y is the high-frequency word that often appears together with x, I(x, y) is the mutual information between x and y, p(x, y) is the distribution probability that x and y appear simultaneously, p(x) is the appearance probability of x alone, p(y) is the appearance probability of y alone, and a ∈ (0.5, 1]. In this application, x is each word in the target text, and y is the first keyword and the second keyword. During the calculation process using the above formula, first, the target text is segmented to obtain each word of the target text, then each word is vectorized and encoded through the word2vec word vector model, and at the same time, the first keyword and the second keyword are vectorized and encoded through the word2vec word vector model. Finally, each vectorized and encoded word, the first keyword, and the second keyword are input into the above formula to calculate the mutual information between each word and the first keyword, and the mutual information between each word and the second keyword respectively. It can be understood that when calculating the mutual information between two word vectors, other methods can also be used, such as calculating through toolkits in tools like python or Matlab, and the specific method is not limited.

[0146] It can be understood that the above uses the word2vec word vector model to vectorize and encode each word, the first keyword, and the second keyword. Of course, other methods can also be used to vectorize the first keyword and the second keyword, and the specific method is not limited.

[0147] 303. Determine the association score of the target text according to the mutual information between each word and the first keyword, the support degree of the first keyword, the mutual information between each word and the second keyword, and the support degree of the second keyword.

[0148] In this embodiment, after the text processing device determines the mutual information between each word and the first key, the support degree of the first keyword, the mutual information between each word and the second keyword, and the support degree of the second keyword, it can determine the association score of the target text according to the mutual information between each word and the first keyword, the support degree of the first keyword, the mutual information between each word and the second keyword, and the support degree of the second keyword. Specifically as follows:

[0149] The text processing device determines the first association score of each word according to the mutual information between each word and the first key and the support degree of the first keyword;

[0150] Determine the second association score of each word according to the mutual information between each word and the second key and the support degree of the second keyword;

[0151] Determine the association score of the target text according to the first association score of each word and the second association score of each word.

[0152] That is, the text processing device can multiply the mutual information between each word and the first keyword by the support degree of the first keyword to obtain the first association score of each word. It can be understood that when there are multiple first keywords, the support degree of each first keyword is multiplied by the mutual information between each word and the first keyword respectively and then added together to obtain the first association score of each word. Similarly, the calculation method of the second association score is also like this. After that, the first association score and the second association score of each word can be obtained. Since the second association score is the association score of the negative sample matching relationship of each word relative to the first subject, it is necessary to change it to a negative number after calculating the second association score of each word, and add the first association score and the second association score of each word to obtain the association score of the target text. It can be understood that for the convenience of calculation, after obtaining the association score of the first word and the association score of the second word, the association score of the first word and the association score of the second word can be standardized, and its range is set to [-1, 1]. The first association score is a positive number, and the second association score is a negative number. The first association score and the second association score of each word are added together to obtain the association score of the target text.

[0153] 304. If the association score of the target text meets the text distribution condition, the target text is distributed to the first subject.

[0154] In this embodiment, after obtaining the association score of the target text, the text processing device can determine whether the association score of the target text meets the text distribution condition, and when the association score of the target text meets the text distribution condition, distribute the target text to the first subject. As described in the above example, the association score of the target text is the probability value of the target text belonging to each subject. The target text is assigned to the subject with the highest association score among each subject, or an association score threshold can be set, and the target text is distributed to the subject whose association score meets the association score threshold among each subject. For example, the association score of the target text relative to the subject "Municipal Exclusive Institution" is 0.8, the association score relative to the subject "Municipal Science and Technology Bureau" is 0.3, and the association score relative to the subject "Municipal Education Bureau" is 0.6. The target text can be directly distributed to the subject "Municipal Exclusive Institution". If the association score threshold is 0.6, the target text can be distributed to the subject "Municipal Exclusive Institution" and the subject "Municipal Education Bureau".

[0155] In one embodiment, if each word in the target text is matched based on the positive sample matching relationship and the negative sample matching relationship of the first subject, and when the matching is successful, the text processing device also performs the following operations:

[0156] Determine the first keyword set that matches each word in the positive sample matching relationship, and determine the second keyword set that matches each word in the negative sample matching relationship;

[0157] Determine the first target keyword with the most characters in the positive sample matching relationship that each first keyword in the first keyword set hits, and determine the second target keyword with the most characters in the negative sample matching relationship that each second keyword in the second keyword set hits;

[0158] Determine the number of first sample clauses hit by the first target keyword in the sample clause set associated with the positive sample matching relationship, determine the number of second sample clauses hit by the second target keyword in the sample clause set associated with the negative sample matching relationship, and the target number of all sample clauses in the sample clause set associated with the positive sample matching relationship;

[0159] Determine the support weight of the target text according to the number of first sample clauses, the number of second sample clauses, and the target number;

[0160] Distribute the target text according to the support weight of the target text.

[0161] In this embodiment, for the target text, each word in the target text is respectively matched with the positive sample matching relationship and the negative sample matching relationship of the first subject. When the matching is successful, the first keyword set that matches each word in the positive sample matching relationship can be determined. Taking the positive sample matching relationship as Table 12 as an example, after matching each word in the target text with the keywords of each character length in Table 12, the first keyword set obtained is "#", "#report", "#alarm", "#case", "#report the alarm", "#report a case", and "#alarm case". Similarly, each word in the target text can be matched with the negative sample matching relationship to obtain the second keyword set;

[0162] After that, the first target keyword with the most characters in the first keyword set and the second target keyword with the most characters in the second keyword set can be determined. Continuing with the above example, the first target keywords are "#report the alarm", "#report a case", and "#alarm case", and determine the number of first sample clauses hit by the first target keyword in the sample clause set associated with the positive sample matching relationship. As shown in Table 7, "report the alarm", "#report a case", and "#alarm case" each hit two sample clauses. Similarly, the number of second sample clauses hit by the second target keyword in the sample clause set associated with the negative sample matching relationship can be determined, and the target number of all sample clauses in the sample clause set associated with the positive correlation relationship can be determined. As shown in Table 7, the target number is 6;

[0163] Finally, determine the support weight of the target text based on the number of first sample clauses, the number of second sample clauses, and the target quantity. Specifically, the positive support weight of the target text can be determined according to the number of first sample clauses and the target quantity, and the negative support weight of the target text can be determined according to the number of second sample clauses and the target quantity. Specifically, the positive support weight and negative support weight of the target text can be calculated through the following formula:

[0164]

[0165] After that, determine the support weight of the target text based on the positive support weight and negative support weight of the target text.

[0166] It should be noted that when there are multiple first target keywords, there will also be multiple corresponding numbers of first sample clauses. The positive support weights calculated by each number of first sample clauses and the target quantity can be added together to obtain the positive support weight of the target text. Similarly, the same is true when there are multiple second target keywords. After calculating the positive support weight and negative support weight of the target text, the positive support weight and negative support weight of the target text can be standardized and set within the range of [-1, 1]. Among them, the negative support weight of the target text is set to a negative number. In this way, the positive support weight of the target text minus the negative support weight of the target text can obtain the support weight of the target text. After that, distribute the target text according to the support weight of the target text. The distribution of the target text has been described in detail in step 304. Here, only the associated score threshold needs to be set as the support weight threshold.

[0167] In summary, it can be seen that in the embodiments provided by the present application, when performing text distribution, there is no need to manually formulate rules and templates, but instead, by introducing the positive sample matching relationship and negative sample matching relationship of the subject, calculate the associated score of the target text through the mutual information and support of each word and the keyword with the most characters, and perform the distribution of the target text based on the associated score, which can avoid the problem that the official document text cannot be accurately distributed due to the insufficient generalization ability of manually formulating rules and templates.

[0168] The above describes the embodiments of the present application from the perspective of the text processing method. The following describes the embodiments of the present application from the perspective of the text processing device.

[0169] Please refer to Figure 4 , in the embodiments of the present application, a text processing device is provided. The text processing device 400 includes:

[0170] A matching unit 401 is configured to match each word in the target text based on the positive sample matching relationship and the negative sample matching relationship of the first subject, where the positive sample matching relationship includes the matching relationship between the keyword of the first subject and the support degree, and the negative sample matching relationship includes the matching relationship between the keyword of the second subject and the support degree;

[0171] A first determination unit 402 is configured to, if the matching fails, determine the mutual information between each word and the first keyword, and determine the mutual information between each word and the second keyword, where the first keyword is the keyword with the most characters among the keywords of the first subject, and the second keyword is the keyword with the most characters among the keywords of the second subject;

[0172] A second determination unit 403 is configured to determine the association score of the target text according to the mutual information between each word and the first keyword, the support degree of the first keyword, the mutual information between each word and the second keyword, and the support degree of the second keyword, where the association score represents the degree of association between the target text and the first subject;

[0173] A distribution unit 404 is configured to, if the association score of the target text meets the text distribution condition, distribute the target text to the first subject.

[0174] In a possible design, the second determination unit 403 is specifically configured to:

[0175] Determine the first association score of each word according to the mutual information between each word and the first keyword and the support degree of the first keyword;

[0176] Determine the second association score of each word according to the mutual information between each word and the second keyword and the support degree of the second keyword;

[0177] Determine the association score of the target text according to the first association score of each word and the second association score of each word.

[0178] In a possible design, the first determination unit 402 is further configured to:

[0179] If the matching is successful, determine the first keyword set that matches each word in the positive sample matching relationship, and determine the second keyword set that matches each word in the negative sample matching relationship;

[0180] Determine that each first keyword in the first keyword set hits the first target keyword with the most characters in the positive sample matching relationship, and determine that each second keyword in the second keyword set hits the second target keyword with the most characters in the negative sample matching relationship;

[0181] Determine the number of the first sample clauses hit by the first target keyword in the set of sample clauses associated with the positive sample matching relationship, and determine the number of the second sample clauses hit by the second target keyword in the set of sample clauses associated with the negative sample matching relationship and the target number of all the sample clauses in the set of sample clauses associated with the positive sample matching relationship;

[0182] Determine the support weight of the target text according to the number of the first sample clauses, the number of the second sample clauses and the target number;

[0183] Distribute the target text according to the support weight of the target text.

[0184] In a possible design, the first determination unit 402 determines the support weight of the target text according to the number of the first sample clauses, the number of the second sample clauses and the target number, including:

[0185] Determine the positive support weight of the target text according to the number of the first sample clauses and the target number;

[0186] Determine the negative support weight of the target text according to the number of the second sample clauses and the target number;

[0187] Determine the support weight of the target text according to the positive support weight of the target text and the negative support weight of the target text.

[0188] In a possible design, the device further includes:

[0189] A third determination unit 405, and the third determination unit 405 is configured to:

[0190] Obtain a training text set, where the training text set includes the training texts associated with the first subject and the training texts associated with the second subject;

[0191] Perform clause segmentation on each text in the training text set to obtain a set of clauses corresponding to each text;

[0192] Process the set of clauses corresponding to each text to obtain a first character sequence corresponding to each text;

[0193] Eliminate the keywords in the first character sequence that are less than the support threshold to obtain a second character sequence corresponding to each text;

[0194] Determine the keywords in the second character sequence and the support corresponding to the keywords;

[0195] Determine the keywords in the second character sequence corresponding to the first subject and the support corresponding to the keywords as the positive sample matching relationship of the first subject;

[0196] Determine the keywords in the second word sequence corresponding to the second subject and the support degree corresponding to the keywords as the negative sample matching relationship of the first subject.

[0197] In a possible design, the third determination unit 405 processes the clause set corresponding to each text to obtain the first word sequence corresponding to each text, including:

[0198] Perform stop word filtering on the clause set corresponding to each text based on a preset stop word library;

[0199] Perform named entity recognition on the clause set corresponding to each text after stop word filtering;

[0200] Split the clause set corresponding to each text after named entity recognition into word units to obtain the first word sequence.

[0201] In a possible design, the third determination unit 405 determines the keywords in the second word sequence, including:

[0202] Determine the third keyword with the character number i in the second word sequence and the keyword set associated with the third keyword in the target word unit set. The target word unit set is the sample clause set corresponding to at least one of the first subject and the second subject. The value of i is the numerical value of the character number in the second word sequence;

[0203] Eliminate the keywords in the keyword set that are less than the support degree threshold;

[0204] Recursively based on the character number of the keywords in the second word sequence until the keyword with the most characters in the second word sequence and the keyword set corresponding to the keyword with the most characters in the target word unit set are determined, so as to obtain the keywords with various character numbers in the second word sequence.

[0205] In summary, it can be seen that in the embodiments provided by the present application, when performing text distribution, there is no need to manually formulate rules and templates. Instead, the positive sample matching relationship and negative sample matching relationship of the introduced subject are adopted. The mutual information and support degree between each word and the keyword with the most characters are used to calculate the association score of the target text, and the target text is distributed based on the association score, which can avoid the problem that the official document text cannot be accurately distributed due to the insufficient generalization ability of manually formulating rules and templates.

[0206] The embodiments of the present application also provide another text processing device, which is deployed on the server. Please refer to Figure 5 , Figure 5FIG. 0 is a schematic structural diagram of a server provided by an embodiment of the present invention. The server 500 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 522 (for example, one or more processors) and a memory 532, and one or more storage media 530 for storing application programs 542 or data 544 (for example, one or more mass storage devices). Among them, the memory 532 and the storage media 530 may be transient storage or persistent storage. The program stored in the storage media 530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 522 may be configured to communicate with the storage media 530 and execute a series of instruction operations in the storage media 530 on the server 500.

[0207] The server 500 may further include one or more power supplies 526, one or more wired or wireless network interfaces 550, one or more input / output interfaces 558, and / or one or more operating systems 541, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0208] In the above embodiments, the steps executed by the text processing device may be based on the Figure 5 shown server structure.

[0209] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computer, it implements the method flow related to the text processing device in any of the above method embodiments. Correspondingly, the computer may be the above text processing device.

[0210] An embodiment of the present application further provides a computer program or a computer program product including the computer program. When the computer program is executed on a certain computer, it will cause the computer to implement the method flow related to the text processing device in any of the above method embodiments. Correspondingly, the computer may be the above text processing device.

[0211] In the above Figure 2 or the embodiments corresponding to 3, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product.

[0212] As the text processing device and the server disclosed in the present application, multiple servers may be grouped into a blockchain, and the server is a node on the blockchain.

[0213] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0214] It should be understood that the processor mentioned in the present application may be a central processing unit (CPU), and may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0215] It should also be understood that the number of processors in the present application may be one or more, which can be specifically adjusted according to the actual application scenario. This is only an exemplary illustration here and is not limited. The number of memories in the embodiments of the present application may be one or more, which can be specifically adjusted according to the actual application scenario. This is only an exemplary illustration here and is not limited.

[0216] It should also be noted that when the text processing device includes a processor (or processing unit) and a memory, the processor in the present application may be integrated with the memory or the processor and the memory may be connected through an interface, which can be specifically adjusted according to the actual application scenario and is not limited.

[0217] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0218] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0219] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0220] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0221] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or other devices, etc.) to execute all or part of the steps of the method described in the embodiments of the present application Figure 2 or all or part of the steps of the method described in Embodiment 3 of the present application.

[0222] It should be understood that the storage medium or memory mentioned in this application may include volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synch link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0223] It should be noted that the memories described herein are intended to include but not be limited to these and any other suitable types of memories.

[0224] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A method for processing text, characterized in that, it includes: Based on the positive sample matching relationship and negative sample matching relationship of the first subject, each word in the target text is matched, wherein the positive sample matching relationship includes the matching relationship between the keyword of the first subject and the support degree, and the negative sample matching relationship includes the matching relationship between the keyword of the second subject and the support degree; If the matching fails, the mutual information between each word and the first keyword is determined, and the mutual information between each word and the second keyword is determined, wherein the first keyword is the keyword with the most characters among the keywords of the first subject, and the second keyword is the keyword with the most characters among the keywords of the second subject; According to the mutual information between each word and the first keyword, the support degree of the first keyword, the mutual information between each word and the second keyword, and the support degree of the second keyword, the association score of the target text is determined, wherein the association score represents the degree of association between the target text and the first subject; If the association score of the target text meets the text distribution condition, the target text is distributed to the first subject.

2. The method according to claim 1, characterized in that, The determining the association score of the target text according to the mutual information between each word and the first keyword, the support degree of the first keyword, the mutual information between each word and the second keyword, and the support degree of the second keyword includes: Determining the first association score of each word according to the mutual information between each word and the first keyword and the support degree of the first keyword; Determining the second association score of each word according to the mutual information between each word and the second keyword and the support degree of the second keyword; Determining the association score of the target text according to the first association score of each word and the second association score of each word.

3. The method according to claim 1, characterized in that, The method further includes: If the matching is successful, determining the first keyword set that matches each word in the positive sample matching relationship, and determining the second keyword set that matches each word in the negative sample matching relationship; Determining that each first keyword in the first keyword set hits the first target keyword with the most characters in the positive sample matching relationship, and determining that each second keyword in the second keyword set hits the second target keyword with the most characters in the negative sample matching relationship; Determining the number of first sample clauses hit by the first target keyword in the sample clause set associated with the positive sample matching relationship, and determining the number of second sample clauses hit by the second target keyword in the sample clause set associated with the negative sample matching relationship and the target number of all sample clauses in the sample clause set associated with the positive sample matching relationship; Determining the support degree weight of the target text according to the number of first sample clauses, the number of second sample clauses, and the target number; Distributing the target text according to the support degree weight of the target text.

4. The method according to claim 3, wherein, the determining the support weight of the target text according to the number of clauses in the first sample, the number of clauses in the second sample, and the target number includes: determining the positive support weight of the target text according to the number of clauses in the first sample and the target number; determining the negative support weight of the target text according to the number of clauses in the second sample and the target number; determining the support weight of the target text according to the positive support weight of the target text and the negative support weight of the target text.

5. The method according to any one of claims 1 to 4, wherein, before matching each word in the target text based on the positive sample matching relationship and the negative sample matching relationship of the first subject, the method further includes: obtaining a training text set, the training text set including the training texts associated with the first subject and the training texts associated with the second subject; clause-segmenting each text in the training text set to obtain a clause set corresponding to each text; processing the clause set corresponding to each text to obtain a first character sequence corresponding to each text; removing keywords in the first character sequence that are less than the support threshold to obtain a second character sequence corresponding to each text; determining the keywords in the second character sequence and the support corresponding to the keywords; determining the keywords in the second character sequence corresponding to the first subject and the support corresponding to the keywords as the positive sample matching relationship of the first subject; determining the keywords in the second character sequence corresponding to the second subject and the support corresponding to the keywords as the negative sample matching relationship of the first subject.

6. The method according to claim 5, wherein, the processing the clause set corresponding to each text to obtain a first character sequence corresponding to each text includes: performing stop word filtering on the clause set corresponding to each text based on a preset stop word library; performing named entity recognition on the clause set corresponding to each text after stop word filtering; splitting the clause set corresponding to each text after named entity recognition into word units to obtain the first character sequence.

7. The method according to claim 5, wherein, the determining the keywords in the second character sequence includes: determining a third keyword with the number of characters being i in the second character sequence and a keyword set associated with the third keyword in a target word unit set, the target word unit set being a word unit set corresponding to at least one of the first subject and the second subject, and the value of i being the numerical value of the number of characters in the second character sequence; removing keywords in the keyword set that are less than the support threshold; Recursion is performed based on the number of characters of the keywords in the second word sequence until the keyword with the most characters in the second word sequence and the keyword set corresponding to the keyword with the most characters in the target word unit set are determined, and keywords with each number of characters in the second word sequence are obtained.

8. A text processing device characterized in that it includes: a matching unit configured to match each word in a target text based on a positive sample matching relationship and a negative sample matching relationship of a first subject, wherein the positive sample matching relationship includes a matching relationship between a keyword of the first subject and a support degree, and the negative sample matching relationship includes a matching relationship between a keyword of a second subject and a support degree; a first determination unit configured to, if the matching fails, determine the mutual information between each word and a first keyword, and determine the mutual information between each word and a second keyword, wherein the first keyword is the keyword with the most characters among the keywords of the first subject, and the second keyword is the keyword with the most characters among the keywords of the second subject; a second determination unit configured to determine an association score of the target text according to the mutual information between each word and the first keyword, the support degree of the first keyword, the mutual information between each word and the second keyword, and the support degree of the second keyword, wherein the association score represents the degree of association between the target text and the first subject; a distribution unit configured to, if the association score of the target text meets a text distribution condition, distribute the target text to the first subject.

9. A computer device characterized in that it includes: a memory, a processor, and a bus system; wherein, the memory is used for storing programs, and the bus system is used for connecting the memory and the processor to enable communication between the memory and the processor; the processor is used for executing the programs in the memory, and the processor is used for performing the text processing method according to the instructions in the program code in any one of claims 1 to 7.

10. A computer storage medium characterized in that it includes instructions which, when running on a computer, cause the computer to perform the text processing method according to any one of claims 1 - 7.

Citation Information

Patent Citations

  • Template-based classification system of electronic official documents

    CN108399164A

  • Pseudo-correlation feedback information retrieval method and system based on BM25+ALBERT model and storage medium

    CN111625624A