A Co-evolution Method for Text Classification and the Growth of Term Networks

By combining term networks and text classifiers in text classification, the interpretability and knowledge accumulation of text classification in subject areas are achieved, the problem of lack of interpretability and knowledge accumulation in the existing technology is solved, and the high-precision and migratory text classification effect is achieved.

CN114416997BActive Publication Date: 2025-05-30JIZHI ACAD (BEIJING) TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210078144.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-24
Publication Date
2025-05-30
Estimated Expiration
2042-01-24

AI Technical Summary

Technical Problem

The existing text classification algorithm lacks interpretability and knowledge accumulation functions in subject field classification, and it is difficult to meet actual needs.

Method used

By combining the term network and the text classifier, the co-evolution of text classification and term network is realized, the contribution of terms to classification results is identified, and the term network in the target field is constructed simultaneously.

Benefits of technology

It realizes the interpretability of text classification, improves the accuracy of text classification, and continuously improves the knowledge base through the growth of term network, which is suitable for multi-field and interdisciplinary fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114416997B_ABST
    Figure CN114416997B_ABST
Patent Text Reader

Abstract

The present invention discloses a co-evolution method for text classification and term network growth. On the one hand, a term subgraph is constructed for the text, and the term subgraph is scored based on the characteristics of the term network, so as to achieve text classification; on the other hand, the term subgraph is extracted from the classified text, and the term subgraph is used to expand and optimize a domain term network. The co-optimization between the two tasks of text classification and term network generation in this method can achieve a text classifier applicable to a certain domain and a growable domain term network based on a small amount of domain text and a large general term network. Further, this method can be used to establish a knowledge graph of a certain domain, realize real demands such as article recommendation in a certain domain, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of natural language processing and knowledge acquisition in knowledge engineering. It mainly relates to the binary classification problem of a given text for a target field and a method for establishing a domain-specific term network from a corpus. Background Art

[0002] Text classification is a commonly used algorithm for classifying a certain text into a specified category, and there are many current text classification algorithms. When only determining whether a text belongs to a certain category, it is a binary classification algorithm for text. Almost all algorithms that can be used for binary classification can be used for text classification, and in the case of having labels, these algorithms can achieve a relatively high accuracy (CN107908635A).

[0003] However, in many scenarios, existing algorithms still cannot meet the actual needs. The reason is that in actual scenarios, in addition to the given classification results, there are often other requirements. For example, extracting key information in the text can help readers quickly understand the text content and play an explanatory role for the classification result. Another example is the requirement for obtaining a knowledge base. People hope to accumulate knowledge by obtaining more texts, especially to establish a term library in the relevant field.

[0004] Existing text classification algorithms only incorporate interpretability in some simple cases, such as key words indicating positive or negative emotions in sentiment classification (CN102760153A). In more general classification scenarios, especially in the scenario of subject field classification, there are no such algorithms or applications. The present invention realizes text classification in the subject field using a term network. At the same time, the present invention combines a text classifier based on a term network with the iteration of the term network, and synchronously realizes a growable domain term network. It provides a solution for the demand scenarios of text classification and knowledge accumulation.

[0005] The background knowledge and background technology involved in the present invention mainly include term extraction, knowledge representation, and graph theory. Term extraction is a natural language processing technology for extracting scientific concepts from a corpus. The term extraction technology of the present invention uses the method proposed in publication number CN112966508A. Knowledge representation is used for the structured representation of a single text, and graph theory is used for constructing the features of a classifier. Summary of the Invention

[0006] The present invention provides a co-evolution method for text classification and term network growth. As shown in the appendix Figure 1 It is characterized in that this method consists of two parts: text classification and term network growth. These two parts are input to each other and optimize each other, and can achieve the effect of co-evolution.

[0007] The present invention has the following characteristics: When classifying text, the method can identify the contribution of each term in a single text to the classification result, thus making the text classification in the target field interpretable; after applying this text classification method to multiple texts, a term network in the target field can be constructed synchronously; using this method, only a general term network, a small-scale domain term network, and the text to be classified are needed to obtain a high-precision text classifier and the term network in the corresponding field; as the corpus size expands, the extracted term network in the specific field will become increasingly rich, and the text classification effect based on the term network will also become increasingly good.

[0008] In the present invention, the general term network refers to a large-scale network with terms in multiple fields as nodes and semantic or associative connections between terms as edges, denoted by G = (V G , E G ), where V G represents the node set of the general term network, and E G represents the edge set of the general term network. In the present invention, the role of the general term network is to provide background knowledge of terms. Specifically, we will use G to query related terms of a certain term and common related terms of multiple terms. The general term network G is an external input in the present invention and needs to be given in advance. In practical applications, it can be obtained through various methods, such as expert designation, algorithm construction, etc.

[0009] The domain term network refers to a network with terms in a specific field as nodes and term semantic associations as edges, denoted by G * = (V G* , E G* ), where V G* represents the node set of the domain term network, and E G* represents the edge set of the domain term network. In the present invention, the role of the domain term network is to provide domain knowledge and calibrate the domain relevance of terms. When the associated terms of a term are mostly concentrated in the target field, then the current term also has a relatively high probability of belonging to the target field. In the present invention, the domain term network G * is an external input during initialization and will be updated and expanded through the algorithm in the present invention and finally used as the output of the present invention, that is, the knowledge base of the target field.

[0010] The term subgraph refers to a term network constructed from a single text, denoted by g = (V g , E g ), where V g represents the node set of the term subgraph, and E gRepresents the set of edges of the term subgraph. The representation of the term subgraph of the text captures the main semantics in the text - terms and the co-occurrence relationships of terms, and is transformed into a graph structure that is easy for computers to process, which is the basis of the present invention.

[0011] I. Text classification algorithm based on term subgraph and term network.

[0012] Based on the term network classification method in this article, the accuracy of the text classifier can be continuously improved. The main steps include:

[0013] 1. Construct the input

[0014] The input of the algorithm includes a general term network G=(V G , E G ), a small amount of text T=[t 1 , t 2 ,..., t n with domain labels, where n≥1, and the text U to be classified =[u 1 , u 2 ,..., u n , where n≥1.

[0015] Among them, the construction method of the general term network used in the present invention is: ① Extract a certain scale of terms from a large corpus, and the technology used to extract terms is TERMATE (patent publication number CN112966508A); ② Establish a network using the co-occurrence relationship of terms in the document.

[0016] 2. Initialize the domain term network

[0017] For the domain term network, the initialized domain term network can be given by experts or constructed by an algorithm. In the present invention, a small amount of text T with labels is used to extract terms, and edges are constructed according to the co-occurrence relationship of terms, so as to obtain the initialized domain term network.

[0018] 3. Construct the term subgraph

[0019] For a single text u i , the term subgraph is established according to the following steps: ① Divide the text by unit (paragraph or sentence). ② Connect pairwise the terms co-occurring in each text unit, and connect the last term of the previous unit and the starting term of the next unit. ③ The weight of each connection is 1, and the weights can be accumulated, so as to obtain an undirected weighted term subgraph g=(V g , E g ).

[0020] 4. Domain relevance scoring of term subgraph nodes

[0021] Given the general term network G, the domain term network G *, the term sub-graph g, node n in g. The neighbor nodes of n in G are N G (n), in G * The neighbor nodes are N G* (n). In the present invention, it is stipulated that the neighbors of a node include itself.

[0022] We use the set of neighbor nodes of n in the term network to represent the knowledge background of the term. If the proportion of terms in the target field in the knowledge background is higher, the possibility that n belongs to the target field is also greater.

[0023] Furthermore, we divide the knowledge background into general fields and target fields: the relevant knowledge obtained in the general field is used to characterize the extensional relevance of the term to the target field; the relevant knowledge obtained within the target field is used to characterize the intensional relevance of the term to the target field.

[0024] The intensional relevance of node n to the target field is defined as

[0025] The extensional relevance of node n to the target field is defined as

[0026] The field relevance of a node is defined as the weighted sum of α_1 and α_2:

[0027] α = w * α 1 +(1 - w) * α 2 , 0 < w < 1

[0028] Among them, the larger α is, the higher the relevance of the node to the target field;

[0029] 5. Domain relevance scoring of the edges in the term sub-graph.

[0030] Given the general term network G, the domain term network G * , the term sub-graph g, node e = (n 1 , n 2 ).

[0031] Similar to the domain relevance of nodes, we use the common neighbors of n 1 , n 2 to represent the knowledge background of the edge n 1 , n 2 , and use the range of terms in the target field covered by this knowledge background to characterize the domain relevance of the edge.

[0032] Furthermore, the background knowledge in the general term network is used to characterize the extensional relevance of the edge e to the target field, and the knowledge background in the domain term network is used to characterize the intensional relevance of e to the target field.

[0033] The connotation relevance of the connecting edge e to the target domain is defined as

[0034] The extension relevance of the connecting edge e to the target domain is defined as

[0035] The domain relevance of the connecting edge is defined as:

[0036] β = β 1 +β 2

[0037] where the larger β is, the higher the relevance of the connecting edge to the target domain;

[0038] 6. Domain relevance scoring of the three - order hypergraph in the term sub - graph.

[0039] Given the general term network G and the domain term network G * , regarding the term sub - graph g as a three - order hypergraph, the hyper - edge h in g=(n 1 ,n 2 ,n 3 ).

[0040] Similar to the domain relevance of the connecting edge, we use the common neighbors of n 1 ,n 2 ,n 3 to represent the knowledge background of the hyper - edge n 1 ,n 2 ,n 3 , and use the scope of terms in the target domain covered by this knowledge background to characterize the domain relevance of the hyper - edge.

[0041] In a small non - dense network, when the number of nodes increases, the common neighbors of nodes are likely to become 0. Therefore, for hyper - edges, we use the background knowledge in the general term network G to characterize the relevance of h to the target domain.

[0042] The relevance of the hyper - edge h to the target domain is defined as where the larger γ is, the higher the relevance of the hyper - edge to the target domain;

[0043] 7. Classification of the term sub - graph

[0044] According to the scores of the nodes, connecting edges and hyper - edges in the term sub - graph to determine whether the text to be classified belongs to the target domain D, the present invention provides two determination methods, namely the unsupervised classification method and the supervised classification method, and either one can be selected.

[0045] 1) Unsupervised classification

[0046] Given the term sub - graph g=(V g ,E g ), the domain relevance of the nodes Domain relevance of connected edges Vector and The mean values of are denoted as and

[0047] When the number of nodes |V g | in g is ≥ 3, calculate the domain relevance of the hyperedge h i is the hyperedge when g is regarded as a third - order hypergraph, and the mean value of the vector is denoted as

[0048] Take as the input, construct the text classifier C: Set the threshold thresh. When , determine that the text belongs to the target domain D. When , determine that the text does not belong to the target domain D.

[0049] 2) Supervised classification

[0050] Given the term subgraph g=(V g , E g ), the domain relevance of nodes Domain relevance of connected edges When the number of nodes |V g | in g is ≥ 3, the domain relevance of hyperedge h i is the hyperedge when g is regarded as a third - order hypergraph.

[0051] Use a multi - layer feed - forward neural network to splice to obtain as the input, construct and train the text classifier C, and the output is the probability p of the text term belonging to the target domain D. When p ≥ 0.5, determine that the text belongs to the target domain D. When p < 0.5, determine that the text does not belong to the target domain D.

[0052] II. Growth algorithm of domain term network

[0053] The growth of the domain term network is a process of continuous expansion and update of a certain domain term network. As shown in the appendix Figure 2 , the main steps include:

[0054] 1. Construct sample subgraphs

[0055] Sample subgraphs include positive sample subgraphs and negative sample subgraphs. The advantage of constructing positive and negative sample subgraphs is that it can not only make full use of sample information but also aggregate the topological features of the network to enhance robustness. The specific steps are as follows:

[0056] 1) Use the trained text classifier C to classify the classified text U and get the positive sample Pos = {u i |C(u i )≥thresh} and negative samples;

[0057] 2) constructing a sample subgraph for each positive sample and negative sample according to the step of constructing a term subgraph in step 3 of the text classification algorithm based on term subgraph and term network;

[0058] 3) Aggregate all positive sample subgraphs into positive sample subgraphs

[0059] 4) Aggregate all negative sample subgraphs into negative sample subgraphs

[0060] 2. Sample subgraph regularization

[0061] For the sample subgraph, we use the following method for regularization: calculate the kcore value of each node in the sample subgraph, delete the nodes with kcore less than 2 and their edges, and retain the nodes with kcore greater than or equal to 2 and their edges. kcore ≥ 2 ensures that each node is in at least one local triangle motif, which is equivalent to strengthening the conditions for nodes to enter the positive sample subgraph and negative sample subgraph: nodes that appear only once in the positive sample, or appear multiple times but do not form a pairwise co-occurrence triangle motif, will not enter the positive sample subgraph; the same applies to the negative sample subgraph. Positive sample subgraph G P After regularization, we get Negative sample subgraph G N After regularization, we get

[0062]

[0063] 3. Terminology Network Update

[0064] The updating rule of the term network is as follows: the positive sample subgraph obtained in step 2 Add to existing domain term networkG * In, and from G * Subtract the negative sample sub-image obtained in step 2 from

[0065] Subtraction means subtracting the weight values ​​of the corresponding edges. When the weight of the edge after subtraction is less than or equal to 0, the edge is deleted. When the degree of the node is 0 after deleting the edge, the node is deleted.

[0066] Beneficial Effects

[0067] The present invention combines text classification with terminology network technology to develop a co-evolution algorithm, which has the following advantages:

[0068] 1) Improve the domain term network. By accumulating multiple corpora, a corresponding term network can be established. Traditional text classifiers acting on multiple samples cannot generate knowledge accumulation. In the present invention, by combining the classifier with the term network, both positive and negative samples can contribute to the generation of the domain term network, and as the scale of the corpus expands, the generated term network will be more perfect.

[0069] 2) Require fewer samples. This method requires fewer samples and can achieve better and better text classification effects in the process of co-evolving with the growth of the term network.

[0070] 3) High transferability. This method can be horizontally transferred to various fields and vertically can also be applied to custom domain hierarchies. It is applicable to some fields with relatively blurred disciplinary boundaries, especially interdisciplinary fields such as complex sciences, or specific topic classifications in a field.

[0071] 4) The text classification algorithm is interpretable. While the algorithm realizes text classification, it gives relevant term information and the relative contributions of terms and term co-occurrences to this classification, which provides interpretability for text classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 It is a schematic diagram of text classification and the growth of the term network.

[0073] Figure 2 It is a schematic diagram of text classification.

[0074] Figure 3 It is a schematic diagram of a general term network.

[0075] Figure 4 It is a schematic diagram of a term subgraph of labeled texts t1, t2.

[0076] Figure 5 It is a schematic diagram of the initialized domain term network in the example.

[0077] Figure 6 It is a schematic diagram of a term subgraph of unclassified texts u1 - u4 in the example.

[0078] Figure 7 It is a schematic diagram of positive and negative sample subgraphs before regularization in the example.

[0079] Figure 8 It is a schematic diagram of positive and negative sample subgraphs after regularization in the example.

[0080] Figure 9 It is a schematic diagram of the updated domain term network in the example. DETAILED DESCRIPTION OF THE INVENTION

[0081] To further illustrate the technical solution of the present invention, it will be specifically described below in conjunction with the accompanying drawings and examples.

[0082] In this section, the field of network science is selected to illustrate the implementation method of the present invention. It should be noted that the present invention is more suitable for processing a larger-scale collection of texts in actual use. In the cases of this section, only a small number of texts are selected, and the results can reflect the characteristics and relative advantages of the present invention, but they are not the best results expected by the present invention. In actual applications, with the increase in data, the effect of this method will also be improved.

[0083] I. Text Classification Based on Term Subgraphs and Term Networks

[0084] 1. Construct the input

[0085] The input of the algorithm includes a general term network G = (V G , E G ), a small amount of texts T = [t 1 , t 2 ,..., t n with domain labels, where n ≥ 1, and the text U to be classified = [u 1 , u 2 ,..., u n , where n ≥ 1. Fig. Figure 3 is a schematic diagram of the general term network.

[0086] The samples in the field of network science are T = [t 1 , t 2

[0087]

[0088]

[0089] The text to be classified is U = [u 1 , u 2 , u 3 , u 4

[0090]

[0091]

[0092] 2. Initialize the domain term network

[0093] The domain term network is given by experts or constructed by an algorithm. In the present invention, terms are extracted from a small amount of labeled texts T, and edges are constructed based on the co-occurrence relationship of the terms to obtain the initialized G * .

[0094] 1) Build a term subgraph. Using the terms in the general term network as the term library for this example, in t1, we match the terms {generation model, scale-free networks, power-law}, and the corresponding edges are (generation model, scale-free networks, 1), (scale-free network, power-law, 1). In t2, the terms we match are {complex networks, scale-free network}, and the corresponding edge is (complex network, scale-free network, 1). The term subgraphs corresponding to t1 and t2 are denoted as g t1 , g t2 , as shown in the appendix Figure 4 .

[0095] 2) Aggregate the term subgraphs into a domain term network

[0096] G * = g t1 ∪ g t2

[0097] G * 's nodes are the union of the nodes of g t1 , g t2 . The edges of G * are the union of the edges of g t1 , g t2 . If there are common edges, the weights of the edges should be added. Appendix Figure 5 is a schematic diagram of the initialized domain term network.

[0098] 3. Construct a term subgraph

[0099] Construct the term subgraph of the text according to the term nodes and term edges of the text. Appendix Figure 6 are the term subgraphs constructed for u1, u2, u3, u4 in the unclassified text U respectively.

[0100]

[0101] 4. Domain relevance scoring of nodes in the term subgraph (illustrated with g u1 as an example)

[0102] g u1The nodes are {statistical physics, social network, scale-free network}. The neighbor nodes of statistical physics in G are {statistical physics, many-body system, generation model}, and the number is 3, but these nodes are not in G * ; statistical physic has no neighbor nodes in G * According to the definitions of intension relevance and extension relevance, calculate the intension relevance α 1 and α 2 :

[0103]

[0104] Take w = 1 / 2, then the relevance of the node statistical physics to the field of network science is:

[0105]

[0106] Similarly, the neighbor nodes of social network in G are {social network, scale-free, complex network, social interaction}, and there are 2 nodes, scale-free and complex network, in G * ; social network has no neighbor nodes in G *

[0107] The intension relevance and extension relevance of social network are respectively:

[0108]

[0109] The relevance of the node social network to the field of network science is:

[0110]

[0111] The neighbor nodes of scale-free in G are {scale-free, power-law, complex network, social network, epidemic spreading, evolutionary dynamics}, and there are 3 nodes in G * ; scale-free in G​* The neighbor nodes in

[0112] The intension correlation and extension correlation of scale-free are respectively:

[0113]

[0114] The correlation of the scale-free node to the field of network science is:

[0115]

[0116] Therefore, for the term subgraph g u1 The correlation of the nodes to the field of network science is: [0, 1 / 4, 3 / 4]. Similarly, the correlations of the nodes of the term subgraphs g u2 , g u3 , g u4 to the field of network science are [1 / 2, 3 / 4, 1 / 4, 0], 1 / 8, 0, 0] respectively. From the above results, it can be seen that the node correlations of g u1 and g u2 are relatively high, while the node correlations of g u3 and g u4 are relatively low.

[0117] 5. Domain correlation scoring of the edges in the term subgraph (taking g u1 as an example for illustration)

[0118] The edges of g u1 are (statistical physics, social network, 1), (social network, scale-free, 1). For the edge (statistical physics, social network, 1), its two nodes have no common neighbors in G and also have none in G * . According to the definition, β 1 = 0, β 2 = 0, β = 0

[0119] For the edge (social network, scale-free, 1), the two nodes have common neighbors complex network, social network, scale-free in G, and among them, complex network and scale-free have no common neighbors in G * and the two nodes have no common neighbors in G * .

[0120] The connotation relevance and extension relevance of the edge (social network, scale-free, 1) are respectively:

[0121] Thus, g u1 The domain relevance of the edge to network science is [0, 1 / 2].

[0122] Similarly, the domain relevance of the edges of the subgraphs g u2 , g u3 , g u4 to network science can be calculated as: [1 / 2, 1, 1 / 2, 0], [0, 0, 0], [0]. We again find that the edges of the subgraphs g u1 and g u2 have a better relevance to network science, while the edges of g u3 and g u4 have a lower relevance to the field of network science. Compared with nodes, the domain relevance of edges utilizes more information and thus has better discriminability.

[0123] 6. Domain relevance scoring of the third-order hypergraph in the term subgraph (illustrated by taking g u1 as an example)

[0124] In g u1 , the three nodes of the hyperedge (statistical physics, social network, scale-free) have no common neighbors in G, so the relevance of the hyperedge to the field of network science is 0.

[0125] Similarly, the domain relevance of the hyperedges of the subgraphs g u2 , g u3 , g u4 to network science can be calculated as [1 / 4, 0, 0], [0], [0] respectively. Since a hyperedge requires three nodes to have common neighbors, when the network scale is small, the possibility of three nodes having common neighbors is relatively low. Therefore, if the domain relevance of a hyperedge is not 0, it means that the text has a relatively high possibility of relevance to network science.

[0126] 7. Classification of the term subgraph (illustrated by taking g u1 as an example)

[0127] The classification of the term subgraph can adopt unsupervised methods or supervised methods. In this example, the unsupervised method is adopted.

[0128] Let the classification threshold thresh = 1 / 4, and the relevance score vector of the nodes in g u1 is The relevance score vector of the edges is The correlation score vector of the hyperedge is Calculated as

[0129] Therefore, it is determined that the sample u1 belongs to the field of network science. Similarly, for the subgraph g u2 , g u3 , g u4 The calculation results are as follows: g u2 : 1 / 2 + 2 / 3 + 1 / 12 > 1 / 4

[0130] g u3 : 1 / 24 + 0 + 0 < 1 / 4, g u4 : 0 + 0 + 0 < 1 / 4.

[0131] Therefore, we determine that the samples u1 and u2 belong to the field of network science, while u3 and u4 do not belong to the field of network science, thus achieving the text classification of the samples.

[0132] II. Growth of the term network

[0133] 1. Construction of sample subgraphs

[0134] According to the classification results of the above text classification algorithm based on term subgraphs and term networks, the texts u1 and u2 that belong to the field of network science are used as positive samples, and the texts u3 and u4 that do not belong to the field of network science are used as negative samples, denoted as Pos = {u 1 , u 2}, Neg = {u 3 , u 4}.

[0135] Aggregate g u1 , g u2 to obtain the positive sample subgraph G P . Aggregate g u3 , g u4 to obtain the negative sample subgraph G N , as shown in Appendix Figure 7 as shown.

[0136] 2. Regularization of sample subgraphs

[0137] For G P and G N , calculate the kcore value of the nodes respectively, and take the nodes with kcore ≥ 2 as the regularized positive and negative sample subgraphs as shown in Appendix Figure 8 as shown.

[0138] 3. Update of the term network

[0139] Update the domain term network according to the formula . Add the edges in to G* Among them, the social network is a newly added node, (social network, scale-free) and (social network, power-law) are newly added edges, and the weight of (scale-free, power-law) is increased by 1 on the original basis.

[0140] Subtract the edges in from G * Since G * does not contain the edges in , it remains unchanged. The term network obtained after one-step update is shown in the appendix Figure 9 . It can be seen that G * has grown by one step on the original basis.

[0141] We can also use G * to re-classify the samples. Due to the change of G * , the node scores and edge scores of the term subgraph may change, and thus can be re-classified and the network can be updated again until our expectations are met or the set stopping conditions are reached.

[0142] The above-described embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A co-evolution method for text classification and the growth of a term network, characterized in that, this method organically combines the text classification and the process of term network growth, and the specific steps include: 1) Classify the text based on the term subgraph and the term network; 1-1) Input of the construction algorithm, including the general term network G = (V G , E G ), where V G represents the node set of the general term network, and E G represents the edge set of the general term network, a small amount of text T = [t 1 , t 2 ,..., t n with target domain labels, n ≥ 1, and the text U to be classified = [u 1 , u 2 ,..., u n , n ≥ 1, the initialized domain term network where represents the node set of the domain term network, represents the edge set of the domain term network; 1-2) According to the text to be classified U = [u 1 , u 2 ,..., u n , n ≥ 1, construct the term subgraph g = (V g , E g ); 1-3) Domain relevance scoring of the term subgraph, including the domain relevance scoring of nodes in the term subgraph, the domain relevance scoring of edges in the term subgraph, and the domain relevance scoring of three-order hypergraphs in the term subgraph; 1-4) Text classification. Determine whether the text to be classified belongs to the target domain D according to the scores of the nodes, edges, and hyperedges of the term subgraph. The determination method can be unsupervised classification or supervised classification; 2) Extract the term subgraph from the classified text and update the domain term network; 2-1) Construct sample subgraphs, and classify the text to be classified U = [u 1 , u 2 ,..., u n , n≥1 using the above text classification algorithm, to obtain positive samples Pos = {u i | C(u i ) ≥ thresh} and negative samples Neg = {u i | C(u i ) < thresh}. Aggregate the subgraphs of all positive samples into a positive sample subgraph Aggregate the subgraphs of all negative samples into a negative sample subgraph 2-2) Sample subgraph regularization. Calculate the kcore value of each node for the positive and negative sample subgraphs respectively, delete the nodes with kcore less than 2 and their connecting edges, and retain the nodes with kcore greater than or equal to 2 and their connecting edges. The positive sample subgraph G P After regularization, we get The negative sample subgraph G N After regularization, we get 2-3) Update the term network by adding the positive sample subgraph obtained in step 2-2) to the existing domain term network G * and subtracting the negative sample subgraph obtained in step 2 * from G. Subtraction means subtracting the weight values of the corresponding connected edges. When the weight of a connected edge after subtraction is less than or equal to 0, delete that connected edge. When the degree of a node becomes 0 after deleting a connected edge, delete that node; ​ 3) Optimize the text classifier based on the updated domain term network; 4) Classify the text using the optimized text classifier; Iteratively performing the above steps can achieve the co-evolution of the text classifier and the term network, and obtain a text classifier with higher classification accuracy and a larger-scale domain term network.

2. The method according to claim 1, characterized in that, For a single text u i , the specific steps for constructing a term subgraph include: 1-2-1) Divide the text into paragraphs or sentences; 1-2-2) Connect pairwise the terms co-occurring in each text unit, and establish a connection between the last term of the previous unit and the starting term of the next unit; 1-2-2) The weight of each connection is 1, and the weights can be accumulated to obtain an undirected weighted term subgraph g = (V g , E g ).

3. The method according to claim 2, characterized in that The specific steps for the relevance scoring of the nodes, edges, and hyperedges of the term subgraph are as follows: 1-3-1) Domain relevance scoring of nodes in the term subgraph Given the general term network G and the domain term network G * , the term subgraph g, the node n in g, and the neighbor nodes of n in G are N G (n) in G * The neighbor nodes are Use intension relevance and extension relevance to characterize the relevance of the nodes of the term subgraph to the target domain: The connotation relevance of node n to the target domain is defined as The extensional relevance of node n to the target domain is defined as The domain relevance of a node is defined as α 1 and α 2 weighted sum of: α = w * α 1 +(1 - w) * α 2 , 0 < w < 1; where the larger α is, the higher the relevance of the node to the target domain; 1-3-2) Domain relevance scoring of edges in the term subgraph Given the general term network G and the domain term network G * , the term subgraph g, and the node e=(n 1 , n 2 ) in g; Use intension relevance and extension relevance to characterize the relevance of the edges of the term subgraph to the target domain: The connotation relevance of the connecting edge e to the target domain is defined as The extensional relevance of the connecting edge e to the target domain is defined as The domain relevance of the edge is defined as: β=β 1 +β 2 ; where the larger β is, the higher the relevance of the edge to the target domain; 1-3-3) Domain relevance scoring of three-order hypergraphs in the term subgraph Given the general term network G and the domain term network G * , consider the term subgraph g as a third-order hypergraph, and the hyperedge h in g = (n 1 , n 2 , n 3 ); The relevance of the hyperedge h to the target domain is defined as where the larger γ is, the higher the relevance of the hyperedge to the target domain.

4. The method according to claim 2, characterized in that, The specific steps of unsupervised classification are as follows: Given the term subgraph g=(V g ,E g ), the domain relevance of nodes the domain relevance of edges vectors and The means of are denoted as and When the number of nodes |V g | in g is ≥ 3, calculate the neighborhood correlation of the hyperedges h i is the hyperedge when g is regarded as a third-order hypergraph, and the mean value of the vector is denoted as Take as input to construct a text classifier C: Set a threshold thresh. When , it is determined that the text belongs to the target domain D. When , it is determined that the text does not belong to the target domain D.

5. The method according to claim 2, characterized in that, The specific steps of supervised classification are as follows: Given a term subgraph g = (V g , E g ), the domain relevance of nodes The domain relevance of connected edges When the number of nodes |V g | ≥ 3, the domain relevance of hyperedges h i is the hyperedge when g is regarded as a third-order hypergraph; Using a multi-layer feedforward neural network, concatenate to obtain As the input, construct and train a text classifier C, and the output is the probability p of the text term belonging to the target domain D. When p ≥ 0.5, it is determined that the text belongs to the target domain D; when p < 0.5, it is determined that the text does not belong to the target domain D.

Citation Information

Patent Citations

  • Incorporating lexicon knowledge to improve sentiment classification

    CN102760153A

  • Text classification model building and text classification methods and apparatuses

    CN107908635A

  • Universal automatic term extraction method

    CN112966508A

  • Domain term automatic extraction method based on abnormal sub-graph detection

    CN112528640A

  • Knowledge graph construction method based on adaptive few-sample relation extraction

    CN113254675A