Theme modeling method and device based on graph convolutional network and reinforcement learning

By combining the theme modeling method of graph convolution network and reinforcement learning, the problem of insufficient utilization of semantic and structural information in traditional methods is solved, and more efficient theme modeling and better theme classification are achieved.

CN120542379APending Publication Date: 2025-08-26ZHEJIANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510620285.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Traditional theme modeling methods fail to make full use of the semantic information and structural information of text data, and are inefficient and have limited modeling accuracy when processing large-scale text data. The existing graph neural networks lose important semantic structures in text-intensive networks, limiting the representation ability.

Method used

Combining graph convolution networks and reinforcement learning, word embeddings and document-word matrix are obtained through graph convolution networks, and the policy network is trained using the REINFORCE algorithm to generate joint embeddings and infer topic distributions, while calculating topic consistency and diversity to evaluate model quality.

Benefits of technology

It improves the ability of topic modeling to obtain mutual reference relationships between text sets, improves the independent training ability of topic models and the mining ability of topic quality, and provides higher quality topic classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120542379A_ABST
    Figure CN120542379A_ABST
Patent Text Reader

Abstract

The invention discloses a topic modeling method and device based on a graph convolutional network and reinforcement learning, and the method comprises the following steps: 1, collecting text data from a plurality of sources, carrying out the preprocessing of the text data, and constructing a data set; 2, constructing a joint model based on the graph convolutional network, and generating joint embedding; 3, based on joint embedding, training a strategy network of a continuous action space by using a REINFORCE algorithm to deduce topic distribution; and 4, calculating a theme quality value to evaluate the quality of the generated theme. According to the method, semantic association and structural information among documents are fully mined by using the graph convolutional network, so that the generated theme is more accurate and fits text content; the reinforcement learning algorithm used in the invention enables the model to more efficiently explore and learn theme distribution in the training process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a topic modeling method, and in particular to a topic modeling method and device based on graph convolutional networks and reinforcement learning, belonging to the fields of natural language processing and machine learning. Background Art

[0002] In recent years, with the rapid development of the internet, text data has exploded in size. Extracting valuable information from this massive amount of text data has become a crucial task in natural language processing. Topic modeling, a technique for automatically discovering latent topics from text data, has been widely used in fields such as text classification, information retrieval, and data mining. Traditional topic modeling methods, such as Latent Dirichlet Allocation (LDA) and its variant, ProdLDA, are primarily based on probabilistic statistical models, discovering topics by assuming distributional relationships between documents, topics, and vocabulary.

[0003] However, traditional topic modeling methods have many limitations: they fail to fully utilize the semantic and structural information in text data; they are inefficient when processing large amounts of text data; and their modeling accuracy is limited. In recent years, the development of deep learning technology has brought new breakthroughs in the field of natural language processing. Deep learning-based topic modeling methods have gradually emerged. The combination of graph convolutional networks (GCNs) and reinforcement learning (RL) has provided new ideas for addressing the shortcomings of traditional topic modeling methods.

[0004] GCN is a deep learning model specifically designed for processing graph-structured data. It updates node representations layer by layer through information transfer and aggregation between nodes and their neighbors, effectively capturing both local and global features within the graph structure. It has demonstrated outstanding performance in tasks such as social network analysis, recommender systems, and text classification. Reinforcement learning, which learns optimal policies through the interaction between an agent and its environment and possesses both exploration and exploitation properties, has been widely applied in natural language processing tasks such as text summarization, question answering, and machine translation.

[0005] However, most existing graph neural networks (GNNs) targeting text-intensive networks typically treat text as node attributes. This approach inevitably loses important semantic structure and limits the GNN's representational capabilities. Furthermore, due to the long history of ProdLDA research and its outdated structure, a more modern transformation of topic models is urgently needed.

[0006] Therefore, in view of the above-mentioned shortcomings of topic modeling and the advantages of GCN and reinforcement learning, it is necessary to study a new topic modeling method to improve the ability of topic modeling to obtain document semantic associations and its autonomous training capabilities. Summary of the Invention

[0007] The purpose of the present invention is to provide a topic modeling method and device based on graph convolutional networks and reinforcement learning. The present invention improves the performance and accuracy of topic modeling by combining graph convolutional networks and reinforcement learning technology.

[0008] To achieve the above objectives, this patent provides the following technical solutions: a topic modeling method based on graph convolutional networks and reinforcement learning, comprising the following steps:

[0009] S1: Collect text data from the Internet and other sources and clean the data appropriately;

[0010] S2: Use graph convolutional networks to obtain word embeddings for the dataset, use bag-of-words models to obtain the document-word matrix for the dataset, and obtain the joint embedding generated during training;

[0011] S3: Using joint embedding and the REINFORCE algorithm, we train a policy network in the continuous action space to infer the topic distribution and complete the topic modeling process of the reinforcement learning strategy.

[0012] S4: Calculate topic consistency, topic diversity, and topic quality to evaluate the quality of model-generated topics.

[0013] In step S1, the data cleaning method is to remove stop words, remove punctuation marks, unify Chinese and English characters, and convert the text data into a format suitable for model input.

[0014] Preferably, the step S2 specifically includes:

[0015] S2.1 Generate joint embedding using graph convolutional network encoder;

[0016] S2.2 Calculate the loss value of the graph convolutional network.

[0017] Preferably, the process of generating the joint embedding in step S2 is:

[0018] S2.1.1 treats the dataset as an undirected, unweighted graph, where each document is considered a node, forming the adjacency matrix A and the document-word matrix W;

[0019] S2.1.2 Calculate the normalized adjacency matrix in, is the adjacency matrix with self-connection added, is the degree matrix, i N is the identity matrix;

[0020] S2.1.3 Calculate the normalized document-word matrix Where |W| is the modulus of the document-word matrix W;

[0021] S2.1.4 Constructing a two-layer graph convolutional network g based on document similarity φ :

[0022]

[0023] Generate latent joint embedding z i .

[0024] Where σ is the activation function (such as ReLU), ⊙ represents element-wise multiplication, is the weight matrix of the first layer, H (1) is the output of the first layer, is the weight matrix of the second layer, H (2) is the output of the second layer.

[0025] In step S2, the loss value of the graph convolutional network is calculated as follows:

[0026] S2.2.1 Generate graph topology embedding based on joint embedding: η i =h (G) (z i ), where h (G) is a neural network;

[0027] S2.2.2 Use the distance function to calculate the probability of connection between nodes i and j and reconstruct the adjacency matrix: where f τ (η j ,η j )=σ(τ-‖η i -η j ‖ 2 ), σ represents the logistic sigmoid function, and τ is a learnable parameter;

[0028] S2.2.3 Generate document embedding based on joint embedding: δ i =h (T) (z i ), h (T) is a neural network;

[0029] S2.2.4 Generate topic ratio of documents based on document embedding: θ i =softmax(δ i );

[0030] S2.2.5 Initialize the topic embedding α and word embedding matrix ρ to generate the topic-word distribution: β = softmax(α T ρ), reconstruct the document-word matrix: Among them, M i is the number of words contained in the i-th node;

[0031] S2.2.6 Calculate graph reconstruction loss:

[0032] S2.2.7 Calculate document reconstruction loss:

[0033] S2.2.8 Calculate the joint embedding loss function: L total =aL graph +bL doc , a and b are weight parameters.

[0034] In step S3, the process of topic modeling is as follows:

[0035] S3.1 Initialize the neural network parameters θ and construct the neural network;

[0036] S3.2 will embed the joint input neural network to obtain the mean and standard deviation of the action, and construct the Gaussian distribution policy function according to the REINFORCE algorithm principle;

[0037] S3.3 samples actions from Gaussian distribution (i.e., topic distribution);

[0038] S3.4 calculates the weighted ELBO (evidence lower bound) as the reward based on the sampled actions;

[0039] S3.5 updates the network parameters based on the accumulated rewards and policy gradients.

[0040] In step S3, the contents of the REINFORCE algorithm used are as follows:

[0041] Input: A differentiable parameterized policy function π(a|s,θ).

[0042] Algorithm parameters: step size α>0, discount factor γ<1.

[0043] Here are the steps:

[0044] 1. Initialize the parameter θ;

[0045] 2 In each round, the loop executes:

[0046] 2.1 Generate a round according to strategy π, including: s0, a0, r1, ..., s T-1 ,a T-1 ,r T

[0047] 2.2 At each step t in this round from 0 to T-1, do the following calculation:

[0048] 2.2.1 Calculating Cumulative Rewards

[0049] 2.2.2 Update Parameters

[0050] 2.3 ends.

[0051] 3End.

[0052] In step S3, the symbols of the REINFORCE algorithm are explained as follows:

[0053] a represents the topic distribution of the document, s represents the embedding representation of the document, and r represents the reward for selecting the action.

[0054] In step S3, the calculation method of the REINFORCE strategy function π is: Where μ(s,θ) is the mean and σ(s,θ) is the standard deviation, both of which are generated by the neural network based on the state s and parameter θ.

[0055] In step S3, the reinforcement learning strategy used is explained as follows:

[0056] The document embedding is considered a state in the round, and the document topic distribution is considered an action taken in the state. During training, the policy network selects an action based on the current embedding (i.e., state), that is, samples a topic distribution, and then transfers to a new state based on the action and gives a reward, and so on, until the end of the round, obtaining the complete sequence s0, a0, r1, ..., s T-1 ,a T-1 ,r T , completing a complete exploration in the distribution.

[0057] For each time step t, the cumulative reward R is calculated to reflect the long-term return obtained from taking subsequent actions starting from the current time step, and then based on the cumulative reward R and the gradient of the policy function Update the parameters.

[0058] In step S3, the constructed neural network structure is as follows:

[0059] The input layer uses the joint embedding obtained in step S2 as input, replacing the traditional bag-of-words (BoW) embedding;

[0060] The hidden layer uses the GELU activation function and layer normalization, and dropout is added after the fully connected layer to prevent overfitting;

[0061] ρ is initialized with ρ~N(0,0.02), and weight decay regularization is added to each layer;

[0062] The output layer generates the mean μ and standard deviation σ, which in turn generates a Gaussian distribution.

[0063] In step S3, the calculation method of the weighted evidence lower bound is as follows:

[0064] S3.4.1 Compute the KL divergence D between the posterior distribution P and the prior distribution Q KL (P‖Q);

[0065] S3.4.2 Calculate the log-likelihood of the model;

[0066] S3.4.3 Use the hyperparameter λ to adjust the weight of the KL divergence term and calculate the weighted evidence lower bound: ELBO weighted =λ·D KL (P‖Q)-log-likelihood.

[0067] In step S3, the reward value r at each time step is ELBO weighted .

[0068] In step S4, the process of calculating the subject quality is as follows:

[0069] S4.1 Calculate the subject consistency value of the model;

[0070] S4.2 Calculate the topic diversity value of the model;

[0071] S4.3 calculates topic quality using topic consistency and topic diversity values.

[0072] Preferably, in step S4, the method for calculating the topic consistency is:

[0073] S4.1.1 Calculate word w i The probability P(w i );

[0074] S4.1.2 Calculate word w i and w j The probability P(w i ,w j );

[0075] S4.1.3 Calculate word w by formula i Normalized point mutual information of:

[0076]

[0077] S4.1.4 For the generated topic, select the n words with the highest frequency as the high-frequency words of the topic. The average value of the normalized point mutual information of these n words is the topic consistency value of the topic:

[0078] S4.1.5 Calculate the model's topic consistency value: the average of all topic consistency values: TC = AvgTC t .

[0079] Preferably, in step S4, the method for calculating topic diversity is:

[0080] S4.2.1 Count the number of words N in the first N high-frequency words of a topic that do not appear in other topics. u ;

[0081] S4.2.2 Calculate the topic diversity value for each topic:

[0082] S4.2.3 Calculate the topic diversity value of the model: TD = AvgTD t .

[0083] In step S4, the method for calculating the subject quality is:

[0084] Multiply the topic consistency value by the topic diversity value to obtain the topic quality value: TQ = TC × TD.

[0085] The second aspect of the present invention relates to a topic modeling device based on a graph convolutional network and reinforcement learning, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement a topic modeling method based on a graph convolutional network and reinforcement learning of the present invention.

[0086] The third aspect of the present invention relates to a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements a topic modeling method based on graph convolutional networks and reinforcement learning of the present invention.

[0087] The innovation of this invention is: a topic modeling method based on graph convolutional networks and reinforcement learning is proposed. This method uses graph convolutional networks and the REINFORCE algorithm in reinforcement learning to build a model, generates a joint embedding containing the graph structure of the text set, and trains a policy network in the continuous action space based on the joint embedding to infer the topic distribution, and at the same time calculates the topic quality value to evaluate the quality of the generated topics.

[0088] The advantages of the present invention are as follows: the present invention improves the ability of the topic model to obtain the mutual reference relationship of the text set by replacing the word embedding in the traditional topic model with a joint embedding including document topology and document-term topology generated by a graph convolutional network; the present invention improves the model structure by replacing the variational autoencoder (VAE) of ProdLDA with reinforcement learning, thereby improving the topic model's ability to mine topic quality, especially topic consistency. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] Figure 1 is a flow chart of the present invention;

[0090] Figure 2 This is the model structure of the present invention.

[0091] Figure 3 It is a schematic diagram of the device of the present invention. DETAILED DESCRIPTION

[0092] To make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be described in further detail below with reference to the accompanying drawings. It should be understood that these embodiments are intended only to illustrate the present invention and are not intended to limit the scope of the present invention. In addition, it should be understood that after reading the contents of the present invention, those skilled in the art may make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0093] Example 1

[0094] This embodiment relates to a method for scientific research literature classification and citation classification based on a topic modeling method of graph convolutional network and reinforcement learning provided by the present invention, such as Figure 1 and Figure 2 As shown, the following steps are included: S1: collect text data from the Internet and other sources, and clean the data appropriately; S2: use the graph convolutional network to obtain the word embedding of the dataset, use the bag-of-words model to obtain the document-word matrix of the dataset, and obtain the joint embedding generated during training; S3: use the joint embedding and the REINFORCE algorithm to train the policy network of the continuous action space to infer the topic distribution and complete the topic modeling process of the reinforcement learning strategy; S4: calculate the topic consistency, topic diversity and topic quality, and evaluate the quality of the topics generated by the model; step S5: classify scientific research documents according to user needs.

[0095] Step S1: Collect text data from the Internet and other sources and perform appropriate data cleaning.

[0096] Specifically, in an academic literature classification application scenario, the Cora dataset is used as an example. This dataset contains 2,708 scientific papers categorized into seven categories: case-based, genetic algorithm, neural network, probabilistic methods, reinforcement learning, rule learning, and theory. It covers 5,429 citation relationships and 1,433 basic vocabulary. Each document uses a word vector with a 0 / 1 value to represent the occurrence of dictionary terms. The Cora-enrich dataset expands the text features of the Cora dataset to 25,955 words, forming a network. This dataset maintains the same document collection, classification system, and citation relationships as Cora.

[0097] After obtaining the text set, preprocessing work is performed on the text set, including word segmentation, removal of stop words, removal of low-frequency words, etc.

[0098] Step S2: Use the graph convolutional network to obtain the word embedding of the dataset, use the bag-of-words model to obtain the document-word matrix of the dataset, and obtain the joint embedding generated during training;

[0099] S2.1. Construct an N × N adjacency matrix A from an undirected, unweighted graph G consisting of N nodes.

[0100] S2.2 Select different words from the entire data set to form a vocabulary V;

[0101] S2.3 forms a document-term matrix W of size N × V;

[0102] S2.4 Calculate the self-connection adjacency matrix

[0103] S2.5 Calculation of degree matrix in When i≠j,

[0104] S2.6 Calculate the normalized adjacency matrix

[0105] S2.7 Calculate the normalized document-word matrix

[0106] S2.8 Construct a two-layer document similarity-based graph convolutional network g φ :

[0107]

[0108] Generate latent joint embedding Z;

[0109] S2.9 Generate graph topology embedding based on joint embedding: η i =h (G) (z i ), h (G) is a neural network;

[0110] S2.10 reconstructs the adjacency matrix and uses the distance function to calculate the probability of connection between nodes i and j. where f τ (η i ,η j )=σ(τ-‖η i -η j ‖ 2 ), σ represents the logistic sigmoid function, and τ is a learnable parameter;

[0111] S2.11 Generate document embedding based on joint embedding: δ i =h (T) (z i ), h (T) is a neural network;

[0112] S2.12 Generate the topic ratio of the document based on the document embedding: θ i =softmax(δ i );

[0113] S2.13 Initialize the topic embedding α and word embedding matrix ρ to generate the topic-word distribution: β = softmax(α T ρ);

[0114] S2.14 Reconstruct the document-word matrix: Among them, M i is the number of words contained in the i-th node;

[0115] S2.15 Calculate graph reconstruction loss:

[0116] S2.16 Calculate document reconstruction loss:

[0117] S2.17 Calculate the joint embedding loss function: L total =aL graph +bL doc , a and b are weight parameters, and the training model obtains the optimal joint embedding Z.

[0118] Step S3: Using joint embedding, the REINFORCE algorithm is used to train the policy network in the continuous action space to infer the topic distribution and complete the topic modeling process of the reinforcement learning strategy;

[0119] Following the principles of ProdLDA, we use the weighted product of expert models for word sampling. Following the principles of the REINFORCE algorithm, we continuously select topic distributions within the action space. The reward function is implemented using the minimum evidence lower bound obtained by summing the KL divergence and the log-likelihood estimate, completing the topic modeling process of the reinforcement learning strategy. The process is as follows:

[0120] S3.1 Initialize the neural network parameters θ, state s, step size α and discount factor γ;

[0121] S3.2 will embed Z into the neural network, get the mean μ and standard deviation σ of the action, and construct the Gaussian distribution strategy function according to the formula

[0122] S3.3 At time step t from 0 to T-1, perform the following operations:

[0123] S3.3.1 Sampling actions from the policy function (i.e., topic distribution): a t ~π(a|s,θ);

[0124] S3.3.2 Calculate the KL divergence D between the posterior distribution P and the prior distribution Q KL (P‖Q);

[0125] S3.3.3 Calculate the log-likelihood of the model;

[0126] S3.3.4 Calculate the reward value: r t =ELBO weighted =λ·D KL (P‖Q)-log-likelihood, where λ is the weight parameter;

[0127] S3.3.5 Calculation of cumulative rewards:

[0128] S3.3.6 Update parameters

[0129] S3.3.7 Execution state transitions t →s t+1 ;

[0130] Repeat the above process to train the model until convergence and obtain the best topic distribution inference.

[0131] Step S4: Calculate topic consistency, topic diversity, and topic quality to evaluate the quality of the topics generated by the model. The process is as follows:

[0132] S4.1 Calculate word w i The probability P(w i );

[0133] S4.2 Calculate word w i and w j The probability P(w i ,w j );

[0134] S4.3 Calculate word w i Normalized point mutual information of:

[0135]

[0136] S4.4 In the generated topics, select the n words with the highest frequency as the high-frequency words of the topic, and calculate the average normalized point mutual information of the high-frequency words as the topic consistency value of the topic:

[0137] S4.5 Calculate the topic consistency value of all topics: TC = AvgTC topic ;

[0138] S4.6 Count the number of words N in the top N high-frequency words of a topic that do not appear in other topics. u ;

[0139] S4.7 Calculate the topic diversity value for each topic:

[0140] S4.8 Calculate the overall topic diversity value of the model: TD = AvgTD topic ;

[0141] S4.9 Calculate the topic quality value: TQ = TC × TD.

[0142] The topic quality of topic modeling on the Cora dataset using different embeddings under reinforcement learning strategies is shown in Table 1 below:

[0143] Table 1

[0144]

[0145]

[0146] It can be seen that the reinforcement learning topic modeling method using joint embedding provided by the present disclosure can provide more reasonable and higher-quality classification results for the scenario of classifying scientific research documents with citation relationships, thereby achieving better-quality document classification.

[0147] Step S5: Classify scientific research documents according to the needs of scientific researchers;

[0148] When the user specifies the number of classification categories before model training, the model performs topic modeling under the reinforcement learning strategy according to the content and mutual citation relationship of the input documents, and then completes the classification of the input documents according to the number of categories specified by the user.

[0149] The key symbols in this embodiment are shown in Table 2:

[0150] Table 2

[0151]

[0152]

[0153] Example 2

[0154] Reference Figure 3 This embodiment relates to a topic modeling device based on graph convolutional networks and reinforcement learning, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the scientific research literature classification and citation classification method of Example 1 using the topic modeling method based on graph convolutional networks and reinforcement learning of the present invention.

[0155] Example 3

[0156] This embodiment relates to a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the scientific research literature classification and citation classification method of Example 1 using the topic modeling method based on graph convolutional network and reinforcement learning of the present invention is implemented.

[0157] Any details not described in detail herein are well known to those skilled in the art. Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present invention may be modified or replaced with equivalents without departing from the spirit and scope of the technical solutions of the present invention, and such modifications or equivalents should be encompassed by the claims of the present invention.

Claims

1. A topic modeling method based on graph convolutional networks and reinforcement learning, characterized by: The method comprises the following steps: S1: Collect text data from multiple sources, preprocess the text data, build datasets, and perform data cleaning; S2: Build a model based on the graph convolutional network encoder to generate joint embeddings; S3: Based on joint embedding, the REINFORCE algorithm is used to train a policy network in the continuous action space to infer the topic distribution and complete the topic modeling process of the reinforcement learning strategy; S4: Calculate the topic consistency, topic diversity and topic quality values ​​to evaluate the quality of the generated topics.

2. The method according to claim 1, wherein: The data cleaning described in step S1 includes removing stop words, removing punctuation marks, unifying Chinese and English characters, and converting the text data into a format suitable for model input.

3. The method according to claim 1, wherein: In step S2, the model is built using the graph convolutional network encoder. The process of generating the joint embedding includes: S2.1: Treat the dataset as an undirected, unweighted graph, each document as a node, and calculate the normalized adjacency matrix and the normalized document-term matrix S2.2: Construct a two-layer document similarity-based graph convolutional network g φ , generating a joint embedding Z.

4. The method according to claim 1, wherein: The topic modeling process described in step S3 includes: S3.1: Initialize neural network parameters and build neural network; S3.2: Input the joint embedding into the neural network and construct a Gaussian distribution policy function according to the REINFORCE algorithm principle; S3.3: Sample actions from a Gaussian distribution and compute a weighted evidence lower bound as the reward value. S3.4: Update parameters based on the reward value.

5. The method according to claim 4, characterized in that: The process of training the policy network using the REINFORCE algorithm in step S3.2 includes: S3.2.1: Input the policy function π(a|s,θ), whose parameters are the mean μ and the standard deviation σ; S3.2.2: At time steps t from 0 to T-1, perform the following operations: 1: Sampling action a t ; 2: Calculate the action reward value r t ; 3: Calculate the cumulative reward R (i.e. loss function); 4: Update parameters θ; 5: Execute state transfer s t →s t+1 ; S3.2.3: Repeat the above process to train the model until convergence and obtain the best topic distribution inference.

6. The method according to claim 5, characterized in that: The generation method of the REINFORCE algorithm strategy function in step S3.2.1 is 7. The method according to claim 4, characterized in that: The calculation method of the weighted evidence lower bound in step S3.3 includes: S3.3.1: Compute the KL divergence D between the posterior distribution P and the prior distribution Q KL (P‖Q); S3.3.2: Calculate the log-likelihood of the model; S3.3.3: Computing the Weighted Evidence Lower Bound: ELBO weighted =λ·D KL (P‖Q)-log-likelihood.

8. The method according to claim 1, wherein: The method for calculating the topic quality value in step S4 is as follows: S4.1: Calculate word w i The probability P(w i ), calculate word w i and w j The probability P(w i ,w j ); S4.2: Calculate word w i Normalized point mutual information of: S4.3: In the generated topics, select the n words with the highest frequency as the high-frequency words of the topic, and calculate the average normalized point mutual information of the high-frequency words as the topic consistency value of the topic: S4.4: Calculate the topic consistency value of all topics: TC = AvgTC topic ; S4.5: Count the number of words N in the first n high-frequency words of the topic that do not appear in other topics u ; S4.6: Calculate the topic diversity value for each topic: S4.7: Calculate the overall topic diversity value of the model: TD = AvgTD topic ; S4.8: Calculate the topic quality value: TQ = TC × TD.

9. A topic modeling device based on graph convolutional network and reinforcement learning, characterized in that: The method comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement a topic modeling method based on graph convolutional network and reinforcement learning as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that A program is stored thereon, and when the program is executed by a processor, a topic modeling method based on a graph convolutional network and reinforcement learning as described in any one of claims 1 to 8 is implemented.