Text label prediction model training method, text label prediction method, device, equipment and medium

By initializing text collection using the Bert model and combining iterative training of the Text-level-BertGCN model and Text-level-GCN model, the problem that the Text-Level-GCN text classification method is difficult to extract semantic features is solved, achieving higher text classification accuracy.

CN115391525BActive Publication Date: 2025-07-18GUANGXI ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210941916.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-08
Publication Date
2025-07-18
Estimated Expiration
2042-08-08

AI Technical Summary

Technical Problem

The existing Text-Level-GCN text classification method is difficult to extract text features containing semantics, resulting in poor classification results.

Method used

The text word segmentation and word mapping module based on large-scale pre-trained Bert model is initialized to initialize the text set, and text-level graph is constructed, and iterative training of the Text-level-BertGCN model and Text-level-GCN model is combined with Bert classification prediction and GCN classification prediction for weighted synthesis, and the model is trained using cross entropy loss.

Benefits of technology

It improves the accuracy of text classification, can effectively extract the deep semantic and structural characteristics of the text, and improves the classification effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115391525B_ABST
    Figure CN115391525B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for training a text label prediction model, a text label prediction method, an apparatus, a device, and a medium, which relate to the field of artificial intelligence technology. Among them, the method for training a text label prediction model includes the following steps: S110, obtaining a training text set and its corresponding true labels; S120, initializing the texts in the training text set using the text tokenization and word mapping module of the first Bert model to obtain the feature representations of each word in the training text set, and constructing a text-level graph with the feature representations of each word as nodes; S130, training a prediction model. The method provided by the present invention can solve the problem that the existing Text-Level-GCN text classification method is difficult to extract text features containing semantics, resulting in poor classification effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and in particular to a method for training a text label prediction model, a text label prediction method, a device, a device and a medium. Background Art

[0002] Text classification is a very important task in natural language processing and has been applied in many places in reality, such as email detection, opinion mining, etc. GCN is a convolutional neural network based on a graph structure and is a text classification method that has been proven to be very effective in recent years.

[0003] When processing document data, GCN performs overall graph construction on the text set. That is, GCN regards a corpus as a large graph, and each text is a node of the graph. This graph construction method can obtain the global information of the corpus, but it ignores the relationship between words and the relationship between words and texts. Therefore, for a corpus of documents containing a large number of long texts, it still cannot achieve a satisfactory classification effect. For this reason, the neural network Text-Level-GCN based on the text-level graph has been developed to a certain extent, but the existing text-level graph neural network Text-Level-GCN is difficult to extract text features with semantics from the nodes of the text-level graph, so there is still a problem of poor classification effect. Summary of the Invention

[0004] Aiming at the problems existing in the prior art, the present invention provides a text label prediction method based on large-scale pre-training, which solves the problem that the existing Text-Level-GCN text classification method is difficult to extract text features with semantics, resulting in poor classification effect.

[0005] To achieve the above invention purpose, the technical solution of the present invention is as follows:

[0006] The present invention provides a method for training a text label prediction model, including the following steps:

[0007] S110, obtaining a training text set and the true labels corresponding to the training text set;

[0008] S120, initializing the texts in the training text set using the text tokenization and word mapping module of the first Bert model to obtain the feature representations of each word in the training text set, and constructing a text-level graph with the feature representations of each word as nodes, where: the feature representation of each word carries the semantic feature information of the corresponding word extracted by the first Bert model;

[0009] S130, training a prediction model, and the steps of training the prediction model are as follows:

[0010] The prediction model outputs a classification prediction;

[0011] Iteratively train the prediction model according to the classification prediction and the true label to obtain a trained prediction model;

[0012] Wherein: the prediction model includes a Text-level-BertGCN model, and the Text-level-BertGCN model includes a second Bert model and a Text-level-GCN model,

[0013] The classification prediction output by the prediction model includes the following steps:

[0014] Use the text classification module of the second Bert model to calculate the feature representation of each node in the text-level graph to obtain the text classification feature of each node;

[0015] The Text-level-GCN model aggregates the text classification features and uses the softmax module to obtain the GCN classification prediction y gcn 。

[0016] Further, in the step S130, the prediction model further includes a third Bert model,

[0017] The classification prediction output by the prediction model further includes the following steps:

[0018] The third Bert model processes the text in the training text set to obtain text features, and then processes them through the softmax module to obtain the Bert classification prediction y bert ;

[0019] Combine the Bert classification prediction y bert and the GCN classification prediction y gcn through weighted synthesis to obtain a classification prediction.

[0020] Further, in the step S130, the iterative training of the prediction model according to the classification prediction and the true label includes:

[0021] Calculate the cross-entropy loss according to the classification prediction and the true label. When the cross-entropy loss value is greater than or equal to a preset loss threshold, iteratively train the prediction model until the cross-entropy loss value is less than the loss threshold, and then stop training to obtain a trained prediction model.

[0022] Further, the weighted synthesis method is: Wherein:

[0023] is a variable parameter, with a value range of [0.1, 1], used to balance the Bert classification prediction y bert and the GCN classification prediction y gcn The role of y is the classification prediction.

[0024] Furthermore, the cross-entropy loss calculation formula is as follows:

[0025] loss = -g log y, where: Loss is the cross-entropy loss, g is the true label, and y is the classification prediction.

[0026] The present invention also provides a text label prediction method, including:

[0027] S210, obtaining a text set to be label-predicted;

[0028] S220, initializing the text in the text set to be label-predicted using the text tokenization and word mapping module of the first Bert model, obtaining the feature representation of each word in the text set to be label-predicted, and constructing a text-level graph with the feature representation of each word as a node, where: the feature representation of each word carries the semantic feature information of the corresponding word extracted by the first Bert model;

[0029] S230, inputting the text-level graph and the text in the text set to be label-predicted into the prediction model to obtain a text label prediction, and the prediction model is trained by the training method of the text label prediction model.

[0030] The present invention also provides a text label prediction training device, including the following modules:

[0031] An acquisition module, used to acquire a training text set and the true label corresponding to the training text set;

[0032] A graph construction module, used to initialize the text in the training text set using the text tokenization and word mapping module of the first Bert model, obtaining the feature representation of each word in the training text set, and constructing a text-level graph with the feature representation of each word as a node, where: the feature representation of each word carries the semantic feature information of the corresponding word extracted by the first Bert model;

[0033] A training module, used to train the prediction model, and the steps of training the prediction model are as follows:

[0034] The prediction model outputs a classification prediction;

[0035] Iteratively training the prediction model according to the classification prediction and the true label to obtain a trained prediction model;

[0036] Among them: the prediction model includes a Text-level-BertGCN model, and the Text-level-BertGCN model includes a second Bert model and a Text-level-GCN model.

[0037] The classification prediction output by the prediction model includes the following steps:

[0038] Use the text classification module of the second Bert model to calculate the feature representation of each node in the text-level graph, and obtain the text classification features of each node.

[0039] The Text-level-GCN model aggregates the text classification features and uses the softmax module to obtain the GCN classification prediction y gcn .

[0040] The present invention also provides a text label prediction device, including:

[0041] An acquisition module for acquiring a text set to be predicted with labels.

[0042] A graph construction module for initializing the texts in the text set to be predicted with labels using the text tokenization and word mapping module of the first Bert model, obtaining the feature representation of each word in the text set to be predicted with labels, and constructing a text-level graph with the feature representation of each word as a node, where: the feature representation of each word carries the semantic feature information of the corresponding word extracted by the first Bert model.

[0043] A prediction module for inputting the text-level graph and the texts in the text set to be predicted with labels into a prediction model to obtain a text label prediction, and the prediction model is trained by the training method of the text label prediction model.

[0044] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the text label prediction model training method or the text label prediction method.

[0045] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the text label prediction model training method or the text label prediction method.

[0046] Compared with the prior art, the present invention has the following beneficial effects:

[0047] 1. The training method for the text label prediction model provided by the present invention initializes the texts in the training text set using the text tokenization and word mapping module of the first Bert model, obtains the feature representations of each word in the training text set, and constructs a text-level graph with the feature representations of each word as nodes, where: the feature representation of each word carries the semantic feature information of the corresponding word extracted by the first Bert model; then, classification prediction training is performed using the prediction model Text-level-BertGCN model. First, the text classification module of the second Bert model calculates the feature representations of each node in the text-level graph to obtain the text classification features of each node. The Text-level-GCN model then aggregates the text classification features and uses the softmax module to obtain the GCN classification prediction; finally, the prediction model is iteratively trained according to the classification prediction and the true label output by the Text-level-BertGCN model to obtain the trained prediction model. By using the first Bert model to initialize the texts in the training text set, the feature representations of each word in the training text set are obtained, and a text-level graph is constructed with the feature representations of each word as nodes, where: the feature representation of each word carries the semantic feature information of the corresponding word extracted by the first Bert model; then, the text classification module of the second Bert model calculates the feature representations of each node in the text-level graph to obtain the text classification features of each node, and then inputs the text classification features into the Text-level-GCN model, solving the problem that the existing Text-Level-GCN text classification method is difficult to extract text features containing semantics, resulting in poor classification effects.

[0048] 2. Jointly train two models, the Text-level-BertGCN model and the third Bert model, and perform weighted synthesis on the Bert classification prediction y bert and the GCN classification prediction y gcn to obtain the classification prediction. Iteratively train the two models, the Text-level-BertGCN model and the third Bert model, according to the classification prediction and the true label to obtain the trained prediction model, further improving the accuracy of text classification. Description of the Drawings

[0049] Figure 1 It is a flowchart of a training method for a text label prediction model according to an embodiment of the present invention.

[0050] Figure 2 It is an overall model diagram of a training method for a text label prediction model according to an embodiment of the present invention.

[0051] Figure 3This is a model flow chart of a method for training a text label prediction model according to an embodiment of the present invention.

[0052] Figure 4 This is a flow chart of a text label prediction method according to an embodiment of the present invention.

[0053] Figure 5 This is a schematic structural diagram of a text label prediction training device according to an embodiment of the present invention.

[0054] Figure 6 This is a schematic structural diagram of a text label prediction device according to an embodiment of the present invention.

[0055] Figure 7 This is a schematic structural diagram of a computer device according to an embodiment of the present invention. Detailed implementation manners

[0056] The following further elaborates on the present invention in conjunction with embodiments, so that those skilled in the art can implement it with reference to the description in the specification.

[0057] To elaborate in detail on the technical content, achieved objectives, and effects of the present invention, the following is described in conjunction with embodiments and with reference to the accompanying drawings.

[0058] Embodiment 1

[0059] As Figure 1 shown, the method for training a text label prediction model includes the following steps:

[0060] S110. Obtain a training text set and the corresponding true labels for the training text set;

[0061] S120. Initialize the texts in the training text set using the text tokenization and word mapping module of the first Bert model to obtain the feature representations of each word in the training text set, and construct a text-level graph with the feature representations of each word as nodes, where: the feature representation of each word carries the semantic feature information of the corresponding word extracted by the first Bert model;

[0062] Let D be the training text set, and T be the text in the training text set D. For text T, T = {v1, v2,..., v n}, where v is a word in text T, that is, text T contains n words; after initializing text T using the text tokenization and word mapping module of the first Bert model, there is T' = Bert(T), and the feature representation of text T' is X b ∈R |v|*d , where R represents the Euclidean space of the features of text T', and |v| represents the number of feature representations of words in text T', which is numerically equal to v nwhere n is equal, and d is the dimensionality size after mapping words to vectors. is the feature representation corresponding to the i-th word in text T, carrying the semantic feature information of the i-th word extracted by the first Bert model. The text-level graph is defined as G T‘ ={V, E, A}, where V is the set of nodes in the graph, E is the set of edges, and A is the global shared mutual information matrix. The global shared mutual information matrix A = PMI(D) is constructed for the training text set D using the PMI algorithm. When the PMI value is greater than 0, it indicates a high semantic association degree between words. A[i, j] represents the weight between word i and word j; the feature representations of all words in text T are used as the nodes of the text-level graph. In this way, the nodes in the text-level graph also carry the semantic feature information of the corresponding words extracted by the first Bert model; at the same time, in the text-level graph G T′ ={V, E, A}, an edge e is connected between node i and node j ij , and the edge set The weight value of edge e ij comes from A, that is represents the number of connections between adjacent words in the text-level graph, which is also the size of the sliding window. When constructing the text-level graph, a fixed sliding small window is used to construct the text-level graph.

[0063] S130. Train the prediction model. The steps for training the prediction model are as follows:

[0064] The prediction model outputs a classification prediction;

[0065] Iteratively train the prediction model according to the classification prediction and the true label to obtain a trained prediction model;

[0066] where: The prediction model includes a Text-level-BertGCN model, and the Text-level-BertGCN model includes a second Bert model and a Text-level-GCN model. The Text-level-BertGCN model is used to output a GCN classification prediction y gcn ,

[0067] The steps for the prediction model to output a classification prediction include:

[0068] Use the text classification module of the second Bert model to calculate the feature representation of each node in the text-level graph to obtain the text classification feature of each node;

[0069] The Text-level-GCN model aggregates the text classification features and uses the softmax module to obtain the GCN classification prediction y gcn .

[0070] The text classification features output by the classification module using the second Bert model serve as the input for the Text-level-GCN model. The second Bert model calculates each classification feature in the text-level graph and uses the message passing mechanism (MPM) (Gilmer et al., 2017) for feature update, where nodes can extract structural features locally. The message passing mechanism performs convolution through a non-spectral method. Nodes (i.e., source nodes) in the text-level graph first collect information from neighboring nodes, and then aggregate and update the source node information and neighboring node information.

[0071]

[0072]

[0073] m i ∈R d is the feature information aggregated by node i from neighboring nodes. Rd represents the d-dimensional Euclidean space, and max is a function that combines the maximum values in each dimension into a new feature vector as the output. represents the nearest neighbor nodes to node i. represents the previous feature representation of source node i, represents the feature representation of neighboring nodes, x′ i represents the new feature representation of node i aggregating neighboring information, that is, the updated text classification feature of node i in the text-level graph, η i ∈R 1 is the trainable parameter of node i, indicating how much information should be retained.

[0074] The Text-level-GCN model aggregates the text classification features of all points in the text-level graph of the training text set D and uses the text classification features of all nodes in the text-level graph to predict the label of the text, that is, obtaining the GCN classification prediction y gcn :

[0075]

[0076] where W is the matrix that maps the text classification features to the output space. N T is the node set of text T, b is the bias, softmax is the classification function, a conventional function used in the last step of deep learning, which turns the classification result into a probability value, and relu is the activation function.

[0077] In summary, for the method provided by the present invention, the text in the training text set D is initialized using the text tokenization and word mapping module of the first Bert model to obtain the feature representation of each word in the training text set D, and a text-level graph is constructed with the feature representation of each word as a node, where: the feature representation of each word carries the semantic feature information of the corresponding word extracted by the first Bert model; then, classification prediction training is performed using the prediction model Text-level-BertGCN model. First, the text classification module of the second Bert model is used to calculate the feature representation of each node in the text-level graph to obtain the text classification feature of each node. The Text-level-GCN model then aggregates the text classification features and uses the softmax module to obtain the GCN classification prediction; finally, the prediction model is iteratively trained according to the classification prediction and the true label output by the Text-level-BertGCN model to obtain a trained prediction model. By using the first Bert model, the text in the training text set D is initialized to obtain the feature representation of each word in the training text set D, and a text-level graph is constructed with the feature representation of each word as a node, where: the feature representation of each word carries the semantic feature information of the corresponding word extracted by the first Bert model; then, the text classification module of the second Bert model is used to calculate the feature representation of each node in the text-level graph to obtain the text classification feature of each node, and the text classification feature is input into the Text-level-GCN model to solve the problem that it is difficult for the existing Text-Level-GCN text classification method to extract text features containing semantics, resulting in poor classification effects.

[0078] In this embodiment, in step S130, the iterative training of the prediction model according to the classification prediction and the true label includes:

[0079] Calculate the cross-entropy loss according to the classification prediction and the true label. When the cross-entropy loss value is greater than or equal to a preset loss threshold, perform iterative training on the prediction model until the cross-entropy loss value is less than the loss threshold, and then stop training to obtain a trained prediction model.

[0080] The cross-entropy loss calculation formula is as follows:

[0081] loss = -glogy, where: Loss is the cross-entropy loss, g is the true label, and y is the classification prediction.

[0082] Comparative Example 1

[0083] Based on initializing the text in the training text set in the text tokenization and word mapping module of the first Bert model, the Text-level-BertGCN model is iteratively trained, improving the accuracy of text classification. Compared with the existing Text-GCN model and Text-Level-GCN model, the results are shown in Table 1.

[0084]

[0085] Table 1 Comparison of classification accuracies (%) of Text-level-BertGCN model, Text-GCN model and Text-Level-GCN model

[0086] Among them, R8, R52 and Ohsumed are common English text classification data sets. As can be seen from Table 1, the classification accuracy of the Text-level-BertGCN model is higher than that of the Text-Level-GCN model because the Text-level-BertGCN model combines the advantages of the Bert model and the Text-Level-GCN model and can extract the deep semantics of the text and the text structure features.

[0087] Example two

[0088] Based on Example one, in step S130, the prediction model further includes a third Bert model for outputting the Bert classification prediction y bert ,

[0089] The prediction model outputting the classification prediction further includes the following steps:

[0090] The third Bert model processes the text in the training text set to obtain the text features of the text, and then through the softmax module processing, obtains the Bert classification prediction y bert ;

[0091] Weight the Bert classification prediction y bert and the GCN classification prediction y gcn for weighted synthesis to obtain the classification prediction.

[0092] Specifically, the third Bert model uses its own classifier to act on the text T, inputs the text T in the training text set D into the softmax layer of the third Bert model, and obtains the classification prediction y bert . The calculation method is as follows:

[0093] Q = W Q *X T

[0094] K = W K*X T

[0095] V = W v *X T

[0096]

[0097]

[0098] Q, K, and V are matrices of queries, keys, and values respectively. W Q , W K and W V are parameter matrices that linearly map X T to Q, K, and V. X T is the vector representation of text T. K T represents the transpose matrix of K. D k is the dimension of the K matrix, represents the text features at the i-th iteration, is a trainable weight matrix.

[0099] Jointly train two models, the Text-level-BertGCN model and the third Bert model, and perform weighted synthesis on the Bert classification prediction y bert and the GCN classification prediction y gcn to obtain a classification prediction. Determine the cross-entropy loss based on the classification prediction and the true label. When the cross-entropy loss value is greater than or equal to a preset loss threshold, perform iterative training on the prediction model until the cross-entropy loss value is less than the loss threshold, and then stop training to obtain a trained prediction model;

[0100] In this implementation, the weighted synthesis method is: where:

[0101] is a variable parameter with a value range of [0.1, 1], used to balance the effects of the Bert classification prediction y bert and the GCN classification prediction y gcn . y is the classification prediction. The optimal value of can be determined through iterative training.

[0102] The overall model diagram and the model flow chart are shown in Figure 2 , Figure 3 respectively.

[0103] Comparative Example 2

[0104] The Text-level-BertGCN model and the third Bert model are fused and iteratively trained, further improving the accuracy of text classification. Compared with the Text-level-BertGCN model, the results are shown in Table 2.

[0105]

[0106] Table 2 Comparison of classification accuracy rates (%) between the fusion model and the Text-Level-GCN model

[0107] Among them, R8, R52, and Ohsumed are common English text classification datasets, and the fusion model is the fusion model of the Text-level-BertGCN model and the third Bert model after iterative training.

[0108] The amount of network information is increasing rapidly, and most of this information is unstructured massive text. This method can effectively process unstructured text in various applications through graph neural networks, improving text processing efficiency. The current solution model is compared with non-graph structured models, and the results are shown in Table 3:

[0109]

[0110] Table 3 Comparison of classification accuracy rates (%) between the fusion model, the Bi-LSTM model, and the fast-Text model

[0111] Example 3

[0112] The present invention provides a text label prediction method, as Figure 4 shown, including:

[0113] S210, obtaining a text set for which label prediction is to be performed;

[0114] S220, initializing the texts in the text set for which label prediction is to be performed using the text tokenization and word mapping module of the first Bert model, obtaining the feature representations of each word in the text set for which label prediction is to be performed, and constructing a text-level graph with the feature representations of each word as nodes, where: the feature representation of each word carries the semantic feature information of the corresponding word extracted by the first Bert model;

[0115] S230, inputting the text-level graph and the texts in the text set for which label prediction is to be performed into a prediction model to obtain text label prediction, and the prediction model is trained by the text label prediction model training method in Example 1 or Example 2.

[0116] Example 4

[0117] The present invention provides a text label prediction training device, asFigure 5 As shown, the device includes:

[0118] An acquisition module, configured to acquire a training text set and true labels corresponding to the training text set;

[0119] A graph construction module, configured to initialize the texts in the training text set using the text tokenization and word mapping module of the first Bert model, obtain the feature representations of each word in the training text set, and construct a text-level graph with the feature representations of each word as nodes, where: the feature representation of each word carries the semantic feature information of the corresponding word extracted by the first Bert model;

[0120] A training module, configured to train a prediction model. The steps of training the prediction model are as follows:

[0121] The prediction model outputs a classification prediction;

[0122] Iteratively train the prediction model according to the classification prediction and the true labels to obtain a trained prediction model;

[0123] Wherein: the prediction model includes a Text-level-BertGCN model, and the Text-level-BertGCN model includes a second Bert model and a Text-level-GCN model. The Text-level-BertGCN model is used to output a GCN classification prediction y gcn ,

[0124] The steps for the prediction model to output a classification prediction include:

[0125] Use the text classification module of the second Bert model to calculate the feature representations of each node in the text-level graph to obtain the text classification features of each node;

[0126] The Text-level-GCN model aggregates the text classification features and uses the softmax module to obtain the GCN classification prediction y gcn .

[0127] Further, in the training module, the prediction model further includes a third Bert model, which is used to output a Bert classification prediction y bert ,

[0128] The steps for the prediction model to output a classification prediction further include:

[0129] The third Bert model processes the texts in the training text set to obtain text features, and then through the softmax module, obtains the Bert classification prediction y bert ;

[0130] Perform weighted synthesis on the Bert classification prediction y bert and the GCN classification prediction y gcn to obtain a classification prediction through weighted synthesis.

[0131] Example 5

[0132] As Figure 6 shown, the present invention also provides a text label prediction device, including:

[0133] An acquisition module, configured to acquire a text set to be predicted for labels;

[0134] A graph construction module, configured to initialize the texts in the text set to be predicted for labels by using the text tokenization and word mapping module of the first Bert model, obtain the feature representation of each word in the text set to be predicted for labels, and construct a text-level graph with the feature representation of each word as nodes, where: the feature representation of each word carries the semantic feature information of the corresponding word extracted by the first Bert model;

[0135] A prediction module, configured to input the text-level graph and the texts in the text set to be predicted for labels into a prediction model to obtain a text label prediction, where the prediction model is trained by the text label prediction model training method described in Example 1 or Example 2.

[0136] Example 6

[0137] The present invention provides a computer device, which can be a server. As Figure 7 shown, it includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor is used to provide computing and control capabilities and run the computer program. The memory includes a non-volatile storage medium, an internal memory, etc. The non-volatile storage medium stores an operating system, a computer program, a database, etc. The memory and the processor are connected through a system bus. When the processor executes the computer program, it implements the text label prediction model training method in Example 1 or Example 2, or the text label prediction method in Example 3.

[0138] Example 7

[0139] The present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the text label prediction model training method in Example 1 or Example 2, or the text label prediction method in Example 3.

[0140] Although the present invention has been described in detail in the foregoing with specific embodiments, some modifications or improvements can be made thereto based on the present invention, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.

Claims

1. A method for training a text label prediction model, characterized in that, It includes the following steps: S110. Obtain a training text set and the true labels corresponding to the training text set; S120. Initialize the texts in the training text set using the text tokenization and word mapping module of the first Bert model to obtain the feature representations of each word in the training text set, and construct a text-level graph with the feature representations of each word as nodes, where: the feature representation of each word carries the semantic feature information of the corresponding word extracted by the first Bert model; S130. Train a prediction model. The steps for training the prediction model are as follows: The prediction model outputs a classification prediction; Iteratively train the prediction model according to the classification prediction and the true labels to obtain a trained prediction model; Wherein: the prediction model includes a Text-level-BertGCN model, and the Text-level-BertGCN model includes a second Bert model and a Text-level-GCN model; The steps for the prediction model to output a classification prediction include: Use the text classification module of the second Bert model to calculate the feature representations of each node in the text-level graph to obtain the text classification features of each node; The Text-level-GCN model aggregates the text classification features and uses the softmax module to obtain the GCN classification prediction y gcn 。 2. The training method of the text label prediction model according to claim 1, characterized in that In step S130, the prediction model further includes a third Bert model; The steps for the prediction model to output a classification prediction further include: The third Bert model processes the text in the training text set to obtain text features, and then processes them through a softmax module to obtain the Bert classification prediction y bert ; Perform weighted synthesis on the Bert classification prediction y bert and the GCN classification prediction y gcn to obtain the classification prediction.

3. The method for training a text label prediction model according to claim 1 or 2, wherein In step S130, the iterative training of the prediction model according to the classification prediction and the true labels includes: Calculate the cross-entropy loss according to the classification prediction and the true labels. When the cross-entropy loss value is greater than or equal to a preset loss threshold, iteratively train the prediction model until the cross-entropy loss value is less than the loss threshold, and then stop training to obtain a trained prediction model.

4. The method for training a text label prediction model according to claim 2, wherein The weighted synthesis method is as follows: Where: is a variable parameter with a value range of [0.1, 1], used to balance the Bert classification prediction y bert and the GCN classification prediction y gcn where y is the classification prediction.

5. The training method of the text label prediction model according to claim 3, characterized in that The calculation formula for the cross-entropy loss is as follows: loss=-glogy, where: Loss is the cross-entropy loss, g is the true label, and y is the classification prediction.

6. A text label prediction method, characterized in that It includes: S210. Obtain a text set for which label prediction is to be made; S220. Initialize the texts in the text set for which label prediction is to be made using the text tokenization and word mapping module of the first Bert model to obtain the feature representations of each word in the text set for which label prediction is to be made, and construct a text-level graph with the feature representations of each word as nodes, where: the feature representation of each word carries the semantic feature information of the corresponding word extracted by the first Bert model; S230. Input the text-level graph and the texts in the text set for which label prediction is to be made into the prediction model to obtain text label predictions, and the prediction model is trained by the text label prediction model training method according to any one of claims 1 to 5.

7. A text label prediction training device, characterized in that, It includes the following modules: An acquisition module, configured to obtain a training text set and the true labels corresponding to the training text set; A composition module, configured to initialize the texts in the training text set by using the text tokenization and word mapping module of the first Bert model, obtain the feature representations of each word in the training text set, and construct a text-level graph with the feature representations of each word as nodes, where: the feature representation of each word carries the semantic feature information of the corresponding word extracted by the first Bert model; A training module, configured to train a prediction model. The steps of training the prediction model are as follows: The prediction model outputs a classification prediction; Iteratively train the prediction model according to the classification prediction and the true label to obtain a trained prediction model; Wherein: the prediction model includes a Text-level-BertGCN model, and the Text-level-BertGCN model includes a second Bert model and a Text-level-GCN model; The output of the classification prediction by the prediction model includes the following steps: Use the text classification module of the second Bert model to calculate the feature representations of each node in the text-level graph to obtain the text classification features of each node; The Text-level-GCN model aggregates the text classification features and uses the softmax module to obtain the GCN classification prediction y gcn .

8. A text label prediction device, characterized in that, Including: An acquisition module, configured to acquire a text set to be predicted for labels; A composition module, configured to initialize the texts in the text set to be predicted for labels by using the text tokenization and word mapping module of the first Bert model, obtain the feature representations of each word in the text set to be predicted for labels, and construct a text-level graph with the feature representations of each word as nodes, where: the feature representation of each word carries the semantic feature information of the corresponding word extracted by the first Bert model; A prediction module, configured to input the text-level graph and the texts in the text set to be predicted for labels into the prediction model to obtain a text label prediction, and the prediction model is trained by the text label prediction model training method according to any one of claims 1 to 5.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the text label prediction model training method according to any one of claims 1 to 5, or the text label prediction method according to claim 6.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the text label prediction model training method according to any one of claims 1 to 5, or the text label prediction method according to claim 6.