Training Method of Classification Model, Text Classification Method, Device and Equipment

By constructing sample vocabulary list and word vector matrix, using neural network model to perform multiple rounds of training, determining the target word vector and text semantic vector, and generating classification models, the problem of low text classification processing efficiency and accuracy in the existing technology is solved, and more efficient and accurate text classification is achieved.

CN115391542BActive Publication Date: 2025-06-13AGRICULTURAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211109798.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-13
Publication Date
2025-06-13
Estimated Expiration
2042-09-13

AI Technical Summary

Technical Problem

The prior art has problems with low efficiency and accuracy in text classification processing.

Method used

By constructing the sample vocabulary list and word vector matrix, using neural network model to perform multiple rounds of training, the target word vector and text semantic vector are determined, and a classification model is generated.

Benefits of technology

The processing efficiency and accuracy of the classification model are improved, by retaining the target word vectors that affect text classification and removing redundant data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115391542B_ABST
    Figure CN115391542B_ABST
Patent Text Reader

Abstract

The present application provides a training method for a classification model, a text classification method, an apparatus, and a device. The training method for the classification model includes: an electronic device constructs a sample vocabulary based on a plurality of sample texts, generates a word vector matrix according to the sample vocabulary, and performs at least one round of training on a neural network model based on at least one sample text and the word vector matrix to obtain a classification model. Among them, any round of the training process includes: determining a plurality of target word vectors corresponding to each class label vector according to the neural network model, the sample text, and the word vector matrix obtained in the previous round of training, and determining a text semantic vector of the sample text according to the neural network model, each class label vector, and the corresponding plurality of target word vectors obtained in the previous round of training, and generating the neural network model obtained in this round of training based on the text semantic vector. In the technical solution, the processing efficiency and accuracy of the trained classification model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and particularly relates to a training method for a classification model, a text classification method, an apparatus, and a device. Background Art

[0002] With the explosive growth of information, a large number of texts such as news, announcements, and technical blogs are uploaded to information publishing platforms every day for users to browse and download.

[0003] In order to facilitate users to search for texts, texts are usually classified according to their themes so that users can browse texts under different theme modules according to their own purposes. Currently, text classification mainly splits a text into multiple sentences, and respectively uses an encoder to obtain corresponding encoding representations for each word in each sentence. Then, according to the encoding representations of the words, a hierarchical attention network is used to obtain the feature vector of the text. Finally, according to the feature vector of the text, a multi-layer perceptron is used to output the classification result of the text.

[0004] However, the prior art has problems of low efficiency and accuracy in classification processing. Summary of the Invention

[0005] The present application provides a training method for a classification model, a text classification method, an apparatus, and a device to solve the problems of low efficiency and accuracy in classification processing existing in the prior art.

[0006] In a first aspect, an embodiment of the present application provides a training method for a classification model, including:

[0007] Construct a sample vocabulary according to a plurality of sample texts, each sample text including a plurality of sample sub-texts obtained through word segmentation processing and at least one class label, and the sample vocabulary including the plurality of sample sub-texts and the plurality of class labels;

[0008] Generate a word vector matrix according to the sample vocabulary, the word vector matrix including word vectors respectively corresponding to the sample sub-texts in the sample vocabulary and class label vectors respectively corresponding to the class labels;

[0009] Perform at least one round of training on a neural network model according to at least one sample text and the word vector matrix to obtain a classification model; any round of training process includes: determining a plurality of target word vectors corresponding to the class label vectors according to the neural network model obtained in the previous round of training, the sample text, and the word vector matrix, and determining the text semantic vector of the sample text according to the neural network model obtained in the previous round of training, the class label vectors, and the corresponding plurality of target word vectors, and generating the neural network model obtained in this round of training based on the text semantic vector.

[0010] In a possible design of the first aspect, determining multiple target word vectors corresponding to each type of label vector based on the neural network model obtained from the previous round of training, the sample text, and the word vector matrix includes:

[0011] Determine the word vectors of the sample text according to the neural network model obtained from the previous round of training, the sample text, and the word vector matrix;

[0012] For any type of label vector, calculate the correlation between each word vector in the sample text and the label vector respectively;

[0013] Determine the first preset number of word vectors with the maximum correlation as the multiple target word vectors corresponding to the label vector.

[0014] In another possible design of the first aspect, determining the text semantic vector of the sample text according to the neural network model obtained from the previous round of training, each type of label vector, and the corresponding multiple target word vectors includes:

[0015] For the multiple target word vectors corresponding to any type of label vector, sort the multiple target word vectors according to the neural network model obtained from the previous round of training and the positions of the sample sub-texts corresponding to each target word vector in the sample text to determine the target word sequence vector;

[0016] Determine the sample sub-text semantic vector corresponding to the label vector according to the target word sequence vector;

[0017] Average the sample sub-text semantic vectors corresponding to each type of label vector to determine the text semantic vector of the sample text.

[0018] Optionally, the step of sorting the multiple target word vectors corresponding to any type of label vector according to the neural network model obtained from the previous round of training and the positions of the sample sub-texts corresponding to each target word vector in the sample text to determine the target word sequence vector includes:

[0019] For the multiple target word vectors corresponding to any type of label vector, sort the multiple target word vectors according to the neural network model obtained from the previous round of training and the positions of the sample sub-texts corresponding to each target word vector in the sample text to determine the initial word sequence vector;

[0020] Construct a window according to the distance between the sample sub-text corresponding to each target word vector and other sample sub-texts in the sample text and the second preset number. The window includes the target word vector and the second preset number of word vectors closest to the sample sub-text corresponding to the target word vector;

[0021] Calculate the average vector of each vector within each window, and construct a sample sub-text sequence corresponding to the target word vector according to the average vector of each window.

[0022] Optionally, determining the sample sub-text semantic vector corresponding to the class label vector according to the target word sequence vector includes:

[0023] Extract the text features of the target word sequence vector;

[0024] Determine the semantic vectors of each target word vector in the target word sequence vector according to the Self-Attention mechanism and the text features;

[0025] Determine the sample sub-text semantic vector corresponding to the class label vector according to the semantic vectors of each target word vector and the contribution degree of the target word vector to the class label vector.

[0026] In another possible design of the first aspect, the sample word table further includes each sample sub-text serial number and class label serial number; the word vector matrix further includes the sample sub-text serial number corresponding to each word vector and the class label serial number corresponding to each class label vector. Training the neural network model for at least one round according to at least one sample text and the word vector matrix to obtain a classification model includes:

[0027] Train the neural network model for at least one round according to the multiple sample sub-text serial numbers of at least one sample text and at least one class label serial number to obtain the classification model.

[0028] In a second aspect, an embodiment of the present application provides a text classification method, including:

[0029] Construct a word table according to multiple texts to be classified. The texts to be classified include multiple sub-texts obtained by word segmentation processing. The word table includes multiple sub-texts and pre-set category information;

[0030] Input at least one file to be classified and the word table into the classification model, and obtain the classification result of the file to be classified. The classification model is trained by using the method described in any item of the first aspect.

[0031] In a third aspect, an embodiment of the present application provides a training device for a classification model, including:

[0032] A construction module, configured to construct a sample word table according to multiple sample texts. Each sample text includes multiple sample sub-texts obtained by word segmentation processing and at least one class label. The sample word table includes multiple sample sub-texts and multiple class labels;

[0033] A generation module, configured to generate a word vector matrix according to the sample word list, where the word vector matrix includes word vectors corresponding to respective sample sub-texts in the sample word list and class label vectors corresponding to respective class labels;

[0034] A training module, configured to perform at least one round of training on a neural network model according to at least one sample text and the word vector matrix to obtain a classification model; Any round of training process includes: determining a plurality of target word vectors corresponding to respective class label vectors according to the neural network model obtained in the previous round of training, the sample text, and the word vector matrix, and determining a text semantic vector of the sample text according to the neural network model obtained in the previous round of training, respective class label vectors, and the corresponding plurality of target word vectors, and generating the neural network model obtained in this round of training based on the text semantic vector.

[0035] In a possible design of the third aspect, the training module is specifically configured to:

[0036] Determine word vectors of the sample text according to the neural network model obtained in the previous round of training, the sample text, and the word vector matrix;

[0037] For any class label vector, calculate the relevance between each word vector in the sample text and the class label vector respectively;

[0038] Determine the first preset number of word vectors with the maximum relevance as the plurality of target word vectors corresponding to the class label vector.

[0039] In another possible design of the third aspect, the training module is specifically configured to:

[0040] For the plurality of target word vectors corresponding to any class label vector, sort the plurality of target word vectors according to the neural network model obtained in the previous round of training and the positions where the sample sub-texts corresponding to the respective target word vectors appear in the sample text to determine a target word sequence vector;

[0041] Determine a sample sub-text semantic vector corresponding to the class label vector according to the target word sequence vector;

[0042] Average the sample sub-text semantic vectors corresponding to respective class label vectors to determine the text semantic vector of the sample text.

[0043] Optionally, the training module is specifically configured to, for the plurality of target word vectors corresponding to any class label vector, sort the plurality of target word vectors according to the neural network model obtained in the previous round of training and the positions where the sample sub-texts corresponding to the respective target word vectors appear in the sample text to determine an initial word sequence vector;

[0044] Construct a window according to the distance between the sample sub-text corresponding to each target word vector and other sample sub-texts in the sample text and a second preset quantity, where the window includes the target word vector and the second preset quantity of word vectors that are the closest to the sample sub-text corresponding to the target word vector;

[0045] Calculate the average vector of each vector within each window, and construct the sample sub-text sequence corresponding to the target word vector according to the average vector of each window.

[0046] Optionally, the training module is specifically configured to:

[0047] Extract the text features of the target word sequence vector;

[0048] Determine the semantic vectors of each target word vector in the target word sequence vector according to the Self-Attention mechanism and the text features;

[0049] Determine the sample sub-text semantic vector corresponding to the class label vector according to the semantic vectors of each target word vector and the contribution degree of the target word vector to the class label vector.

[0050] In yet another possible design of the third aspect, the sample word table further includes each sample sub-text serial number and class label serial number; the word vector matrix further includes the sample sub-text serial number corresponding to each word vector and the class label serial number corresponding to each class label vector, and the training module is specifically configured to:

[0051] Perform at least one round of training on the neural network model according to the multiple sample sub-text serial numbers of at least one sample text and at least one class label serial number to obtain the classification model.

[0052] In a fourth aspect, an embodiment of the present application provides a text classification device, including:

[0053] A construction module, configured to construct a word table according to a plurality of texts to be classified, where the texts to be classified include a plurality of sub-texts obtained by word segmentation processing, and the word table includes a plurality of sub-texts and preset category information;

[0054] An input module, configured to input at least one file to be classified and the word table into the classification model to obtain the classification result of the file to be classified, where the classification model is trained by using the method described in any item of the first aspect.

[0055] In a fifth aspect, an embodiment of the present application provides an electronic device, including: a processor, a memory, and computer program instructions stored on the memory and executable on the processor, where the processor is configured to implement the methods provided by the first aspect, the second aspect, and each possible design when executing the computer program instructions.

[0056] The training method, text classification method, device and equipment of the classification model provided by the embodiments of the present application. In the training method of the classification model, the electronic device constructs a sample vocabulary according to multiple sample texts, generates a word vector matrix according to the sample vocabulary, and performs at least one round of training on the neural network model according to at least one sample text and the word vector matrix to obtain the classification model. Among them, any round of training process includes: determining multiple target word vectors corresponding to each class vector according to the neural network model, sample text and word vector matrix obtained in the previous round of training, and determining the text semantic vector of the sample text according to the neural network model, each class vector and the corresponding multiple target word vectors obtained in the previous round of training, and generating the neural network model obtained in this round of training based on the text semantic vector. In this technical solution, the target word vectors that affect text classification are retained during the model training process, and redundant data is removed, thereby improving the processing efficiency and accuracy of the trained classification model. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0058] Figure 1 It is a schematic diagram of the scenario of the training method of the classification model and the text classification method provided by the embodiments of the present application;

[0059] Figure 2 It is a schematic flowchart of Embodiment 1 of the training method of the classification model provided by the embodiments of the present application;

[0060] Figure 3 It is a schematic flowchart of Embodiment 2 of the training method of the classification model provided by the embodiments of the present application;

[0061] Figure 4 It is a schematic flowchart of Embodiment 1 of the text classification method provided by the embodiments of the present application;

[0062] Figure 5 It is a schematic structural diagram of the training device of the classification model provided by the embodiments of the present application;

[0063] Figure 6 It is a schematic structural diagram of the text classification device provided by the embodiments of the present application;

[0064] Figure 7 It is a schematic structural diagram of the electronic device provided by the embodiments of the present application.

[0065] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and more detailed descriptions will be provided hereinafter. These drawings and written descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Embodiments

[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0067] Before introducing the embodiments of the present application, the application background related to the embodiments of the present application will be introduced first:

[0068] When a user needs to share information with others, the user can register and log in to an information publishing platform and upload a text carrying the sharing information in the text format specified by the information publishing platform. Similarly, when a user wants to obtain information, the user can also check whether there is information uploaded by other users that they need in the information publishing platform. However, since there are a large number of people using the information publishing platform, the types of uploaded texts are diverse. Therefore, introducing natural language processing technology to intelligently classify and display the texts on the platform can greatly save labor costs and improve the efficiency of users' indexing information.

[0069] Currently, the classification of texts through natural language processing technology is mainly carried out in the following ways:

[0070] (1) Classifying texts through a deep learning model. However, due to the rich and complex semantic information of long texts, and the fact that they may contain multiple turns, metaphors, and information irrelevant to the theme, it is easy to cause the deep learning model to focus too much on local texts and unable to make accurate judgments on the whole, increasing the difficulty for the deep learning model to extract complete text semantic information. At the same time, in the deep learning model, there will also be phenomena such as vanishing gradients or exploding gradients caused by overly long texts, making it difficult for model training to converge.

[0071] Among them, vanishing gradients refer to the phenomenon that when using a deep learning model to train based on a long text, the gradient values generated by distant errors become 0 due to multiple consecutive multiplications; exploding gradients refer to the phenomenon that when using a deep learning model to train based on a long text, the gradient values generated by distant errors become extremely large due to multiple consecutive multiplications.

[0072] (2) For texts with relatively long lengths such as news and blogs, the first paragraph of the text is intercepted as the representative of the text for classification processing. However, since news and blogs are both written around a theme and have a clear hierarchical logical structure, they often belong to more than one theme category. Therefore, if the text is classified only based on the first paragraph of the text, information may be omitted, and it may not be possible to accurately determine the category to which the text belongs, resulting in the absence of texts under a certain theme model.

[0073] (3) The text is split into multiple sentences, and the encoder is used to obtain the corresponding encoded representations for each word in each sentence. Then, based on the encoded representations of the words, a hierarchical attention network is adopted to obtain the feature vector of the text. Finally, based on the feature vector of the text, a multi-layer perceptron is used to output the classification result of the text. However, when splitting the text, only the sentence length is considered, and a lot of redundant information in the text is still retained.

[0074] In summary, the existing methods for classifying texts have problems of low processing efficiency and accuracy.

[0075] Based on the above problems, the present application proposes a training method for a classification model, which can construct a training set and a word vector matrix according to multiple sample texts. The training set includes at least one sample text, and the word vector matrix includes the word vectors and class label vectors of each sample text. During any round of model training according to the training set and the word vector matrix, the neural network model obtained from the previous round of training can determine multiple target word vectors through the sample texts in the training set and the input word vector matrix, and generate the neural network model obtained from the current round of training based on the sample sub-texts corresponding to the target word vectors, so as to obtain a classification model generated through at least one round of training. Compared with the prior art, the target word vectors affecting text classification are retained during the model training process, and redundant data is removed, thereby improving the processing efficiency and accuracy of the trained classification model.

[0076] Exemplarily, the training method for the classification model and the text classification method provided in the embodiments of the present application can be applied to Figure 1 the scene schematic diagram shown. Figure 1 This is the scene schematic diagram of the training method for the classification model and the text classification method provided in the embodiments of the present application. As Figure 1 shown, this scene includes a sample text database 11, an electronic device 12, and an information publishing platform 13.

[0077] In this embodiment, the sample text database 11 includes a plurality of sample texts for training the classification model, and the electronic device 12 is used to train the neural network model according to the plurality of sample texts in the sample text database 11 to obtain the classification model. Among them, the neural network model can be pre-stored in the electronic device 12 by relevant staff, or can be obtained by the electronic device 12 from the network or other databases storing the neural network model. The embodiments of the present application do not make specific limitations on this.

[0078] In practical applications, the electronic device 12 can obtain a plurality of texts to be classified from the information publishing platform 13, input the plurality of texts to be classified into the above classification model for classification processing, and thus return the classification results of the plurality of texts to be classified to the information publishing platform 13.

[0079] It should be noted that Figure 1 is only a schematic diagram of an application scenario provided by the embodiments of the present application. The embodiments of the present application do not Figure 1 limit the devices included therein, nor do they Figure 1 limit the positional relationship between the devices therein. For example, in Figure 1 , the sample text database 11 can be an external memory relative to the information publishing platform 13. In other cases, the sample text database 11 can also be placed in the information publishing platform 13.

[0080] Optionally, the logical functions of the electronic device 12 can be integrated on the same physical device or on different physical devices, which can be determined according to the actual situation. The embodiments of the present application do not make specific limitations on this.

[0081] It can be understood that the execution subject of the embodiments of the present application can be a terminal device, such as a computer, a tablet computer, etc., or a server, such as a background processing platform, etc. Therefore, in this embodiment, the terminal device and the server are collectively referred to as the electronic device for explanation. Regarding whether the electronic device is specifically a terminal device or a server, it can be determined according to the actual situation.

[0082] Next, the technical solution of the present application will be described in detail through specific embodiments.

[0083] It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0084] Figure 2 is a schematic flowchart of the first embodiment of the training method of the classification model provided by the embodiments of the present application. As Figure 2 shown, the training method of the classification model may include the following steps:

[0085] S201. Construct a sample vocabulary based on multiple sample texts.

[0086] Among them, each sample text includes multiple sample sub - texts obtained through word segmentation processing and at least one class label. The sample vocabulary includes multiple sample sub - texts and multiple class labels.

[0087] Exemplarily, assume the sample text is "Today the weather is very good. The teacher organized everyone to go on an autumn outing,...", then the multiple sample sub - texts obtained through word segmentation processing are respectively "Today, weather, very good, teacher, organize, everyone, together, go, autumn outing,...". It should be understood that existing methods can be used to perform word segmentation processing on the sample text, and this application does not limit the specific word segmentation processing method.

[0088] Among them, the class label is used to indicate the category to which the sample text belongs. For example, in the above example, the class label of this sample text can be a travel note. Optionally, the class label can be obtained by manually annotating the sample text.

[0089] In practical applications, since a sample text can belong to different categories, the number of class labels for each sample text can be one or multiple, which can be determined according to the actual situation. The embodiments of this application do not specifically limit the number of class labels for each sample text.

[0090] Exemplarily, the sample vocabulary can be represented by Table 1.

[0091] Table 1

[0092] Serial number Content Type 1 Sample sub-text 1 Sample sub-text 2 Class label 1 Class label 3 Sample sub-text 2 Sample sub-text

[0093] Exemplarily, the sample vocabulary contains a total of 2 sample sub - texts and 1 class label. Among them, the 2 sample sub - texts are respectively sample sub - text 1 and sample sub - text 2, and the 1 class label is class label 1.

[0094] S202. Generate a word vector matrix according to the sample vocabulary.

[0095] Optionally, the sample sub - texts in the sample vocabulary can be converted into word vectors, and the class labels can be converted into class label vectors, so as to generate a word vector matrix.

[0096] Among them, the above conversion processing can be implemented through pre - trained Word2Vec. When converting the class label into a class label vector, for class labels with specific meanings, the conversion processing can be directly implemented through Word2Vec, while for class labels without specific meanings, they are set as random vectors with the same dimension as this type. It should be understood that class labels without specific meanings can be class labels represented by characters such as "1", "2", "3", etc. Similar to class labels with specific meanings, each sample text with a class label without specific meaning belongs to the same category.

[0097] It should be understood that in the actual application process, other methods can also be used to implement the above conversion process, which can be specifically determined according to the actual situation, and the embodiments of the present application do not specifically limit this.

[0098] Among them, the word vector matrix includes word vectors corresponding to each sample sub-text in the sample word list and class label vectors corresponding to each class label.

[0099] S203. According to at least one sample text and the word vector matrix, perform at least one round of training on the neural network model to obtain a classification model.

[0100] Optionally, the above at least one sample text is a training set required for model training.

[0101] Optionally, the above neural network model can be pre-stored in an electronic device by relevant staff, or can be obtained from other databases storing the neural network model, which can be determined according to the actual situation, and the present application does not specifically limit the method of obtaining the neural network model.

[0102] Among them, any round of training process includes: determining multiple target word vectors corresponding to each class label vector according to the neural network model, sample text, and class label vectors obtained in the previous round of training, and determining the text semantic vector of the sample text according to the neural network model, class label vectors, and corresponding multiple target word vectors obtained in the previous round of training, and generating the neural network model obtained in this round of training based on the text semantic vector.

[0103] In the training method of the classification model provided by the embodiments of the present application, the electronic device constructs a sample word list according to multiple sample texts, generates a word vector matrix according to the sample word list, and performs at least one round of training on the neural network model according to at least one sample text and the word vector matrix to obtain a classification model. Among them, any round of training process includes: determining multiple target word vectors corresponding to each class label vector according to the neural network model, sample text, and word vector matrix obtained in the previous round of training, and determining the text semantic vector of the sample text according to the neural network model, class label vectors, and corresponding multiple target word vectors obtained in the previous round of training, and generating the neural network model obtained in this round of training based on the text semantic vector. In this technical solution, the target word vectors affecting text classification are retained during the model training process, redundant data is removed, the length of the text is effectively shortened, the purpose of long text recognition is achieved, and the processing efficiency and accuracy of the trained classification model are improved.

[0104] Optionally, based on Figure 2In the illustrated embodiment, to determine multiple target word vectors corresponding to various label vectors based on the neural network model, sample text, and word vector matrix obtained from the previous round of training, the following steps may be implemented:

[0105] Step 1: Determine the word vectors of the sample text based on the neural network model, sample text, and word vector matrix obtained from the previous round of training.

[0106] Optionally, through the neural network model obtained from the previous round of training, according to the word vector matrix, multiple sample sub-texts of each sample text in the training set are respectively converted into word vectors, and at least one label of the above sample texts is converted into a label vector.

[0107] Step 2: For any label vector, calculate the correlation between each word vector in the sample text and the label vector.

[0108] Optionally, the correlation between the word vector and the label vector can be calculated by the following formula:

[0109]

[0110]

[0111] Where C m represents the label vector of the m-th class, E t is the t-th word vector. By performing an inner product operation on C m and E t , the correlation between C m and E t is obtained

[0112] Where, in formula (1), represents the result obtained by mapping the label vector of the M-th class to the Query space. Similarly, that is, mapping the word vector E t to the Key vector space. The specific mapping process is shown in formula (2). In formula (2), W Q is the weight of the neural network corresponding to Query, b Q is the bias of the neural network corresponding to Query; W K is the weight of the neural network corresponding to Key, b K is the bias of the neural network corresponding to Key.

[0113]

[0114] Optionally, I M is the set of correlations.

[0115] It should be understood that the correlation between the word vector and the class label vector can also be calculated by other existing methods, and the embodiments of the present application do not specifically limit this.

[0116] Step 3: Determine the first preset number of word vectors with the highest correlation as the multiple target word vectors corresponding to the class label vector.

[0117] Exemplarily, the preset number can be 3, 4, 5, etc., which can be determined in advance according to the actual situation, and the embodiments of the present application do not specifically limit this.

[0118]

[0119] Among them, is the target word vector.

[0120] In the above embodiment, by calculating the correlation between each word vector in the sample text of the training set and the class label vector, the target word vector corresponding to the class label vector is obtained, reducing the redundant data in the training set, and improving the accuracy and efficiency of subsequent model training.

[0121] Optionally, in some embodiments, the above-mentioned determining the text semantic vector of the sample text according to the neural network model, various class label vectors, and the corresponding multiple target word vectors obtained in the previous round of training can be implemented through the following steps:

[0122] Step 1: For the multiple target word vectors corresponding to any class label vector, according to the neural network model obtained in the previous round of training, and the positions of the sample sub-texts corresponding to each target word vector in the sample text, sort the multiple target word vectors to determine the target word sequence vector.

[0123] Exemplarily, assume that the positions of the sample sub-texts corresponding to each target word vector in the sample text are respectively: target word vector 1, the 32nd sample sub-text in the sample text; target word vector 2, the 25th sample sub-text in the sample text; target word vector 3, the 102nd sample sub-text in the sample text; target word vector 4, the 128th sample sub-text in the sample text. Then, sort the above 4 target word vectors, and the target word sequence vector is target word vector 2 - target word vector 1 - target word vector 3 - target word vector 4.

[0124] Step 2: Determine the sample sub-text semantic vector corresponding to the class label vector according to the target word sequence vector.

[0125] Optionally, the text features of the target word sequence vector can be extracted through the bidirectional long short-term memory network (LSTM) in the neural network model obtained from the previous round of training. The semantic vectors of each target word vector in the target word sequence vector are calculated based on the above text features, and the sample sub-text semantic vector corresponding to the class label vector is determined according to the semantic vectors of each target word vector and the contribution degree of the target word vector to the class label vector.

[0126] Step 3. Average the sample sub-text semantic vectors corresponding to each class label vector to determine the text semantic vector of the sample text.

[0127] Optionally, it can be implemented through the following formula:

[0128]

[0129] Optionally, L m is the sample sub-text semantic vector of the class label vector of the m-th class, and la_wsn vec is the text semantic vector of the sample text.

[0130] Optionally, la_wsn vec can be sent to the output layer of the neural network model obtained from the previous round of training, and the neural network model obtained from this round of training is generated through the softmax function.

[0131] In the above embodiments, according to the sample sub-text semantic vector of each class label vector, the text semantic vector of the sample text is determined, so as to achieve the purpose of generating the neural network model obtained from this round of training according to the text semantic vector.

[0132] Optionally, in some embodiments, for the multiple target word vectors corresponding to any class label vector, according to the neural network model obtained from the previous round of training and the positions of the sample sub-texts corresponding to each target word vector in the sample text, the multiple target word vectors are sorted to determine the target word sequence vector, which can be implemented through the following steps:

[0133] Step 1. For the multiple target word vectors corresponding to any class label vector, according to the neural network model obtained from the previous round of training and the positions of the sample sub-texts corresponding to each target word vector in the sample text, the multiple target word vectors are sorted to determine the initial word sequence vector.

[0134] Step 2. A window is constructed according to the distance between the sample sub-text corresponding to each target word vector and other sample sub-texts in the sample text and the second preset quantity.

[0135] Among them, the window includes the target word vector and the second preset number of word vectors that are closest to the sample sub-text corresponding to the target word vector.

[0136] Exemplarily, the second preset number can be 3, 4, 5, etc., which can be determined in advance according to the actual situation, and the embodiments of the present application do not specifically limit this.

[0137] Exemplarily, for any target word vector, taking the sample sub-text corresponding to the target word vector as the center, the second preset number of word vectors that are closest to it can be determined from the above-mentioned upper and lower texts of the sample sub-text in the sample text respectively, and a window with a size of 2×the second preset number + 1 is constructed.

[0138] Step three: Calculate the average vector of each vector in each window, and construct a target word sequence vector according to the average vector of each window.

[0139] Optionally, it can be implemented through the following formula:

[0140]

[0141] Among them, r is the second preset number, is the average vector of this window.

[0142] Among them, constructing a target word sequence vector according to the average vector of each window can be achieved through the following steps:

[0143] Replace the target word vector corresponding to the initial word sequence vector with the average vector of each window, so as to generate a target word sequence vector.

[0144] In the above embodiment, a window is constructed with the target word vector as the center, and the context information and the semantic information of the target word vector are compressed into a new vector with the size of a word vector (that is, the average vector) by averaging the vectors in the window. Compared with the traditional machine learning model based on statistics in the prior art, which represents the text as a high-dimensional vector using the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm based on word frequency statistical information, the technical solution of the present invention effectively retains the context information of the keywords, solves the problem of losing context information caused by non-continuous words in the prior art, and improves the processing accuracy of the classification model. Among them, TF-IDF can represent a piece of text as a high-dimensional vector based on the word frequency information of the entire text statistics.

[0145] Optionally, in some embodiments, the above-mentioned determining the sample sub-text semantic vector corresponding to the class label vector according to the target word sequence vector can be achieved according to the following steps:

[0146] Step 1: Extract the text features of the target word sequence vectors.

[0147] The text features of the target word sequence vectors can be extracted by the bidirectional LSTM in the neural network model obtained from the previous round of training, and the hidden state of the LSTM at each sequence position t is obtained. Among them, are the hidden states of the forward LSTM and the backward LSTM at position t in the target word sequence vectors.

[0148] Step 2: Determine the semantic vectors of each target word vector in the target word sequence vector according to the self-attention Self-Attention mechanism and the text features.

[0149] Calculate Self-Attention on the obtained sequence to determine the semantic vectors of each target word vector in the target word sequence vector, so as to further strengthen the understanding of the overall context semantic information of the sample text.

[0150] Step 3: Determine the sample sub-text semantic vector corresponding to the class label vector according to the semantic vectors of each target word vector and the contribution degree of the target word vector to the class label vector.

[0151] Optionally, it can be implemented through the following steps:

[0152]

[0153] Among them, d is the dimension size of the class label vector, which is used to prevent the value obtained by the inner product operation from being too large, resulting in too small gradient value of the softmax function. The meaning of its calculation process is to calculate the contribution degree of each obtained target word vector to class M and the semantic vector obtained through Self-Attention are weighted and summed to obtain the sample sub-text semantic vector L that integrates the information of class M M .

[0154] In the above embodiment, by the semantic vectors of each target word vector and the contribution degree of the target word vector to the class label vector, the sample sub-text semantic vector integrating the class label vector is determined, which strengthens the understanding ability of the classification model for the text semantic information, and each sample text uses all class label vectors in the word vector matrix, so that the classification model can learn from multiple theme directions for the same sample text, and finally integrates the semantic information of multiple different categories, enriching the information contained in a single sample text and enhancing the generalization ability of the classification model.

[0155] Optionally, in some embodiments, the sample vocabulary further includes the serial numbers of each sample sub-text and the serial number of the class label; the word vector matrix further includes the serial numbers of the sample sub-texts corresponding to each word vector and the serial numbers of the class labels corresponding to each class label vector. The above-mentioned at least one round of training of the neural network model based on at least one sample text and the word vector matrix to obtain a classification model can be implemented through the following steps:

[0156] Perform at least one round of training on the neural network model according to the serial numbers of multiple sample sub-texts of at least one sample text and the serial number of at least one class label to obtain a classification model.

[0157] That is to say, in this embodiment, the training set is the serial numbers of multiple sample sub-texts of at least one sample text and the serial number of at least one class label.

[0158] In this embodiment, using the above training set for model training can effectively reduce the data volume of the training set, thereby improving the training efficiency.

[0159] Combined with the training methods of the classification models in the above various embodiments, the following uses a specific example to illustrate this method.

[0160] Figure 3 It is a schematic flowchart of the second embodiment of the training method of the classification model provided by the embodiment of the present application. As Figure 3 shown, taking any round of training process as an example, the serial numbers of multiple sample sub-texts of the sample text (x 1 、x 2 、……、x r ) can be input into the neural network model obtained from the previous round of training. The neural network model obtained from the previous round of training extracts the text features of the target word sequence vector through a bidirectional LSTM. According to the class label vectors (C 1 、C 2 、……、C M-2 、C M-1 、C M ) in the sample vocabulary, the word vectors (E 1 、E 2 、……、E T ) corresponding to multiple sample sub-text serial numbers are calculated respectively.

[0161] Further, taking C M as an example, calculate the correlation between each word vector in the sample text and C M And determine the first preset number of word vectors (Top-N) with the largest correlation as the multiple target word vectors corresponding to the class label vector C M ​​Next, according to the distance between the sample sub-text corresponding to each target word vector and other sample sub-texts in the sample text and the second preset quantity, a window is constructed, the average vector of each vector within each window is calculated, and the target word sequence vector is constructed according to the average vector of each window Extract the text features of the target word sequence vector through a bidirectional LSTM, and the text features include the hidden state of each sequence position Then, according to the Self-Attention mechanism, determine the semantic vectors of each target word vector in the target word sequence vector And according to the semantic vectors of each target word vector and the contribution degree of the target word vector pair to the class label vector C M Degree of contribution Determine the class label vector C M The corresponding sample sub-text semantic vector (L M ).

[0162] According to the above steps, calculate the sample sub-text semantic vectors (L 1 , L 2 , ……, L M-2 , L M-1 ) of other class label vectors respectively, and calculate the average value of the sample sub-text semantic vectors of each class label vector, so as to obtain the text semantic vector (la_wsn vec ), and send la_wsn vec to the output layer, and generate the neural network model obtained by this round of training through the softmax function

[0163] After obtaining the above classification model, the classification model can be used to classify multiple texts to be classified. The following will specifically describe in detail the method for classifying texts to be classified using this classification model. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments

[0164] Specifically, the execution subject of this text classification method can also be an electronic device with processing capabilities such as a terminal or a server. It should be understood that the electronic device executing this text classification method and the electronic device executing the training method of the above classification model can be the same device or different devices

[0165] Figure 4 It is a schematic flowchart of the first embodiment of the text classification method provided by the embodiment of the present application. As Figure 4 shown, this text classification method may include the following steps

[0166] S401. Construct a word list according to multiple texts to be classified

[0167] The text to be classified includes multiple sub-texts obtained through word segmentation processing, and the vocabulary includes multiple sub-texts and pre-set category information.

[0168] Among them, the category information can be pre-set manually. For example, assuming that the information publishing platform is manually set with ten categories, the category information in the vocabulary is the category information of the above ten classifications.

[0169] It should be understood that the present application does not limit the method used for word segmentation processing, which can be determined according to actual needs.

[0170] S402. Input at least one file to be classified and the vocabulary into a classification model to obtain the classification result of the file to be classified.

[0171] The classification model is trained by using the training method of the classification model shown in any of the above embodiments.

[0172] Among them, the above classification result includes at least one class label of the file to be classified.

[0173] The embodiment of the present application provides a text classification method. An electronic device constructs a vocabulary according to multiple texts to be classified, and inputs at least one file to be classified and the vocabulary into a classification model to obtain the classification result of the file to be classified. Among them, the text to be classified includes multiple sub-texts obtained through word segmentation processing, and the vocabulary includes multiple sub-texts and pre-set category information. The classification model is a model established through a neural network. Compared with the prior art, the classification model can effectively shorten the length of the text to be recognized, so as to achieve the purpose of recognizing long texts. At the same time, the classification model can also remove redundant data irrelevant to the classification process, thereby improving the efficiency and accuracy of the classification process.

[0174] The following is the device embodiment of the present application, which can be used to execute the method embodiment of the present application. For the details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.

[0175] Figure 5 It is a schematic structural diagram of a training device for the classification model provided by the embodiment of the present application. As Figure 5 shown, the training device for the classification model includes:

[0176] A construction module 51, configured to construct a sample vocabulary according to multiple sample texts. Each sample text includes multiple sample sub-texts obtained through word segmentation processing and at least one class label, and the sample vocabulary includes multiple sample sub-texts and multiple class labels;

[0177] A generation module 52, configured to generate a word vector matrix according to the sample vocabulary. The word vector matrix includes word vectors corresponding to each sample sub-text in the sample vocabulary and class label vectors corresponding to each class label;

[0178] A training module 53, configured to perform at least one round of training on a neural network model according to at least one sample text and a word vector matrix to obtain a classification model; any round of training process includes: determining a plurality of target word vectors corresponding to each class label vector according to the neural network model, the sample text, and the word vector matrix obtained in the previous round of training, and determining a text semantic vector of the sample text according to the neural network model, each class label vector, and the corresponding plurality of target word vectors obtained in the previous round of training, and generating the neural network model obtained in this round of training based on the text semantic vector.

[0179] In a possible design of an embodiment of the present application, the training module 53 is specifically configured to:

[0180] Determine each word vector of the sample text according to the neural network model, the sample text, and the word vector matrix obtained in the previous round of training;

[0181] For any class label vector, calculate the correlation degree between each word vector in the sample text and the class label vector respectively;

[0182] Determine the first preset number of word vectors with the largest correlation degree as the plurality of target word vectors corresponding to the class label vector.

[0183] In another possible design of an embodiment of the present application, the training module 53 is specifically configured to:

[0184] For the plurality of target word vectors corresponding to any class label vector, sort the plurality of target word vectors according to the neural network model obtained in the previous round of training and the positions where the sample sub-texts corresponding to each target word vector appear in the sample text, and determine a target word sequence vector;

[0185] Determine a sample sub-text semantic vector corresponding to the class label vector according to the target word sequence vector;

[0186] Average the sample sub-text semantic vectors corresponding to each class label vector to determine the text semantic vector of the sample text.

[0187] Optionally, the training module 53 is specifically configured to, for the plurality of target word vectors corresponding to any class label vector, sort the plurality of target word vectors according to the neural network model obtained in the previous round of training and the positions where the sample sub-texts corresponding to each target word vector appear in the sample text, and determine an initial word sequence vector;

[0188] Construct a window according to the distance between the sample sub-text corresponding to each target word vector and other sample sub-texts in the sample text and a second preset number, where the window includes the target word vector and the second preset number of word vectors closest to the sample sub-text corresponding to the target word vector;

[0189] Calculate the average vector of each vector within each window, and construct a sample sub-text sequence corresponding to the target word vector according to the average vector of each window.

[0190] Optionally, the training module 53 is specifically configured to:

[0191] Extract the text features of the target word sequence vector;

[0192] Determine the semantic vectors of each target word vector in the target word sequence vector according to the Self-Attention mechanism and the text features;

[0193] Determine the sample sub-text semantic vector corresponding to the class label vector according to the semantic vectors of each target word vector and the contribution degree of the target word vector to the class label vector.

[0194] In another possible design of the embodiment of the present application, the sample word list further includes each sample sub-text serial number and class label serial number; the word vector matrix further includes the sample sub-text serial number corresponding to each word vector and the class label serial number corresponding to each class label vector. The training module 53 is specifically configured to:

[0195] Perform at least one round of training on the neural network model according to the multiple sample sub-text serial numbers of at least one sample text and at least one class label serial number to obtain a classification model.

[0196] The training device for the classification model provided by the embodiment of the present application can be used to execute the classification model training method in any of the above embodiments. The implementation principle and technical effect are similar and will not be elaborated here.

[0197] Figure 6 It is a schematic structural diagram of the text classification device provided by the embodiment of the present application. As Figure 6 shown, the text classification device includes:

[0198] A construction module 61, configured to construct a word list according to multiple texts to be classified. The texts to be classified include multiple sub-texts obtained through word segmentation processing. The word list includes multiple sub-texts and preset category information;

[0199] An input module 62, configured to input at least one file to be classified and the word list into the classification model to obtain the classification result of the file to be classified. The classification model is trained by using the method in any item of the first aspect.

[0200] The text classification device provided by the embodiment of the present application can be used to execute the text classification method in any of the above embodiments. The implementation principle and technical effect are similar and will not be elaborated here.

[0201] It should be noted that it should be understood that the division of each module of the above device is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. In addition, all or part of these modules can be integrated together or can be independently implemented. Here, the processing element can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or by instructions in the form of software.

[0202] Figure 7 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 7 shown, the electronic device 12 may include: a processor 71, a memory 72, and computer program instructions stored in the memory 72 and executable on the processor 71. When the processor 71 executes the computer program instructions, the training method of the classification model and / or the text classification method provided in any of the foregoing embodiments is implemented.

[0203] Optionally, the above-mentioned various components of the electronic device 12 may be connected through a system bus.

[0204] The memory 72 may be a separate storage unit or may be an integrated storage unit in the processor. The number of processors is one or more.

[0205] Optionally, the electronic device 12 may further include an interface for interacting with other devices.

[0206] The transceiver is used to communicate with other computers, and the transceiver constitutes a communication interface.

[0207] It should be understood that the processor 71 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor.

[0208] The system bus can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory.

[0209] All or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a readable memory. When the program is executed, it executes the steps including the above method embodiments; and the foregoing memory (storage medium) includes: Read-Only Memory (ROM), RAM, flash memory, hard disk, solid state drive, magnetic tape, floppy disk, optical disc, and any combination thereof.

[0210] The electronic device provided in the embodiments of the present application can be used to execute the training method of the classification model and the text classification method provided in any of the above method embodiments. The implementation principles and technical effects are similar and will not be elaborated here.

[0211] The embodiments of the present application provide a computer-readable storage medium. Computer-executable instructions are stored in the computer-readable storage medium. When the computer-executable instructions run on a computer, the computer is caused to execute the training method of the classification model and the text classification method described above.

[0212] For the above computer-readable storage medium, the above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory, Electrically Erasable Programmable Read-Only Memory, Erasable Programmable Read-Only Memory, Programmable Read-Only Memory, Read-Only Memory, magnetic memory, flash memory, disk or optical disc. The readable storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0213] Optionally, a readable storage medium is coupled to the processor such that the processor can read information from, and write information to, the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in the device.

[0214] An embodiment of the present application further provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium, and when the at least one processor executes the computer program, the training method and text classification method of the above classification model can be implemented.

[0215] It should be understood that the present application is not limited to the exact structures that have been described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A training method for a classification model, characterized in that, it includes: Construct a sample vocabulary based on multiple sample texts. Each sample text includes multiple sample sub-texts obtained through word segmentation processing and at least one class label. The sample vocabulary includes multiple sample sub-texts and multiple class labels; Generate a word vector matrix according to the sample vocabulary. The word vector matrix includes word vectors corresponding to each sample sub-text in the sample vocabulary and class label vectors corresponding to each class label; Perform at least one round of training on the neural network model according to at least one sample text and the word vector matrix to obtain a classification model; Any round of training process includes: determining multiple target word vectors corresponding to each class label vector according to the neural network model obtained in the previous round of training, the sample text, and the word vector matrix, and determining the text semantic vector of the sample text according to the neural network model obtained in the previous round of training, each class label vector, and the corresponding multiple target word vectors, and generating the neural network model obtained in this round of training based on the text semantic vector; Among them, the determining multiple target word vectors corresponding to each class label vector according to the neural network model obtained in the previous round of training, the sample text, and the word vector matrix includes: Determine the word vectors of the sample text according to the neural network model obtained in the previous round of training, the sample text, and the word vector matrix; For any class label vector, calculate the correlation between each word vector in the sample text and the class label vector respectively; Determine the first preset number of word vectors with the largest correlation as the multiple target word vectors corresponding to the class label vector; Among them, the determining the text semantic vector of the sample text according to the neural network model obtained in the previous round of training, each class label vector, and the corresponding multiple target word vectors includes: For the multiple target word vectors corresponding to any class label vector, sort the multiple target word vectors according to the neural network model obtained in the previous round of training and the positions where the sample sub-texts corresponding to each target word vector appear in the sample text to determine the target word sequence vector; Determine the sample sub-text semantic vector corresponding to the class label vector according to the target word sequence vector; Average the sample sub-text semantic vectors corresponding to each class label vector to determine the text semantic vector of the sample text.

2. The method according to claim 1, characterized in that, the sorting the multiple target word vectors corresponding to any class label vector according to the neural network model obtained in the previous round of training and the positions where the sample sub-texts corresponding to each target word vector appear in the sample text to determine the target word sequence vector includes: For the multiple target word vectors corresponding to any class label vector, sort the multiple target word vectors according to the neural network model obtained in the previous round of training and the positions where the sample sub-texts corresponding to each target word vector appear in the sample text to determine the initial word sequence vector; Construct a window according to the distance between the sample sub-text corresponding to each target word vector and other sample sub-texts in the sample text and a second preset quantity, where the window includes the target word vector and the second preset quantity of word vectors that are the closest to the sample sub-text corresponding to the target word vector; Calculate the average vector of each vector in each window, and construct a sample sub-text sequence corresponding to the target word vector according to the average vector of each window.

3. The method according to claim 2, wherein, the determining the sample sub-text semantic vector corresponding to the class label vector according to the target word sequence vector includes: extracting the text features of the target word sequence vector; determining the semantic vectors of the target word vectors in the target word sequence vector according to the self-attention mechanism and the text features; determining the sample sub-text semantic vector corresponding to the class label vector according to the semantic vectors of the target word vectors and the contribution degrees of the target word vectors to the class label vector.

4. The method according to claim 1, wherein, the sample word table further includes the serial numbers of each sample sub-text and the serial number of the class label; the word vector matrix further includes the serial numbers of the sample sub-texts corresponding to each word vector and the serial numbers of the class labels corresponding to each class label vector, and the training the neural network model for at least one round according to at least one sample text and the word vector matrix to obtain a classification model includes: training the neural network model for at least one round according to the serial numbers of multiple sample sub-texts of at least one sample text and the serial number of at least one class label to obtain the classification model.

5. A text classification method, wherein, it includes: constructing a word table according to multiple texts to be classified, where the texts to be classified include multiple sub-texts obtained by word segmentation processing, and the word table includes multiple sub-texts and preset category information; inputting at least one text to be classified and the word table into a classification model, and obtaining a classification result of the text to be classified, where the classification model is trained by using the method according to any one of claims 1-4.

6. A training device for a classification model, wherein, it includes: a construction module, configured to construct a sample word table according to multiple sample texts, each sample text includes multiple sample sub-texts obtained by word segmentation processing and at least one class label, and the sample word table includes multiple sample sub-texts and multiple class labels; a generation module, configured to generate a word vector matrix according to the sample word table, where the word vector matrix includes word vectors corresponding to each sample sub-text in the sample word table and class label vectors corresponding to each class label; a training module, configured to train a neural network model for at least one round according to at least one sample text and the word vector matrix to obtain a classification model; Any round of training process includes: determining multiple target word vectors corresponding to various label vectors according to the neural network model obtained from the previous round of training, the sample text, and the word vector matrix, and determining the text semantic vector of the sample text according to the neural network model obtained from the previous round of training, various label vectors, and the corresponding multiple target word vectors, and generating the neural network model obtained from this round of training based on the text semantic vector; Among them, the training module is specifically used for: determining each word vector of the sample text according to the neural network model obtained from the previous round of training, the sample text, and the word vector matrix; for any label vector, calculating the relevance between each word vector in the sample text and the label vector respectively; determining the first preset number of word vectors with the maximum relevance as the multiple target word vectors corresponding to the label vector; Among them, the training module is specifically used for: for the multiple target word vectors corresponding to any label vector, sorting the multiple target word vectors according to the neural network model obtained from the previous round of training and the positions where the sample sub-texts corresponding to each target word vector appear in the sample text, and determining the target word sequence vector; determining the sample sub-text semantic vector corresponding to the label vector according to the target word sequence vector; averaging the sample sub-text semantic vectors corresponding to various label vectors to determine the text semantic vector of the sample text.

7. A text classification device characterized in that it includes: a construction module, configured to construct a word list according to multiple texts to be classified, the texts to be classified include multiple sub-texts obtained through word segmentation processing, and the word list includes multiple sub-texts and pre-set category information; an input module, configured to input at least one text file to be classified and the word list into a classification model, and obtain the classification result of the text file to be classified, and the classification model is trained by using the method described in any one of claims 1-4.

8. An electronic device including: a processor, a memory, and computer program instructions stored on the memory and executable on the processor, characterized in that when the processor executes the computer program instructions, it is used to implement the method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Small sample text classification method and model

    CN114117039A

  • Text classification method and system based on neural network, and computer device

    WO2020224106A1