A method and system for establishing a multi-level classification model for short texts
By establishing a short text multi-level classification model, using the hierarchical annotation data set and the pre-trained model Bert base for model training and parameter migration, the problems of high-level cross-category errors and data sparseness in the existing short text classification model are solved, and higher prediction accuracy and human acceptance are achieved.
Patent Information
- Application Number
- CN202111636972.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-29
AI Technical Summary
The existing short text classification model only learns the lowest-level category classification tasks, resulting in high-level cross-category errors, and the data sparseness caused by the large number of categories, reducing the network prediction accuracy.
By establishing a short text multi-level classification model, the short text is annotated using a hierarchical annotation data set, and the public pre-trained model Bert base is followed by a fully connected layer for model training to generate the optimal multi-level classification model. Migration of some training parameters between adjacent levels is carried out so that information can be fully interacted and utilized.
It effectively solves the problems of high-level cross-category errors and data sparseness, improves the prediction accuracy of short text classification, and limits the prediction error to a single level, which is in line with human acceptance.
Smart Images

Figure CN114579737B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of text classification, and in particular to a method and system for establishing a multi-level classification model for short texts. Background Art
[0002] In recent years, with the explosive growth of online social network applications, short text classification technology has been widely studied, such as microblogs, chat messages, news topics, opinion comments, question texts, mobile phone messages, literature abstracts, etc. Compared with long texts, short texts lack topicality. One way to solve this problem is to expand short text information by extracting knowledge from text libraries, professional dictionaries, and thesauruses. However, due to the domain independence of professional dictionaries and thesauruses, the data distribution of external knowledge is quite different from the test data distribution collected in some special fields, thus affecting the overall performance of classification. With the development of deep learning technology, some deep network models have achieved good results in short text classification applications, such as TextCNN, LSTM, etc. However, the current mainstream network models do not consider the hierarchy of text categories. For example, "cats" and "dogs" belong to animals, and "orchids" and "chrysanthemums" belong to plants. If the high-level categories (animals, plants) are simply ignored and the network only learns the classification tasks of the bottom-level categories, there will be high-level cross-category errors such as predicting "animals" as "plants", and it will also face the problem of data sparsity caused by a large number of categories, thereby reducing the network prediction accuracy. Summary of the Invention
[0003] In order to solve the problems in the prior art that when classifying short texts, only the classification tasks of the bottom-level categories are learned, high-level cross-category errors will occur, and data sparsity caused by too many categories is faced, etc., embodiments of the present invention provide a method and system for establishing a multi-level classification model for short texts.
[0004] According to one aspect of the embodiments of the present invention, a method for establishing a multi-level classification model for short texts is provided. The method includes:
[0005] Step 101, obtaining a first-level labeled data set, where the first-level labeled data set is a data set generated by labeling each short text in the short text data set according to a pre-set first-level category label;
[0006] Step 102: Input the first-level labeled dataset into the initial first-level classification model for model training to generate the optimal first-level classification model. Among them, the initial first-level classification model is the publicly available pre-trained model Bert base followed by the initial first-level fully connected layer. The optimal first-level classification model is the optimal first-level pre-trained model Bert base followed by the optimal first-level fully connected layer. The optimal first-level pre-trained model Bert base is the pre-trained model Bert base obtained by fine-tuning the publicly available pre-trained model Bert base. The optimal first-level fully connected layer is the fully connected layer obtained by adjusting the parameters of the initial first-level fully connected layer;
[0007] Step 103: Obtain the second-level labeled dataset. Among them, the second-level labeled dataset is the dataset generated by labeling each short text in the short text dataset according to the pre-set second-level category labels;
[0008] Step 104: Input the second-level labeled dataset into the initial second-level classification model for model training to generate the optimal second-level classification model. Among them, the initial second-level classification model is the initial second-level pre-trained model Bert base followed by the initial second-level fully connected layer. The initial second-level pre-trained model Bert base is the pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first-level pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base. The optimal second-level classification model is the optimal second pre-trained model Bert base followed by the optimal second fully connected layer. The optimal second-level pre-trained model Bert base is the pre-trained model Bert base obtained by fine-tuning the initial second-level pre-trained model Bert base. The optimal second-level fully connected layer is the fully connected layer obtained by adjusting the parameters of the initial second-level fully connected layer, where N is a natural number;
[0009] Step 105: Obtain the i-level labeled dataset. Among them, the i-level labeled dataset is the dataset generated by labeling each short text in the short text dataset according to the pre-set i-level category labels, where 3 ≤ i ≤ I, the initial value of i is 3, and I is a natural number;
[0010] Step 106: Input the labeled dataset of the i-th level into the initial classification model of the i-th level for model training to generate the optimal classification model of the i-th level. Among them, the initial classification model of the i-th level is the initial pre-training model Bert base of the i-th level followed by the initial fully connected layer of the i-th level. The initial pre-training model Bert base of the i-th level is a pre-training model Bert base obtained by migrating the training parameters of the first N layers of the optimal pre-training model Bert base of the i-1-th level to the first N layers of the initial pre-training model Bert base of the i-th level. The optimal classification model of the i-th level is the optimal pre-training model Bert base of the i-th level followed by the optimal fully connected layer of the i-th level. The optimal pre-training model Bert base of the i-th level is a pre-training model Bert base obtained by fine-tuning the initial pre-training model Bert base of the i-th level. The optimal fully connected layer of the i-th level is a fully connected layer obtained by adjusting the parameters of the initial fully connected layer of the i-th level;
[0011] Step 107: Let i = i + 1. When i ≤ I, return to Step 105. When i > I, go to Step 108;
[0012] Step 108: Use the model generated by combining the optimal first-level classification model to the optimal I-level classification model in the order from the first level to the I-th level as the short text multi-level classification model.
[0013] Optionally, in each of the above method embodiments of the present invention, before obtaining the labeled dataset of the first level, it further includes:
[0014] Set the short text category labels of J levels, and generate the first-level category labels to the J-level category labels respectively. Among them, the classification level of the j-level category label is higher than that of the j + 1-level category label, 1 ≤ j ≤ J, and J is equal to I;
[0015] Collect multiple short texts to generate a short text dataset;
[0016] According to the set first-level category labels to the J-level category labels, label each short text in the short text dataset respectively, and correspondingly generate the first-level labeled dataset to the J-level labeled dataset.
[0017] Optionally, in each of the above method embodiments of the present invention, the publicly available pre-training model Bert base used in the method has a network layer number L = 12, a hidden layer node number H = 768, and a self-attention head number A = 12.
[0018] Optionally, in each of the above method embodiments of the present invention, the initial second-level pre-trained model Bert base is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first-level pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base, where the value of N is 6.
[0019] According to another aspect of the embodiments of the present invention, a system for establishing a short text multi-level classification model is provided. The system includes:
[0020] A first data module, configured to obtain a first-level labeled data set, where the first-level labeled data set is a data set generated by labeling each short text in the short text data set according to a pre-set first-level category label;
[0021] A first model module, configured to input the first-level labeled data set into an initial first-level classification model for model training to generate an optimal first-level classification model, where the initial first-level classification model is a publicly available pre-trained model Bert base followed by an initial first-level fully connected layer, the optimal first-level classification model is an optimal first-level pre-trained model Bert base followed by an optimal first-level fully connected layer, the optimal first-level pre-trained model Bert base is a pre-trained model Bert base obtained by fine-tuning the publicly available pre-trained model Bert base, and the optimal first-level fully connected layer is a fully connected layer obtained by adjusting the parameters of the initial first-level fully connected layer;
[0022] A second data module, configured to obtain a second-level labeled data set, where the second-level labeled data set is a data set generated by labeling each short text in the short text data set according to a pre-set second-level category label;
[0023] The second model module is used to input the second-level labeled dataset into the initial second-level classification model for model training to generate an optimal second-level classification model. Among them, the initial second-level classification model is the initial second-level pre-trained model Bert base followed by the initial second-level fully connected layer. The initial second-level pre-trained model Bert base is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first-level pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base. The optimal second-level classification model is the optimal second pre-trained model Bert base followed by the optimal second fully connected layer. The optimal second-level pre-trained model Bert base is a pre-trained model Bert base obtained by fine-tuning the initial second-level pre-trained model Bert base. The optimal second-level fully connected layer is a fully connected layer obtained by adjusting the parameters of the initial second-level fully connected layer, and N is a natural number;
[0024] The third data module is used to obtain the i-level labeled dataset. Among them, the i-level labeled dataset is a dataset generated by labeling each short text in the short text dataset according to the preset i-level category label, where 3 ≤ i ≤ I, the initial value of i is 3, and I is a natural number;
[0025] The third model module is used to input the i-level labeled dataset into the initial i-level classification model for model training to generate an optimal i-level classification model. Among them, the initial i-level classification model is the initial i-level pre-trained model Bert base followed by the initial i-level fully connected layer. The initial i-level pre-trained model Bert base is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal i-1 level pre-trained model Bert base to the first N layers of the initial i-level pre-trained model Bert base. The optimal i-level classification model is the optimal i-level pre-trained model Bert base followed by the optimal i-level fully connected layer. The optimal i-level pre-trained model Bert base is a pre-trained model Bert base obtained by fine-tuning the initial i-level pre-trained model Bert base. The optimal i-level fully connected layer is a fully connected layer obtained by adjusting the parameters of the initial i-level fully connected layer;
[0026] The iterative calculation module is used to set i = i + 1. When i ≤ I, it returns to the third data module. When i > I, it goes to the model generation module;
[0027] A model generation module, which uses the model generated by combining the optimal first-level classification model to the optimal I-level classification model in the order from the first level to the I level as the short text multi-level classification model.
[0028] Optionally, in each of the above device embodiments of the present invention, the system further includes a data annotation module, which is used to annotate short texts to generate an annotated data set, where:
[0029] A category label unit, which is used to set short text category labels for J levels, and respectively generate first-level category labels to J-level category labels. Among them, the classification level of the j-level category label is higher than that of the j+1-level category label, 1≤j≤J, and J is equal to I;
[0030] A text collection unit, which is used to collect multiple short texts to generate a short text data set;
[0031] A text annotation unit, which is used to annotate each short text in the short text data set according to the set first-level category labels to J-level category labels, and respectively generate a first-level annotated data set to a J-level annotated data set.
[0032] Optionally, in each of the above device embodiments of the present invention, the publicly available pre-trained model Bert base adopted by the first model module has a network layer number L = 12, a hidden layer node number H = 768, and a self-attention head number A = 12.
[0033] Optionally, in each of the above device embodiments of the present invention, the initial second-level pre-trained model Bert base in the second model module is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first-level pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base, where the value of N is 6.
[0034] Based on the method and system for establishing a short text multi-level classification model provided in the above embodiments of the present invention, the method includes: for the same short text dataset, generating different-level labeled datasets after labeling according to the set short text category labels at different levels, and using them as inputs to hierarchically train a classification model established by connecting a fully connected layer after the publicly available pre-trained model Bert base, generating different-level classification models, and when training the next-level classification model, transferring some training parameters of the pre-trained model Bert base fine-tuned at the previous level to the corresponding parts of the initial pre-trained model Bert base at the next level, and finally combining the generated multi-level classification models to generate the final classification model. The beneficial effects of the method and system include: through the multi-level classification model, aiming at the multi-level characteristics of short text classification, designing a multi-level classification model, performing fine-tuning learning for each level, and directly transferring some training parameters between adjacent levels, which can enable the full interaction and effective utilization of information between adjacent levels. On the one hand, for the high level, it can effectively increase the total amount of data under each category, solving the problem of data sparsity in the training of the low-level model; on the other hand, for the classification learning of the low-level classification model, it can guide the learning by transferring the general parameters of the high-level classification model, thereby improving the training effect. Further, compared with the existing short text classification models, the multi-level classification model proposed in this patent can limit the prediction error within a single level, avoiding the problem that the existing classification models simply flatten the multi-level categories into a single level for classification learning, which is prone to prediction errors between different-level categories, making it more acceptable to people.
[0035] The following will further describe the technical solutions of the present invention in detail through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] By describing the embodiments of the present invention in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present invention will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0037] Figure 1 is a schematic flowchart of a method for establishing a short text multi-level classification model provided by an exemplary embodiment of the present invention;
[0038] Figure 2 is a schematic diagram of short text classification using a short text multi-level classification model provided by an exemplary embodiment of the present invention;
[0039] Figure 3It is a schematic structural diagram of a system for establishing a short text multi - level classification model provided by an exemplary embodiment of the present invention. Detailed implementation manners
[0040] Next, exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited by the exemplary embodiments described herein.
[0041] It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present invention.
[0042] Those skilled in the art can understand that terms such as "first", "second", etc. in the embodiments of the present invention are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.
[0043] It should also be understood that in the embodiments of the present invention, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.
[0044] It should also be understood that for any component, data or structure mentioned in the embodiments of the present invention, without clear limitation or contrary indication in the context, it is generally understood as one or more.
[0045] In addition, the term "and / or" in the present invention is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the associated objects before and after.
[0046] It should also be understood that the present invention emphasizes the differences between various embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be described one by one.
[0047] At the same time, it should be understood that for the convenience of description, the dimensions of each part shown in the drawings are not drawn according to the actual proportional relationship.
[0048] The following description of at least one exemplary embodiment is actually only illustrative and in no way restricts the present invention and its application or use.
[0049] Techniques, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods and devices should be regarded as part of the specification.
[0050] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, further discussion thereof is not required in subsequent figures.
[0051] Embodiments of the present invention can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate together with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, small computer systems, large computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.
[0052] Terminal devices, computer systems, servers and other electronic devices can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules may include routines, programs, target programs, components, logics, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment, where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.
[0053] Exemplary method
[0054] Figure 1 is a schematic flowchart of a method for establishing a short text multi-level classification model provided by an exemplary embodiment of the present invention. This embodiment can be applied to an electronic device, such as Figure 1 As shown, the method for establishing a short text multi-level classification model in this embodiment includes the following steps:
[0055] Step 101, obtain a first-level labeled data set, where the first-level labeled data set is a data set generated by labeling each short text in the short text data set according to a pre-set first-level category label.
[0056] Optionally, before obtaining the first-level labeled data set, it further includes:
[0057] Set short text category labels for J levels, and generate the first-level category label to the Jth-level category label respectively, where the classification level of the jth-level category label is higher than that of the j + 1th-level category label, 1 ≤ j ≤ J, and J is equal to I;
[0058] Collect multiple short texts to generate a short text dataset;
[0059] According to the set first-level category labels to the J-level category labels, annotate each short text in the short text dataset respectively, and correspondingly generate the first-level annotation dataset to the J-level annotation dataset.
[0060] In one embodiment, fully considering the hierarchy of the categories to which the short texts belong, different levels of category labels are set for the short texts. For example, for cats and dogs, the category labels of two levels, namely animals, feline and animals, canine, are set respectively. For orchids and chrysanthemums, the category labels of two levels, namely plants, orchidaceae and plants, asteraceae, are set respectively. By establishing two levels of category labels, it is avoided that the classification model only learns the bottom-level classification tasks, thus resulting in cross-category errors such as predicting animals as plants. Moreover, when the classification model only learns the bottom-level classification tasks, it is also prone to the problem of data sparsity caused by a large number of categories, thereby reducing the network prediction accuracy.
[0061] Step 102: Input the first-level annotation dataset into the initial first-level classification model for model training to generate the optimal first-level classification model. Among them, the initial first-level classification model is the publicly available pre-trained model Bert base followed by the initial first-level fully connected layer. The optimal first-level classification model is the optimal first-level pre-trained model Bert base followed by the optimal first-level fully connected layer. The optimal first-level pre-trained model Bert base is the pre-trained model Bert base obtained by fine-tuning the publicly available pre-trained model Bert base. The optimal first-level fully connected layer is the fully connected layer obtained by adjusting the parameters of the initial first-level fully connected layer.
[0062] Optionally, the publicly available pre-trained model Bert base adopted by the method has a network layer number L = 12, a hidden layer node number H = 768, and a self-attention head number A = 12.
[0063] In one embodiment, for multiple publicly available pre-trained models Bert base, in this embodiment, the publicly available pre-trained model Bert base with a network layer number L = 12, a hidden layer node number H = 768, and a self-attention head number A = 12 is selected.
[0064] Since, on the premise of having an annotated dataset, training the initial first-level classification model with the publicly available pre-trained model Bert base followed by a fully connected layer, and by analyzing the output results, thus fine-tuning the training parameters of the publicly available pre-trained model Bert base and the fully connected layer to obtain the optimal first-level classification model is known to those skilled in the art and will not be elaborated here.
[0065] Step 103: Obtain the second-level labeled dataset, where the second-level labeled dataset is a dataset generated by labeling each short text in the short text dataset according to the preset second-level category labels.
[0066] Step 104: Input the second-level labeled dataset into the initial second-level classification model for model training to generate the optimal second-level classification model. The initial second-level classification model is the initial second-level pre-trained model Bert base followed by the initial second-level fully connected layer. The initial second-level pre-trained model Bert base is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first-level pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base. The optimal second-level classification model is the optimal second pre-trained model Bert base followed by the optimal second fully connected layer. The optimal second-level pre-trained model Bert base is a pre-trained model Bert base obtained by fine-tuning the initial second-level pre-trained model Bert base. The optimal second-level fully connected layer is a fully connected layer obtained by adjusting the parameters of the initial second-level fully connected layer. N is a natural number.
[0067] Optionally, the initial second-level pre-trained model Bert base is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first-level pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base, where the value of N is 6.
[0068] In one embodiment, the value of N is selected as 6. The initial first-level classification model is trained based on the first labeled dataset to obtain the optimal first-level classification model, and the general features learned by the first 6 layers of the pre-trained model Bert base in the optimal first-level classification model are migrated to the first 6 layers of the pre-trained model Bert base in the second-level classification model for the second-level classification model to learn more categories involved in the second labeled dataset. By means of migrating network parameters, the problem of data sparsity during the training of the second-level classification model is alleviated.
[0069] Step 105: Obtain the i-level labeled dataset, where the i-level labeled dataset is a dataset generated by labeling each short text in the short text dataset according to the preset i-level category labels, where 3 ≤ i ≤ I, the initial value of i is 3, and I is a natural number.
[0070] Step 106: Input the labeled dataset of the i-th level into the initial i-th level classification model for model training to generate the optimal i-th level classification model. Among them, the initial i-th level classification model is the initial i-th level pre-trained model Bertbase followed by the initial i-th level fully connected layer. The initial i-th level pre-trained model Bert base is the pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal (i - 1)-th level pre-trained model Bert base to the first N layers of the initial i-th level pre-trained model Bert base. The optimal i-th level classification model is the optimal i-th level pre-trained model Bert base followed by the optimal i-th level fully connected layer. The optimal i-th level pre-trained model Bert base is the pre-trained model Bert base obtained by fine-tuning the initial i-th level pre-trained model Bert base. The optimal i-th level fully connected layer is the fully connected layer obtained by adjusting the parameters of the initial i-th level fully connected layer.
[0071] Step 107: Let i = i + 1. When i ≤ I, return to Step 105. When i > I, go to Step 108.
[0072] Step 108: Use the model generated by combining the optimal first level classification model to the optimal I-th level classification model in the order from the first level to the I-th level as the short text multi-level classification model.
[0073] Figure 2 It is a schematic diagram for short text classification of the short text multi-level classification model provided by an exemplary embodiment of the present invention. As Figure 2 shown, in one embodiment, through the generated two-level classification model, for the same text, by respectively inputting the optimal first level classification model and the optimal second level classification model in the short text multi-level classification model, then for the same short text, two levels of classification categories are obtained. Among them, in the first level, there are two categories A and B, and in the second level, there are two sub-categories A1 and A2 in category A, and two categories B1 and B2 in category B.
[0074] As can be seen from the above embodiments, if only one level is used to extract information, since the number of custom entity types to be extracted is large, such as hundreds of types, it will inevitably result in insufficient data required for model training. When using a two-level model to classify short texts, since the number of categories set in the first level is relatively small, the data required for model training is also relatively small. However, since the general features learned from the first 6 layers of Bert base in the first-level classification model are transferred to the first 6 layers of Bert base in the second-level classification in the form of training parameters and used for learning hundreds of sub-categories involved in the second level, it can better alleviate the data sparsity problem during the model training of the second-level classification model and also avoid cross-category errors when classifying short texts.
[0075] Exemplary system
[0076] Figure 3 FIG. is a schematic structural diagram of a system for establishing a multi-level classification model for short texts provided by an exemplary embodiment of the present invention. As Figure 3 shown, the system for establishing a multi-level classification model for short texts described in this embodiment includes:
[0077] A first data module 301, configured to obtain a first-level labeled data set, where the first-level labeled data set is a data set generated by labeling each short text in the short text data set according to a pre-set first-level category label;
[0078] A first model module 302, configured to input the first-level labeled data set into an initial first-level classification model for model training to generate an optimal first-level classification model, where the initial first-level classification model is a publicly available pre-trained model Bert base followed by an initial first-level fully connected layer, the optimal first-level classification model is an optimal first-level pre-trained model Bert base followed by an optimal first-level fully connected layer, the optimal first-level pre-trained model Bert base is a pre-trained model Bert base obtained by fine-tuning the publicly available pre-trained model Bert base, and the optimal first-level fully connected layer is a fully connected layer obtained by adjusting the parameters of the initial first-level fully connected layer;
[0079] A second data module 303, configured to obtain a second-level labeled data set, where the second-level labeled data set is a data set generated by labeling each short text in the short text data set according to a pre-set second-level category label;
[0080] The second model module 304 is used to input the second-level labeled dataset into the initial second-level classification model for model training to generate an optimal second-level classification model. Among them, the initial second-level classification model is an initial second-level pre-trained model Bert base followed by an initial second-level fully connected layer. The initial second-level pre-trained model Bert base is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first-level pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base. The optimal second-level classification model is an optimal second pre-trained model Bert base followed by an optimal second fully connected layer. The optimal second-level pre-trained model Bert base is a pre-trained model Bert base obtained by fine-tuning the initial second-level pre-trained model Bert base. The optimal second-level fully connected layer is a fully connected layer obtained by adjusting the parameters of the initial second-level fully connected layer. N is a natural number;
[0081] The third data module 305 is used to obtain the i-level labeled dataset. Among them, the i-level labeled dataset is a dataset generated by labeling each short text in the short text dataset according to the pre-set i-level category label, where 3 ≤ i ≤ I, the initial value of i is 3, and I is a natural number;
[0082] The third model module 306 is used to input the i-level labeled dataset into the initial i-level classification model for model training to generate an optimal i-level classification model. Among them, the initial i-level classification model is an initial i-level pre-trained model Bert base followed by an initial i-level fully connected layer. The initial i-level pre-trained model Bert base is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal i-1 level pre-trained model Bert base to the first N layers of the initial i-level pre-trained model Bert base. The optimal i-level classification model is an optimal i-level pre-trained model Bert base followed by an optimal i-level fully connected layer. The optimal i-level pre-trained model Bert base is a pre-trained model Bert base obtained by fine-tuning the initial i-level pre-trained model Bert base. The optimal i-level fully connected layer is a fully connected layer obtained by adjusting the parameters of the initial i-level fully connected layer;
[0083] The iterative calculation module 307 is used to set i = i + 1. When i ≤ I, it returns to the third data module. When i > I, it goes to the model generation module;
[0084] The model generation module 308 is configured to use the model generated by combining the optimal first-level classification model to the optimal I-level classification model in the order from the first level to the I level as the short text multi-level classification model.
[0085] Optionally, the system further includes a data annotation module 309 for annotating short texts to generate an annotated data set, where:
[0086] The category label unit 391 is configured to set short text category labels for J levels, and generate the first-level category label to the J-level category label respectively, where the classification level of the j-level category label is higher than that of the j+1-level category label, 1≤j≤J, and J is equal to I;
[0087] The text collection unit 392 is configured to collect multiple short texts to generate a short text data set;
[0088] The text annotation unit 393 is configured to annotate each short text in the short text data set according to the set first-level category label to the J-level category label, and correspondingly generate the first-level annotated data set to the J-level annotated data set.
[0089] Optionally, the publicly available pre-trained model Bert base adopted by the first model module 302 has a network layer number L = 12, a hidden layer node number H = 768, and a self-attention head number A = 12.
[0090] Optionally, the initial second-level pre-trained model Bert base in the second model module 304 is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first-level pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base, where the value of N is 6.
[0091] Exemplary computer program product and computer-readable storage medium
[0092] In addition to the above methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the method for establishing a short text multi-level classification model described in the above "Exemplary Method" section of this specification.
[0093] The computer program product can be written in any combination of one or more programming languages for executing the program code of the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0094] In addition, an embodiment of the present disclosure can also be a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the method for establishing a short text multi-level classification model described in the above "Exemplary Method" section of this specification.
[0095] The computer-readable storage medium can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0096] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for the purposes of illustration and facilitating understanding, and are not limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.
[0097] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment.
[0098] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms that mean "including but not limited to" and can be used interchangeably with each other. The word "or" and "and" used herein refer to the phrase "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.
[0099] The methods and apparatuses of this disclosure can be implemented in many ways. For example, the methods and apparatuses of this disclosure can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, this disclosure can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the methods according to this disclosure. Therefore, this disclosure also covers the recording medium storing the programs for executing the methods according to this disclosure.
[0100] It should also be noted that in the apparatuses, equipment, and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure. The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
[0101] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.
Claims
1. A method for establishing a multi - level classification model for short texts, characterized in that, the method comprises: Step 101, obtain a first - level labeled data set, wherein the first - level labeled data set is a data set generated by labeling each short text in the short - text data set according to a pre - set first - level category label; Step 102, input the first - level labeled data set into an initial first - level classification model for model training to generate an optimal first - level classification model, wherein the initial first - level classification model is a publicly available pre - trained model Bert base followed by an initial first - level fully - connected layer, the optimal first - level classification model is an optimal first - level pre - trained model Bert base followed by an optimal first - level fully - connected layer, the optimal first - level pre - trained model Bert base is a pre - trained model Bert base obtained by fine - tuning the publicly available pre - trained model Bert base, and the optimal first - level fully - connected layer is a fully - connected layer obtained by adjusting the parameters of the initial first - level fully - connected layer; Step 103, obtain a second - level labeled data set, wherein the second - level labeled data set is a data set generated by labeling each short text in the short - text data set according to a pre - set second - level category label; Step 104, input the second - level labeled data set into an initial second - level classification model for model training to generate an optimal second - level classification model, wherein the initial second - level classification model is an initial second - level pre - trained model Bert base followed by an initial second - level fully - connected layer, the initial second - level pre - trained model Bert base is a pre - trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first - level pre - trained model Bert base to the first N layers of the publicly available pre - trained model Bert base, the optimal second - level classification model is an optimal second - level pre - trained model Bert base followed by an optimal second - level fully - connected layer, the optimal second - level pre - trained model Bert base is a pre - trained model Bert base obtained by fine - tuning the initial second - level pre - trained model Bert base, and the optimal second - level fully - connected layer is a fully - connected layer obtained by adjusting the parameters of the initial second - level fully - connected layer, and N is a natural number; Step 105, obtain an i - th - level labeled data set, wherein the i - th - level labeled data set is a data set generated by labeling each short text in the short - text data set according to a pre - set i - th - level category label, wherein 3 ≤ i ≤ I, the initial value of i is 3, and I is a natural number; Step 106: Input the labeled dataset of the i-th level into the initial classification model of the i-th level for model training to generate the optimal classification model of the i-th level. Among them, the initial classification model of the i-th level is the initial pre-trained model Bert base of the i-th level followed by the initial fully connected layer of the i-th level. The initial pre-trained model Bert base of the i-th level is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal pre-trained model Bert base of the (i - 1)-th level to the first N layers of the initial pre-trained model Bert base of the i-th level. The optimal classification model of the i-th level is the optimal pre-trained model Bert base of the i-th level followed by the optimal fully connected layer of the i-th level. The optimal pre-trained model Bert base of the i-th level is a pre-trained model Bert base obtained by fine-tuning the initial pre-trained model Bert base of the i-th level. The optimal fully connected layer of the i-th level is a fully connected layer obtained by adjusting the parameters of the initial fully connected layer of the i-th level; Step 107: Let i = i + 1. When i ≤ I, return to Step 105. When i > I, go to Step 108; Step 108: Use the model generated by combining the optimal classification model of the first level to the optimal classification model of the I-th level in the order from the first level to the I-th level as the short text multi-level classification model.
2. The method according to claim 1, wherein, before obtaining the labeled dataset of the first level, it further includes: Setting J levels of short text category labels, and respectively generating the first level category label to the J-th level category label. Among them, the classification level of the j-th level category label is higher than that of the (j + 1)-th level category label, 1 ≤ j ≤ J, and J is equal to I; Collecting multiple short texts to generate a short text dataset; According to the set first level category label to the J-th level category label, respectively annotating each short text in the short text dataset, and correspondingly generating the first level labeled dataset to the J-th level labeled dataset.
3. The method according to claim 1, wherein, The publicly available pre-trained model Bert base used in the method has a network layer number L = 12, a hidden layer node number H = 768, and a self-attention head number A = 12.
4. The method according to claim 3, wherein, The initial pre-trained model Bert base of the second level is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal pre-trained model Bert base of the first level to the first N layers of the publicly available pre-trained model Bert base, where the value of N is 6.
5. A system for establishing a short text multi-level classification model, wherein, The system includes: A first data module for obtaining the labeled dataset of the first level, where the labeled dataset of the first level is a dataset generated by annotating each short text in the short text dataset according to the pre-set first level category label; The first model module is used to input the first-level labeled dataset into the initial first-level classification model for model training to generate the optimal first-level classification model. Among them, the initial first-level classification model is the publicly available pre-trained model Bert base followed by the initial first-level fully connected layer. The optimal first-level classification model is the optimal first-level pre-trained model Bert base followed by the optimal first-level fully connected layer. The optimal first-level pre-trained model Bert base is the pre-trained model Bert base obtained by fine-tuning the publicly available pre-trained model Bert base. The optimal first-level fully connected layer is the fully connected layer obtained by adjusting the parameters of the initial first-level fully connected layer; The second data module is used to obtain the second-level labeled dataset. Among them, the second-level labeled dataset is the dataset generated by labeling each short text in the short text dataset according to the pre-set second-level category labels; The second model module is used to input the second-level labeled dataset into the initial second-level classification model for model training to generate the optimal second-level classification model. Among them, the initial second-level classification model is the initial second-level pre-trained model Bert base followed by the initial second-level fully connected layer. The initial second-level pre-trained model Bert base is the pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first-level pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base. The optimal second-level classification model is the optimal second-level pre-trained model Bert base followed by the optimal second-level fully connected layer. The optimal second-level pre-trained model Bert base is the pre-trained model Bert base obtained by fine-tuning the initial second-level pre-trained model Bert base. The optimal second-level fully connected layer is the fully connected layer obtained by adjusting the parameters of the initial second-level fully connected layer, where N is a natural number; The third data module is used to obtain the i-level labeled dataset. Among them, the i-level labeled dataset is the dataset generated by labeling each short text in the short text dataset according to the pre-set i-level category labels, where 3 ≤ i ≤ I, the initial value of i is 3, and I is a natural number; The third model module is used to input the labeled dataset at the i-th level into the initial i-th level classification model for model training to generate the optimal i-th level classification model. Among them, the initial i-th level classification model is the initial i-th level pre-trained model Bert base followed by the initial i-th level fully connected layer. The initial i-th level pre-trained model Bert base is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal i-1-th level pre-trained model Bert base to the first N layers of the initial i-th level pre-trained model Bert base. The optimal i-th level classification model is the optimal i-th level pre-trained model Bert base followed by the optimal i-th level fully connected layer. The optimal i-th level pre-trained model Bert base is a pre-trained model Bert base obtained by fine-tuning the initial i-th level pre-trained model Bert base. The optimal i-th level fully connected layer is a fully connected layer obtained by adjusting the parameters of the initial i-th level fully connected layer; The iterative calculation module is used to set i = i + 1. When i ≤ I, it returns to the third data module. When i > I, it goes to the model generation module; The model generation module is used to use the model generated by combining the optimal first level classification model to the optimal I-th level classification model in the order from the first level to the I-th level as the short text multi-level classification model.
6. The system according to claim 5, wherein, the system further includes a data annotation module for annotating short texts to generate an annotated dataset, where: The category label unit is used to set the short text category labels of J levels, and generate the first level category label to the J-th level category label respectively. Among them, the classification level of the j-th level category label is higher than that of the j + 1-th level category label, 1 ≤ j ≤ J, and J is equal to I; The text collection unit is used to collect multiple short texts to generate a short text dataset; The text annotation unit is used to annotate each short text in the short text dataset according to the set first level category label to the J-th level category label, and correspondingly generate the first level annotated dataset to the J-th level annotated dataset.
7. The system according to claim 5, wherein, the publicly available pre-trained model Bert base used by the first model module has a network layer number L = 12, a hidden layer node number H = 768, and a self-attention head number A = 12.
8. The system according to claim 7, wherein, the initial second level pre-trained model Bert base in the second model module is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first level pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base, where the value of N is 6.
Citation Information
Patent Citations
Text classification data processing method and device, storage medium and program product
CN113722493A
Adversarial training of machine learning models
US20210142181A1