Text classification model training method and apparatus, classification method and apparatus, device, medium, and product

The minimum encirclement ball algorithm and SMOTE technology generate enhanced text training vectors, which solves the problem of lack of data in the Q&A text classification of new energy vehicles, and realizes efficient training and ideal classification effects of text classification models.

WO2025171738A1PCT designated stage Publication Date: 2025-08-21ZHEJIANG ZEEKR INTELLIGENT TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/138285
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-18
Filing Date
2024-12-10
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

In the new energy vehicle Q&A text classification, due to the special domain and high annotation cost, the text sample data is lacking, and the training results of the existing classification model are not ideal.

Method used

The minimum encircling sphere algorithm is used to calculate the hyperplanar sphere of each class of initial text training vector, generate enhanced text training vectors, and the number of text training vectors is increased by SMOTE synthesis of a few classes of oversampling technology, and the text classification model is trained using support sample sets and query sample sets.

Benefits of technology

In the absence of training data, the number of text training vectors is increased, the cost of manual labeling is reduced, and the classification effect and generalization ability of the text classification model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024138285_21082025_PF_FP_ABST
    Figure CN2024138285_21082025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are a text classification model training method and apparatus, a classification method and apparatus, a device, a medium, and a product, capable of being applied in but not limited to the technical field of artificial intelligence. The text classification model training method comprises: acquiring initial text training samples, and encoding the acquired initial text training samples to obtain initial text training vectors that are represented in a vectorized manner; on the basis of each category of initial text training vectors, using a minimum enclosing ball algorithm to calculate a hyperplane sphere corresponding to each category, and on the basis of the hyperplane sphere corresponding to each category and the corresponding initial text training vector for generating the hyperplane sphere, obtaining augmented text training vectors corresponding to each category; and training a preset text classification model by using the initial text training samples and augmented text training samples corresponding to each category.
Need to check novelty before this filing date? Find Prior Art

Description

Text classification model training method and classification method, device, equipment, medium and product

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on February 18, 2024, with application number 2024101795902 and application name “Text classification model training method and classification method, device, equipment, medium and product”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to, but is not limited to, the field of artificial intelligence technology, and in particular to a text classification model training method and classification method, device, equipment, medium, and product. Background Art

[0003] With the rapid adoption of new energy vehicles, consumers face numerous choices and confusion when purchasing a car. To better serve consumers, major auto brands have launched their own shopping apps. These apps not only offer vehicle display and purchasing capabilities, but also provide a platform for consumers to ask questions and engage in conversation. To better serve consumers, accurate categorization of automotive Q&A texts is essential. This categorization allows consumers to more easily find questions they are interested in and quickly understand the features, advantages, and disadvantages of different vehicles.

[0004] Currently, training classification models typically involves manually labeling text samples with their respective categories. In the case of new energy vehicle Q&A text classification, the domain specificity and high labeling costs lead to a lack of classified text sample data. Consequently, training classification models on such scarce text samples results in suboptimal training results. Summary of the Invention

[0005] The purpose of this application is to provide a text classification model training method and classification method, device, equipment, medium and product. By adopting the minimum bounding sphere algorithm and determining the enhanced text training vector based on each type of initial text training vector, the number of text training vectors of each type can be increased, avoiding the manual classification and category labeling of large amounts of text data, and reducing labor costs.

[0006] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0007] In a first aspect, the present application discloses a text classification model training method, comprising:

[0008] Obtaining an initial text training sample, and encoding the initial text training sample to obtain an initial text training vector;

[0009] For each type of initial text training vector, the minimum bounding sphere algorithm is used to calculate the hyperplane sphere corresponding to each type of initial text training vector;

[0010] Determine an enhanced text training vector based on each type of initial text training vector and the corresponding hyperplane sphere;

[0011] The preset text classification model is trained using the initial text training vector and the enhanced text training vector of each category.

[0012] Specifically, initial text training samples can be obtained from a database and encoded to obtain vectorized initial text training vectors. Therefore, a minimum bounding sphere algorithm can be used to calculate the hyperplane sphere corresponding to each class based on the initial text training vectors for each class. Furthermore, based on the hyperplane sphere corresponding to each class and the corresponding initial text training vectors that generated the hyperplane spheres, the enhanced text training vectors corresponding to each class can be derived. The initial text training samples and enhanced text training samples corresponding to each class can then be input into a preset text classification model to train the preset text classification model.

[0013] In one implementation, determining the enhanced text training vector based on each type of initial text training vector and the corresponding hyperplane sphere includes:

[0014] The SMOTE synthetic minority class oversampling technique interpolation algorithm is used to generate enhanced text training vectors based on the initial text training vectors of each class and the corresponding hyperplane sphere.

[0015] In one implementation, the training of a preset text classification model using the initial text training vector and the enhanced text training vector for each category includes:

[0016] Generate a support sample set and a query sample set based on each type of initial text training vector and enhanced text training vector;

[0017] Using the support sample set to train the preset text classification model, and using the query sample set to test the trained text classification model;

[0018] When the test pass condition is met, the text classification model that meets the test pass condition is determined as a text classification model that has been trained to convergence.

[0019] In one implementation, before training the preset text classification model using the initial text training vector and the enhanced text training vector of each category, the method further includes:

[0020] Determine the category reference vector corresponding to the center of each hyperplane sphere;

[0021] The enhanced text training vector of the corresponding category is sampled according to the category reference vector to obtain the sampled enhanced text training vector.

[0022] In one implementation, the sampling process of the enhanced text training vector of the corresponding category according to the category reference vector to obtain the sampled enhanced text training vector includes:

[0023] Calculate the cosine distance between each augmented text training vector and the category benchmark vector of the corresponding category;

[0024] The enhanced text training vector is sampled according to the cosine distance to obtain a sampled enhanced text training vector.

[0025] In one implementation, the sampling process of the enhanced text training vector according to the cosine distance to obtain the sampled enhanced text training vector includes:

[0026] Calculate the corresponding sampling weight according to each cosine distance;

[0027] The enhanced text training vector is sampled based on a sampling weight and a preset weight threshold to obtain a sampled enhanced text training vector.

[0028] In one implementation, the sampling process of the enhanced text training vector based on the sampling weight and the preset weight threshold to obtain the sampled enhanced text training vector includes:

[0029] In response to the sampling weight being greater than or equal to the preset weight threshold, the corresponding enhanced text training vector is retained as the sampled enhanced text training vector.

[0030] In one implementation, the preset text classification model is a preset prototype network.

[0031] In a second aspect, the present application provides a text classification method, comprising:

[0032] Get the target text to be classified;

[0033] Encoding the target text to obtain a target text vector;

[0034] The target text vector is classified using a text classification model that has been trained to convergence to obtain a corresponding category, wherein the text classification model that has been trained to convergence is obtained by training a preset text classification model using any of the methods described in the first aspect. In a third aspect, the present application provides a text classification model training device, the device comprising:

[0035] An acquisition module is used to obtain initial text training samples;

[0036] An encoding module, configured to perform encoding processing on the initial text training sample to obtain an initial text training vector;

[0037] A calculation module is used to calculate the hyperplane sphere corresponding to each type of initial text training vector using a minimum bounding sphere algorithm;

[0038] a determination module, configured to determine an enhanced text training vector based on each type of initial text training vector and the corresponding hyperplane sphere;

[0039] The training module is used to train a preset text classification model using the initial text training vector and the enhanced text training vector for each category.

[0040] In a fourth aspect, the present application provides a text classification device, comprising:

[0041] An acquisition module is used to obtain the target text to be classified;

[0042] An encoding module, configured to encode the target text to obtain a target text vector;

[0043] The classification module is used to classify the target text vector using a text classification model that has been trained to convergence to obtain a corresponding category.

[0044] In a fifth aspect, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;

[0045] The memory stores computer-executable instructions;

[0046] The processor executes the computer-executable instructions stored in the memory to implement the method as described in any one of the first aspect and the second aspect.

[0047] In a sixth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method as described in any one of the first and second aspects.

[0048] In a seventh aspect, the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method as described in any one of the first and second aspects.

[0049] Still other aspects will become apparent upon reading and understanding the accompanying drawings and detailed description.

[0050] The text classification model training method and classification method, device, equipment, medium and product provided in the present application obtain initial text training samples and encode the initial text training samples to obtain initial text training vectors; for each type of initial text training vectors, the minimum bounding sphere algorithm is used to calculate the hyperplane sphere corresponding to each type of initial text training vectors; based on each type of initial text training vectors and the corresponding hyperplane spheres, enhanced text training vectors are determined; and the initial text training vectors and enhanced text training vectors of each type are used to train a preset text classification model. Since the obtained initial text training samples are encoded, the vectorized initial text training vectors are obtained. Therefore, the minimum bounding sphere algorithm can be used to calculate the hyperplane sphere corresponding to each type based on the initial text training vectors of each type, and the enhanced text training vector corresponding to each type can be obtained based on the hyperplane sphere corresponding to each type and the corresponding initial text training vectors that generate the hyperplane spheres, thereby increasing the number of text training vectors for each type, thereby avoiding the manual labeling of classification category labels for a large amount of text data and reducing labor costs. By inputting the initial and enhanced text training samples corresponding to each category into a preset text classification model for training, and by ensuring that the enhanced text training vectors obtained based on the hyperplane sphere are strongly correlated with each category, the text classification model achieves good classification results. This makes it possible to train the text classification model even when training data is scarce, while still ensuring ideal classification results. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0052] FIG1 is an application scenario diagram of a text classification model training method and a classification method provided by one embodiment of the present application;

[0053] FIG2 is a flowchart of a text classification model training method provided in one embodiment of the present application;

[0054] FIG3 is a flowchart of a text classification model training method provided by another embodiment of the present application;

[0055] FIG4 is a flowchart of a text classification method provided by another embodiment of the present application;

[0056] FIG5 is a schematic diagram of the structure of a text classification model training device provided in one embodiment of the present application;

[0057] FIG6 is a schematic diagram of the structure of a text classification device provided in one embodiment of the present application;

[0058] FIG7 is a schematic structural diagram of an electronic device provided in one embodiment of the present application.

[0059] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0060] The exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, identical numerals in different figures represent identical or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Instead, they are merely examples of apparatuses and methods consistent with certain aspects of the present application, as detailed in the appended claims. Currently, when training a classification model, a large number of text samples labeled with class labels are generally used to train the classification model in order to achieve good classification results. The number of text samples corresponding to each class can reach tens of thousands. Specifically, before training the classification model, a large number of text samples must be prepared, and each text must be manually classified and labeled with class labels. After the text samples are labeled with class labels, the labeled text samples are fed into an encoder, which encodes the text samples for each class. The encoded text samples are then fed into the classification model for model training. Labeling a large number of text samples with class labels results in wasted time and increased labor costs.

[0061] In the field of new energy vehicles, due to the large number of text categories and the small amount of data for each type of question and answer text, the classification effect is poor and cannot be guaranteed.

[0062] Therefore, the present application provides a text classification model training method and a classification method. In order to enable the classification model trained in a small sample scenario to accurately classify the target text, the amount of text data is increased through data enhancement technology. Specifically, a training sample of the initial text is obtained, and the training sample of the initial text contains multiple categories of text with marked class labels. The training sample of the initial text is encoded by an encoder to obtain a training vector of the initial text. However, because the training sample of the initial text is a small sample, the amount of data for each type of training text vector must be increased. Therefore, the initial text training vector of each category is calculated by the minimum enclosing sphere algorithm to obtain the hyperplane sphere corresponding to each type of initial text training vector. The enhanced text training vector can be determined by each type of initial text training vector and the corresponding hyperplane sphere, thereby realizing data enhancement of each type of text training vector, and then the training of the preset text classification model can be realized by inputting each type of initial text training vector and the enhanced text training vector into the preset text classification model.

[0063] It should be understood by those skilled in the art that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications should be included in the scope of the claims of the present invention.

[0064] FIG1 is an application scenario diagram of a text classification model training method and a classification method provided by an embodiment of the present application, as shown in FIG1 . The system corresponding to the text classification model training method and the classification method in the embodiment of the present application may include: a text classification model training device 101 and a text classification device 102. When the text classification device 102 performs text classification, it uses the text classification model that has been trained to convergence in the text classification model training device 101. Specifically, the text classification model that has been trained to convergence can be obtained from the text classification model training device. When performing model training, the text classification model training device 101 obtains training samples of the initial text and converts them into initial text training vectors represented by vectorization through encoding. For the initial text training vectors corresponding to each category, the minimum bounding sphere algorithm is used to calculate the hyperplane sphere corresponding to each category of the initial text training vectors, thereby generating the enhanced text training vectors corresponding to each category based on the initial text training vectors and the corresponding hyperplane spheres, and then training the preset text classification model based on the enhanced text training vectors corresponding to each category and the initial text training vectors of each category until the text training model training is completed. After the text training model is trained, the text classification device 102 can use the trained text classification model to classify the received text. Specifically, the text classification device 102 obtains the target text to be classified, converts the target text to be classified into a target text vector through an encoder, inputs the target text vector into the trained text classification model to perform text classification, and outputs the final classification result after the classification is completed.

[0065] The following specific embodiments are used to describe the technical solutions of the present application and the technical solutions of the present application in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0066] Figure 2 is a flowchart of a text classification model training method provided by an embodiment of the present application. As shown in Figure 2, the execution subject of this embodiment is a text classification model training device, which is located in an electronic device. The text classification model training method provided by this embodiment includes the following steps:

[0067] S201: Acquire an initial text training sample and perform encoding processing on the initial text training sample to obtain an initial text training vector.

[0068] The initial text training samples are text samples with class labels to be used for classification model training. The initial text training samples can be stored in a text dataset. The initial text training samples can be declarative sentences or question sentences containing information, which is not limited in this embodiment.

[0069] It is understandable that the initial text training samples can be public text data or text data obtained through a proprietary car company APP, which is not limited in this embodiment.

[0070] Optionally, since this method can be applied to small samples, the number of initial text training samples included in each category can be 50 to 100, and can also be set according to needs. It is not set in this embodiment.

[0071] Specifically, in this embodiment, initial text training samples are obtained from a database, encoded using a fully connected neural network, and converted into initial text training vectors. A fully connected neural network (FCNN) is a deep neural network composed of a series of fully connected layers and is the fundamental architecture of deep learning.

[0072] S202 : For each type of initial text training vector, a minimum bounding sphere algorithm is used to calculate a hyperplane sphere corresponding to each type of initial text training vector.

[0073] The minimum bounding sphere algorithm is an algorithm for calculating the minimum bounding sphere, which can be used to calculate the minimum bounding sphere between a set of data points.

[0074] Specifically, in this embodiment, the initial text training vector for generating a hyperplane sphere is selected from the initial text training vectors of each category according to preset rules, and the initial text training vector selected for each category is input into the minimum bounding sphere algorithm to calculate the hyperplane sphere corresponding to each category according to the minimum bounding sphere algorithm, and the center and radius corresponding to the hyperplane sphere can be obtained through the hyperplane sphere.

[0075] The hyperplane sphere can be represented by a formula, and the center of the sphere can be represented by a vector.

[0076] S203 : Determine an enhanced text training vector based on each type of initial text training vector and the corresponding hyperplane sphere.

[0077] The enhanced text training vector is an added text training vector used for classification model training.

[0078] Specifically, in this embodiment, based on the hyperplane sphere of each category, an enhanced text training vector corresponding to each category is obtained by generating an initial text training vector corresponding to the hyperplane sphere according to each category through a preset data enhancement algorithm.

[0079] S204 , using each type of initial text training vector and enhanced text training vector to train a preset text classification model.

[0080] The preset text classification model is a text classification model with pre-set parameters. The specific type is not limited. For example, it can be a preset prototype network. A prototype network is a small-sample learning method whose basic principle is to create a prototype representation for each category. Then, for a query point to be classified, the distance between the prototype vector of the category and the query point is calculated to determine the classification.

[0081] Specifically, in this embodiment, the obtained initial text training vectors and enhanced text training vectors for each class are input into a preset text classification model for training. During the training process, a determination is made as to whether a preset convergence condition is met. If so, the model that meets the preset convergence condition is determined to be a converged text classification model.

[0082] The text classification model training method provided in this embodiment obtains initial text training samples and encodes the initial text training samples to obtain initial text training vectors; for each type of initial text training vectors, a minimum bounding sphere algorithm is used to calculate the hyperplane sphere corresponding to each type of initial text training vector; an enhanced text training vector is determined based on each type of initial text training vector and the corresponding hyperplane sphere; and a preset text classification model is trained using each type of initial text training vector and the enhanced text training vector. Since the obtained initial text training samples are encoded, a vectorized initial text training vector is obtained. Therefore, the minimum bounding sphere algorithm can be used to calculate the hyperplane sphere corresponding to each type based on the initial text training vector of each type, and the enhanced text training vector corresponding to each type can be obtained based on the hyperplane sphere corresponding to each type and the corresponding initial text training vector that generates the hyperplane sphere, thereby increasing the number of text training vectors for each type, avoiding manual labeling of classification category labels for a large amount of text data, and reducing labor costs. By inputting the initial and enhanced text training samples corresponding to each category into a preset text classification model for training, and by ensuring that the enhanced text training vectors obtained based on the hyperplane sphere are strongly correlated with each category, the text classification model achieves good classification results. This makes it possible to train the text classification model even when training data is scarce, while still ensuring ideal classification results.

[0083] As an optional implementation, based on the above embodiment, the enhanced text training vector is determined based on each type of initial text training vector and the corresponding hyperplane sphere, including:

[0084] The SMOTE synthetic minority class oversampling technique interpolation algorithm is used to generate enhanced text training vectors based on the initial text training vectors of each class and the corresponding hyperplane sphere.

[0085] Among them, the full name of SMOTE is Synthetic Minority Oversampling Technique, which is the synthetic minority oversampling technology. Its main idea is to oversample the minority class samples to reach a number equivalent to that of the majority class samples, so as to better classify them.

[0086] Specifically, in this embodiment, a preset number of initial text training vectors are arbitrarily selected from the initial text training vectors in the hyperplane sphere of each category, and then a preset number of enhanced text training vectors corresponding to each category are generated using the SMOTE interpolation algorithm.

[0087] For example, take a hyperplane sphere and arbitrarily select two initial text training vectors within the hyperplane sphere, such as point A and point B. If the coordinates of point A are [0,0,0...,0] and the coordinates of point B are [x1,x2,x3,...,xn], after using the SMOTE interpolation method, the enhanced text training vector C can be obtained. The coordinates of point C are located on the line connecting points A and B, that is, the coordinates of point C are [a·x1,a·x2,a·x3,...,a·xn], where a is a constant. Similarly, a predetermined number of enhanced text training vectors can be obtained, and the enhanced text training vectors are data augmentation samples for the corresponding category of the hyperplane sphere. Perform the same operation on the hyperplane sphere corresponding to each category, and each category will obtain a corresponding enhanced text training vector.

[0088] It can be understood that the vector points corresponding to the enhanced text training vector obtained based on the hyperplane sphere are all contained in the hyperplane sphere.

[0089] The number of enhanced text training vectors obtained for each category may be 5 to 10 times the number of initial text training vectors for each category, and the number may be adjusted based on demand, which is not limited in this embodiment.

[0090] The text classification model training method provided in this embodiment uses the SMOTE synthetic minority class oversampling technique interpolation algorithm to generate enhanced text training vectors based on each class's initial text training vector and the corresponding hyperplane sphere. Because the SMOTE interpolation algorithm is used to generate enhanced text training vectors based on the hyperplane sphere corresponding to each class and the initial text training vector within the hyperplane sphere, the number of text training vectors per class is increased, and the newly added enhanced text training vectors are ensured to be distributed based on the hyperplane sphere corresponding to each class. This generates representative enhanced text training vectors for each class, rather than simply copying existing samples, which helps reduce the risk of overfitting.

[0091] As an optional implementation, based on any of the above embodiments, the preset text classification model is trained using the initial text training vector and the enhanced text training vector of each category, including:

[0092] Generate support sample sets and query sample sets based on each type of initial text training vector and enhanced text training vector;

[0093] Use the support sample set to train the preset text classification model, and use the query sample set to test the trained text classification model;

[0094] When the test pass condition is met, the text classification model that meets the test pass condition is determined as a text classification model that has been trained to convergence.

[0095] The support sample set is the data subset used to train the classification model, and the query sample set is the data subset used to evaluate the performance of the classification model.

[0096] Specifically, in this embodiment, the initial text training vectors for each class are divided into a support sample set and a query sample set according to a preset classification rule, and all the enhanced text training vectors corresponding to each class are divided into the support sample set. The support sample set is trained using a pre-set prototype network, and the loss function of the support sample set is calculated. When the loss function is within a preset range, the trained prototype network classification model is tested using the query sample set, and the loss function of the query sample set is calculated. When the loss function of the query set reaches minimum convergence, the text classification model is considered to have been trained to convergence.

[0097] Among them, when the text classification model is determined to be a text classification model that has been trained to convergence, it can be considered that the training of the text classification model is completed, and the center of the hyperplane sphere corresponding to each category in its text classification model is used as the category reference vector of each category.

[0098] Among them, the category benchmark vector is a typical sample vector representing a category, and the category benchmark vector corresponding to each category is different.

[0099] It can be understood that the loss function of the support sample set can be used to adjust the parameters of the classification model.

[0100] Optionally, the loss function may use a cross entropy loss function, or other loss functions may be used for calculation as required, which is not limited in this embodiment.

[0101] A loss function maps the value of a random event or its related random variables to a non-negative real number to represent the "risk" or "loss" of the random event. The cross-entropy cost function is a method used to measure the predicted value of an artificial neural network against the actual value.

[0102] It should be understood that for the division of the support set sample set and the query sample set, the initial text training vectors of each category can be divided into the support set sample set and the query sample set according to a certain ratio. For example, 80% of the initial text training vectors of each category can be divided into the support set sample set, and 20% of the initial text training vectors of each category can be divided into the query set sample set. It can also be divided according to other ratios as needed, which is not limited in this embodiment.

[0103] Among them, the Protoypal Network is a classification network based on prototype representation learning. Its basic idea is to generate a prototype vector for each category and determine the category to which the sample to be classified belongs by calculating the distance between the sample to be classified and the prototype vector of each category.

[0104] The text classification model training method provided in this embodiment uses each type of initial text training vector and enhanced text training vector to train a preset text classification model, including: generating a support sample set and a query sample set based on each type of initial text training vector and enhanced text training vector; using the support sample set to train the preset text classification model, and using the query sample set to test the trained text classification model; when the test pass condition is met, the text classification model that meets the test pass condition is determined to be a text classification model that has been trained to convergence. Since the text classification model is trained using each type of initial text training vector and enhanced text training vector, and is divided into a support sample set for training the classification model and a query sample set for evaluating the classification model, it helps to reduce the risk of overfitting, and the generalization ability of the text classification model can be verified through the query sample set, thereby timely discovering overfitting and making corresponding adjustments. In addition, testing and evaluating the text classification model through the query sample set can help accurately determine whether the text classification model has reached a convergence state and its performance in practical applications.

[0105] As an optional implementation manner, based on any of the above embodiments, before using each type of initial text training vector and enhanced text training vector to train the preset text classification model, the method further includes:

[0106] Determine the class basis vector corresponding to the center of each hyperplane sphere.

[0107] The enhanced text training vector of the corresponding category is sampled according to the category reference vector to obtain the sampled enhanced text training vector.

[0108] Among them, the category benchmark vector is a typical sample vector representing a category, and the category benchmark vector corresponding to each category is different.

[0109] Specifically, in this embodiment, the center of the hyperplane sphere corresponding to each category is determined as the category reference vector that can represent the category. Therefore, the category reference vector corresponding to each category is determined based on the vector point corresponding to the center of each hyperplane sphere.

[0110] It can be understood that there may be differences between the category benchmark vector in this embodiment and the category benchmark vector corresponding to each category after the text classification model is trained. When the text classification model is subsequently applied for text classification, the category benchmark vector used is the category benchmark vector determined according to the corresponding hyperplane sphere of each category after the text classification model is trained to convergence.

[0111] Specifically, in this embodiment, the vector corresponding to the center of each hyperplane sphere is used as the category reference vector of each category, and the enhanced text training vector of each category is sampled according to its corresponding category reference vector according to the preset sampling rule, so as to obtain the sampled enhanced text training vector.

[0112] The text classification model training method provided in this embodiment, before training the preset text classification model using the initial text training vector and the enhanced text training vector for each category, further includes: determining the category reference vector corresponding to the center of each hyperplane sphere; and sampling the enhanced text training vector for the corresponding category based on the category reference vector to obtain the sampled enhanced text training vector. Since the enhanced text training vector corresponding to each category is sampled based on the category reference vector, noise introduced when obtaining the enhanced text training vector can be reduced.

[0113] As an optional implementation, based on any of the above embodiments, sampling the enhanced text training vector of the corresponding category according to the category reference vector to obtain the sampled enhanced text training vector includes:

[0114] Calculate the cosine distance between each augmented text training vector and the category benchmark vector of the corresponding category.

[0115] The enhanced text training vector is sampled according to the cosine distance to obtain a sampled enhanced text training vector.

[0116] Cosine distance, also known as cosine similarity, uses the cosine of the angle between two vectors in vector space as a measure of the difference between two individuals. The smaller the cosine distance, the higher the similarity between the two texts.

[0117] Specifically, in this embodiment, the cosine distance between the enhanced text training vector corresponding to each category and its category reference vector determined based on the category hyperplane sphere is calculated respectively. The enhanced text training vector can be sampled in an ascending sampling order according to the cosine distance between the enhanced text training vector corresponding to each category and the corresponding category reference vector according to a preset sampling quantity, thereby obtaining the sampled enhanced text training vector.

[0118] The text classification model training method provided in this embodiment samples enhanced text training vectors of corresponding categories based on category benchmark vectors to obtain sampled enhanced text training vectors, including: calculating the cosine distance between each enhanced text training vector and the category benchmark vector of the corresponding category; and sampling the enhanced text training vectors based on the cosine distance to obtain the sampled enhanced text training vectors. Because the cosine distance between the enhanced text training vectors of each category is calculated based on the corresponding category benchmark vector, enhanced text training vectors with high similarity to each category can be sampled based on the characteristics of the cosine distance, thereby accurately reducing the impact of noise and irrelevant information in the enhanced text training vectors.

[0119] As an optional implementation, based on any of the above embodiments, the enhanced text training vector is sampled according to the cosine distance to obtain the sampled enhanced text training vector, including:

[0120] Calculate the corresponding sampling weight according to each cosine distance.

[0121] The enhanced text training vector is sampled based on the sampling weight and the preset weight threshold to obtain the sampled enhanced text training vector.

[0122] The sampling weight is used to calculate the probability of different strong text training vectors being sampled. The larger the sampling weight, the more likely it is to be sampled, and the smaller the sampling weight, the less likely it is to be sampled.

[0123] Specifically, in this embodiment, each cosine distance is normalized, and then sampling weights are calculated based on the normalized cosine distances according to a preset sampling weight formula. The sampling weights are then statistically analyzed, and a weight threshold is set based on a preset number and the characteristics of the sampling weights. Sampling is then performed based on the set weight threshold and the sampling weights of each enhanced text training vector to obtain the sampled enhanced text training vector.

[0124] Among them, the formula for sampling weight can be:

[0125] Specifically, w i represents the sampling weight of the enhanced text training vector point i, d i Represents the normalized cosine distance between the enhanced text training vector point i and the category reference vector of the corresponding category hyperplane sphere.

[0126] The number of samples can be half or one-third of the number of training vectors for each type of enhanced text. This is a fixed number, meaning the same number of training vectors for each type of enhanced text is collected. The number of samples can also be modified as needed and is not limited in this embodiment.

[0127] It can be understood that the setting of the weight threshold can be based on the number of enhanced text training vectors after sampling. By calculating the sampling weight of each enhanced text training vector, the sampling weight value that meets the required sampling number is used as the weight threshold, and then the enhanced text training vector can be sampled according to this weight threshold.

[0128] The text classification model training method provided in this embodiment samples the enhanced text training vector according to the cosine distance to obtain the sampled enhanced text training vector, including: calculating the corresponding sampling weight according to each cosine distance; sampling the enhanced text training vector based on the sampling weight and the preset weight threshold to obtain the sampled enhanced text training vector. Since the corresponding sampling weight is calculated according to the cosine distance of each enhanced text training vector, and the enhanced text training vector of each category is sampled according to the sampling weight according to the preset weight threshold. The enhanced text training vector corresponding to each category after screening by the weight threshold will focus more on the text with a higher similarity to the category benchmark vector, thereby helping the text classification model to better learn and understand the characteristics of each type of text, thereby improving the generalization ability of the model.

[0129] As an optional implementation, based on any of the above embodiments, the enhanced text training vector is sampled based on the sampling weight and the preset weight threshold to obtain the sampled enhanced text training vector, including:

[0130] In response to the sampling weight being greater than or equal to the preset weight threshold, the corresponding enhanced text training vector is retained as the sampled enhanced text training vector.

[0131] Specifically, in this embodiment, according to a preset weight threshold, the enhanced text training vectors whose sampling weights are greater than or equal to the preset sampling weight threshold are retained as sampled enhanced text training vectors, and thus used for subsequent training of the text training model.

[0132] It is understandable that for those enhanced text training vectors whose sampling weights are less than the preset weight threshold, they will not be put into the text classification model for training and will be deleted.

[0133] The text classification model training method provided in this embodiment samples enhanced text training vectors based on sampling weights and a preset weight threshold to obtain sampled enhanced text training vectors, including: in response to the sampling weight being greater than or equal to the preset weight threshold, retaining the corresponding enhanced text training vector as the sampled enhanced text training vector. Since enhanced text training vectors with sampling weights greater than or equal to the preset weight threshold are used as sampled enhanced text training vectors and input into the text classification model for training, while enhanced text training vectors with sampling weights less than the preset weight threshold are deleted, the amount of data processed by the text classification model during training is reduced, the training speed of the text classification model is accelerated, the consumption of computing resources is reduced, and the overall training efficiency is improved.

[0134] FIG3 is a flowchart of a text classification model training method provided by another embodiment of the present application. As shown in FIG3 , the execution subject of this embodiment is a text classification model training device. The text classification model training method provided by this embodiment includes the following steps:

[0135] S301, obtaining initial text training samples.

[0136] S302: Encode the initial text training sample to obtain an initial text training vector.

[0137] S303: Divide the initial text training vector of each category into a support sample set and a query sample set.

[0138] S304: Using the minimum bounding sphere algorithm to calculate the hyperplane sphere corresponding to each category of the support sample set corresponding to each category of the initial text training vector, and determining the center of the hyperplane sphere of each category as the category reference vector of each category.

[0139] S305 , based on the support sample set and the corresponding hyperplane sphere corresponding to the initial text training vector of each category, an enhanced text training vector is generated using the SMOTE interpolation algorithm.

[0140] Among them, the full name of SMOTE is Synthetic Minority Oversampling Technique, which is synthetic minority oversampling technology.

[0141] S306 , calculating the cosine distance between each enhanced text training vector and the category reference vector of the corresponding category.

[0142] S307: Calculate the sampling weight corresponding to each enhanced text training vector according to the cosine distance.

[0143] S308 : Sampling the enhanced text training vector according to the sampling weight and a preset weight threshold to obtain the sampled enhanced text training vector.

[0144] S309 , training the text classification model using the sampled enhanced text training vectors of each category and the support samples corresponding to the initial text training vectors of each category, and calculating the corresponding loss function.

[0145] S310, when the loss function corresponding to the sampled enhanced text training vector and each type of initial text training vector reaches a preset threshold, the text classification model is tested using the query sample set corresponding to each type of initial text training vector, and the corresponding loss function is calculated.

[0146] S311, when the loss function is minimized, the text classification model is determined to be a text classification model that has been trained to convergence, and the training of the text classification model is stopped.

[0147] S312: The center of the sphere corresponding to the hyperplane sphere in the text classification model that has been trained to convergence is used as the final category reference vector.

[0148] The final category benchmark vector is used for text classification in practical applications.

[0149] FIG4 is a flowchart of a text classification method provided by an embodiment of the present application. As shown in FIG4 , the execution subject of this embodiment is a text classification device, which is located in an electronic device. The text classification method provided by this embodiment includes the following steps:

[0150] S401: Obtain target text to be classified.

[0151] The target text is the text to be classified, and the expression can be any one or more of a sentence and a question.

[0152] Specifically, in this embodiment, the target text to be classified obtained may be the text to be classified obtained in real time when the text classification model is applied.

[0153] For example, if applied to a car company's APP, when a car company's APP user posts a text in the car company's APP, the corresponding text can be obtained in real time and classified.

[0154] S402: Encode the target text to obtain a target text vector.

[0155] Specifically, in this embodiment, the target text is encoded by a preset encoder to convert the target text into a target text vector.

[0156] Among them, the preset encoder is the same encoder used in text classification model training, which can be a fully connected neural network.

[0157] S403: Classify the target text vector using the text classification model that has been trained to convergence to obtain a corresponding category.

[0158] Specifically, in this embodiment, the target text vector is input into a text classification model that has been trained to convergence. By calculating the vector distance between the target text vector and the category reference vector of each category, the category reference vector with the shortest vector distance to the target text vector is determined, thereby determining the text category corresponding to the category reference vector, and then outputting the determined text category as the classification result of the target text.

[0159] The vector distance may be a cosine distance.

[0160] In this embodiment, the implementation method of each step is similar to the implementation method of the corresponding solution in the above embodiment, and will not be repeated here.

[0161] Figure 5 is a structural diagram of a text classification model training device provided in an embodiment of the present application. As shown in Figure 5, the text classification model training device provided in this embodiment is located in an electronic device, and the text classification model training device 50 provided in this embodiment includes: an acquisition module 51, an encoding module 52, a calculation module 53, a determination module 54, and a training module 55.

[0162] Among them, the acquisition module 51 is used to obtain the initial text training sample; the encoding module 52 is used to encode the initial text training sample to obtain the initial text training vector; the calculation module 53 is used to use the minimum bounding sphere algorithm to calculate the hyperplane sphere corresponding to each type of initial text training vector for each type of initial text training vector; the determination module 54 is used to determine the enhanced text training vector based on each type of initial text training vector and the corresponding hyperplane sphere; the training module 55 is used to train the preset text classification model using each type of initial text training vector and the enhanced text training vector.

[0163] The text classification model training device provided in this embodiment can execute the method embodiment shown in Figure 2. The specific implementation principles and technical effects are similar and will not be repeated here.

[0164] Optionally, when determining the enhanced text training vector based on each type of initial text training vector and the corresponding hyperplane sphere, the determination module 54 is specifically configured to:

[0165] The SMOTE synthetic minority class oversampling technique interpolation algorithm is used to generate enhanced text training vectors based on the initial text training vectors of each class and the corresponding hyperplane sphere.

[0166] Optionally, the training module 55, when using each type of initial text training vector and enhanced text training vector to train the preset text classification model, is specifically configured to:

[0167] A support sample set and a query sample set are generated based on each category of initial text training vectors and enhanced text training vectors; the support sample set is used to train a preset text classification model, and the query sample set is used to test the trained text classification model; when the test pass conditions are met, the text classification model that meets the test pass conditions is determined as a text classification model that has been trained to convergence.

[0168] Optionally, the text classification model training device provided in this embodiment further includes: a sampling module.

[0169] The determination module 54 is further configured to determine the category reference vector corresponding to the center of each hyperplane sphere before training the preset text classification model using the initial text training vector and the enhanced text training vector for each category. The sampling module is specifically configured to sample the enhanced text training vector for the corresponding category based on the category reference vector to obtain the sampled enhanced text training vector.

[0170] Optionally, the sampling module samples the enhanced text training vector of the corresponding category according to the category reference vector to obtain the sampled enhanced text training vector, and is specifically used to: calculate the cosine distance between each enhanced text training vector and the category reference vector of the corresponding category; and sample the enhanced text training vector according to the cosine distance to obtain the sampled enhanced text training vector.

[0171] Optionally, the sampling module samples the enhanced text training vector according to the cosine distance to obtain the sampled enhanced text training vector, and is specifically used to: calculate the corresponding sampling weight according to each cosine distance; and sample the enhanced text training vector based on the sampling weight and a preset weight threshold to obtain the sampled enhanced text training vector.

[0172] Optionally, the sampling module samples the enhanced text training vector based on the sampling weight and the preset weight threshold to obtain the sampled enhanced text training vector, and is specifically used to: in response to the sampling weight being greater than or equal to the preset weight threshold, retain the corresponding enhanced text training vector as the sampled enhanced text training vector.

[0173] Optionally, the preset text classification model is a preset prototype network.

[0174] The text classification model training device provided in this embodiment can execute any of the method embodiments shown above. The specific implementation principles and technical effects are similar and will not be repeated here.

[0175] FIG6 is a schematic diagram of the structure of a text classification device provided in one embodiment of the present application. As shown in FIG6 , the text classification device provided in this embodiment is located in an electronic device, which may be a text classification device. The text classification device 60 provided in this embodiment includes an acquisition module 61, an encoding module 62, and a classification module 63.

[0176] Among them, the acquisition module 61 is used to obtain the target text to be classified; the encoding module 62 is used to encode the target text to obtain a target text vector; the classification module 63 is used to classify the target text vector using a text classification model that has been trained to convergence to obtain a corresponding category.

[0177] FIG7 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. As shown in FIG7 , the electronic device 70 provided in this embodiment includes: a processor 71 and a memory 72 communicatively connected to the processor.

[0178] Memory 71 stores computer-executable instructions; processor 71 executes the computer-executable instructions stored in memory 72 to implement the text classification model training method and classification method provided in any of the above-described embodiments. For further explanation, please refer to the corresponding descriptions and effects of the steps in the accompanying drawings, and detailed descriptions are omitted here.

[0179] The program may include program code, which includes computer-executable instructions. The memory 72 may include a high-speed RAM memory, or may also include a non-volatile memory, such as at least one disk memory.

[0180] In this embodiment, the memory 72 and the processor 71 are connected via a bus. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, FIG7 shows only one thick line, but this does not imply that there is only one bus or only one type of bus.

[0181] The present application also provides a computer-readable storage medium storing computer-executable instructions. When executed by a processor, the computer-executable instructions implement the text classification model training method and classification method provided in any of the above-described embodiments. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0182] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the text classification model training method and classification method provided in any one of the above embodiments.

[0183] It can be explained that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art will know that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art will also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily applicable to this application.

[0184] It can be explained that although the various steps in the flowchart are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowchart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these sub-steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0185] It will be understood that the above-described device embodiments are merely illustrative, and the device of the present application may also be implemented in other ways. For example, the division of units / modules in the above-described embodiments is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units, modules, or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0186] In addition, unless otherwise specified, the functional units / modules in the various embodiments of the present application may be integrated into a single unit / module, each unit / module may exist physically separately, or two or more units / modules may be integrated together. The aforementioned integrated units / modules may be implemented in the form of hardware or software program modules.

[0187] If the integrated unit / module is implemented in hardware, the hardware may be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the artificial intelligence processor may be any appropriate hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC. Unless otherwise specified, the storage unit may be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.

[0188] If the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0189] In the above embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined in any way. To keep the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0190] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0191] Those skilled in the art will appreciate that all or part of the steps in the above method can be completed by instructing relevant hardware (such as a processor) through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk or an optical disk. Optionally, all or part of the steps in the above embodiment can also be implemented using one or more integrated circuits. Accordingly, each module / unit in the above embodiment can be implemented in the form of hardware, for example, by implementing its corresponding functions through an integrated circuit, or in the form of a software functional module, for example, by executing a program / instruction stored in a memory by a processor to implement its corresponding function. This application is not limited to any particular form of combination of hardware and software.

[0192] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A text classification model training method, comprising: Obtaining an initial text training sample, and encoding the initial text training sample to obtain an initial text training vector; For each type of initial text training vector, the minimum bounding sphere algorithm is used to calculate the hyperplane sphere corresponding to each type of initial text training vector; Determine an enhanced text training vector based on each type of initial text training vector and the corresponding hyperplane sphere; The preset text classification model is trained using the initial text training vector and the enhanced text training vector for each category.

2. The method according to claim 1, wherein: The step of determining the enhanced text training vector based on each type of initial text training vector and the corresponding hyperplane sphere includes: The SMOTE synthetic minority class oversampling technique interpolation algorithm is used to generate enhanced text training vectors based on the initial text training vectors of each class and the corresponding hyperplane sphere.

3. The method according to claim 1, wherein: The training of a preset text classification model using the initial text training vector and the enhanced text training vector of each type includes: Generate a support sample set and a query sample set based on each type of initial text training vector and enhanced text training vector; Using the support sample set to train the preset text classification model, and using the query sample set to test the trained text classification model; When the test pass condition is met, the text classification model that meets the test pass condition is determined as a text classification model that has been trained to convergence.

4. The method according to any one of claims 1 to 3, further comprising: before training a preset text classification model using the initial text training vector and the enhanced text training vector for each category; Determine the category reference vector corresponding to the center of each hyperplane sphere; The enhanced text training vector of the corresponding category is sampled according to the category reference vector to obtain the sampled enhanced text training vector.

5. The method according to claim 4, wherein: The sampling process of the enhanced text training vector of the corresponding category according to the category reference vector to obtain the sampled enhanced text training vector includes: Calculate the cosine distance between each augmented text training vector and the category benchmark vector of the corresponding category; The enhanced text training vector is sampled according to the cosine distance to obtain a sampled enhanced text training vector.

6. The method according to claim 5, wherein: The sampling process of the enhanced text training vector according to the cosine distance to obtain the sampled enhanced text training vector includes: Calculate the corresponding sampling weight according to each cosine distance; The enhanced text training vector is sampled based on a sampling weight and a preset weight threshold to obtain a sampled enhanced text training vector.

7. The method according to claim 6, wherein: The sampling process of the enhanced text training vector based on the sampling weight and the preset weight threshold to obtain the sampled enhanced text training vector includes: In response to the sampling weight being greater than or equal to the preset weight threshold, the corresponding enhanced text training vector is retained as the sampled enhanced text training vector.

8. The method according to any one of claims 1 to 3, wherein: The preset text classification model is a preset prototype network.

9. A text classification method comprising: Get the target text to be classified; Encoding the target text to obtain a target text vector; The target text vector is classified using a text classification model that has been trained to convergence to obtain a corresponding category, and the text classification model that has been trained to convergence is obtained by training a preset text classification model using any one of the methods described in items 1-8.

10. A text classification model training device, comprising: An acquisition module is used to obtain initial text training samples; An encoding module, configured to perform encoding processing on the initial text training sample to obtain an initial text training vector; A calculation module is used to calculate the hyperplane sphere corresponding to each type of initial text training vector using a minimum bounding sphere algorithm; a determination module, configured to determine an enhanced text training vector based on each type of initial text training vector and the corresponding hyperplane sphere; The training module is used to train a preset text classification model using the initial text training vector and the enhanced text training vector for each category.

11. A text classification device comprising: An acquisition module is used to obtain the target text to be classified; An encoding module, configured to encode the target text to obtain a target text vector; The classification module is used to classify the target text vector using a text classification model that has been trained to convergence to obtain a corresponding category.

12. An electronic device comprising: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 9.

13. A computer-readable storage medium, wherein computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of claims 1 to 9.

14. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Weighted hyper-sphere support vector machine algorithm based image classification method

    CN104112143A

  • Text classification model training method, text classification method, text classification device, text classification equipment, medium and product

    CN118051616A

  • Semi-supervised method and apparatus for public opinion text analysis

    US20230351212A1