Method and apparatus for generating a classification model

By obtaining and determining the statistical features of the training samples and using the pre-trained model set to generate a text classification model, the problem of manual parameter adjustment and complex feature engineering is solved in the existing technology that generates text classification models, and the rapid generation of text classification models suitable for the needs is achieved.

CN110457476BActive Publication Date: 2025-06-27BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201910721353.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-08-06
Publication Date
2025-06-27
Estimated Expiration
2039-08-06

AI Technical Summary

Technical Problem

When generating text classification models, the prior art requires manual parameter adjustment and complex feature engineering, resulting in high technical thresholds and it is difficult to quickly generate text classification models that suit their own needs.

Method used

By obtaining the training sample set, the statistical features of the training sample are determined, including formal features that characterize text space, and trained based on these features and the initial models selected from the preset pre-trained model set to generate a text classification model.

Benefits of technology

It realizes the generation of text classification models without manual parameter adjustment, lowers the threshold for users, and can quickly implement text classification applications that meet their own needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110457476B_ABST
    Figure CN110457476B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method and apparatus for generating a classification model. A specific implementation of the method includes: obtaining a training sample set, where the training samples include sample texts and corresponding sample categories; determining statistical features of the training sample set, where the statistical features include formal features representing the text length; generating a text classification model based on the statistical features and training of an initial model, where the text classification model is used to represent the correspondence between text categories and texts, and the initial model is selected from a preset set of pre-trained models. This implementation realizes the automatic generation of a text classification model based on cloud computing technology without the need for manual parameter tuning by the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly, to methods and apparatuses for generating classification models. Background Art

[0002] With the rapid development of artificial intelligence technology (AI) and Internet technology, in the face of the rapidly growing massive text information, how to effectively classify text is an important prerequisite for subsequent content search and information value extraction.

[0003] There are generally two related methods: One is to use feature engineering technology to extract text features, and then classify the text according to the similarity between the extracted features. The other is to use a trained automated text classification model to classify the text. Summary of the Invention

[0004] Embodiments of the present disclosure propose methods and apparatuses for generating classification models.

[0005] In a first aspect, embodiments of the present disclosure provide a method for generating a classification model, the method including: obtaining a training sample set, where a training sample includes a sample text and a sample category corresponding to the sample text; determining statistical features of the training sample set, where the statistical features include formal features representing the text length; generating a text classification model based on the statistical features and training of an initial model, where the text classification model is used to represent the correspondence between the text category and the text, and the initial model is selected from a preset set of pre-trained models.

[0006] In some embodiments, generating a text classification model based on the statistical features and the initial model includes: inputting the formal features into a pre-trained hyperparameter generation model to obtain a hyperparameter group corresponding to the formal features, where the hyperparameter group includes model hyperparameters and training hyperparameters; selecting a pre-trained model that matches the model hyperparameters in the hyperparameter group from the set of pre-trained models as the initial model; training the initial model according to the training hyperparameters in the hyperparameter group to generate a text classification model.

[0007] In some embodiments, generating a text classification model based on the statistical features and the initial model includes: obtaining a set of initial hyperparameter groups, where the initial hyperparameter groups include initial model hyperparameters and initial training hyperparameters; selecting an initial hyperparameter group from the set of initial hyperparameter groups, and performing the following determination steps: selecting a pre-trained model that matches the initial model hyperparameters in the selected initial hyperparameter group from a set of pre-trained models as the initial model; training the selected initial model according to the initial training hyperparameters in the selected initial hyperparameter group to generate a quasi-text classification model corresponding to the initial hyperparameter group; evaluating the generated quasi-text classification model based on a validation text set to generate an evaluation result; in response to determining that the generated evaluation result meets the hyperparameter group determination condition, determining a text classification model from the quasi-text classification models corresponding to the evaluation results that meet the hyperparameter group determination condition; in response to determining that the generated evaluation result does not meet the hyperparameter group determination condition, updating the initial hyperparameter groups in the set of initial hyperparameter groups; selecting an initial hyperparameter group from the updated set of initial hyperparameter groups, and continuing to perform the determination steps.

[0008] In some embodiments, the above-mentioned statistical features further include content features characterizing the text content; and the above-mentioned selecting a pre-trained model that matches the initial model hyperparameters in the selected initial hyperparameter group from a set of pre-trained models as the initial model includes: selecting a pre-trained model that matches the content features and the initial model hyperparameters in the selected initial hyperparameter group from a set of pre-trained models as the initial model, where the pre-trained model corresponds to a semantic label.

[0009] In some embodiments, the above-mentioned training the selected initial model according to the initial training hyperparameters in the selected initial hyperparameter group to generate a quasi-text classification model corresponding to the initial hyperparameter group includes: selecting a training sample from a set of training samples, and performing the following training steps: inputting the sample text of the selected training sample into the selected initial model to generate a text category; determining a difference value according to the generated text category and the sample category corresponding to the input sample text; determining whether the difference value meets the training completion condition, where the difference value and the training completion condition are determined based on the initial training hyperparameters in the selected hyperparameter group; in response to determining that the training completion condition is met, determining the selected initial model as the quasi-text classification model corresponding to the selected hyperparameter group; in response to determining that the training completion condition is not met, adjusting the relevant parameters of the selected initial model, and selecting a training sample from the set of training samples, using the adjusted initial model as the selected initial model, and continuing to perform the training steps.

[0010] In some embodiments, the obtaining of the training sample set includes: receiving an annotated text set sent by a client, where the annotated text includes text and text category annotation information corresponding to the text; partitioning the annotated text set to generate a training sample set and a validation text set, where a training sample includes text as a sample text and text category annotation information as a sample category corresponding to the sample text.

[0011] In some embodiments, the method further includes: receiving a text set to be classified sent by a client; inputting the text set to be classified into a text classification model to generate category information corresponding to the text to be classified in the text set to be classified, where the category information is used to represent the category to which the text to be classified belongs, and the category information matches the sample category; sending the generated category information and the corresponding text information to be classified to the client, where the text information to be classified is used to identify the text to be classified in the text set to be classified.

[0012] In a second aspect, an embodiment of the present disclosure provides an apparatus for generating a classification model. The apparatus includes: an obtaining unit configured to obtain a training sample set, where a training sample includes a sample text and a sample category corresponding to the sample text; a determining unit configured to determine statistical features of the training sample set, where the statistical features include formal features representing the text length; a first generating unit configured to generate a text classification model based on the statistical features and training of an initial model, where the text classification model is used to represent the correspondence between a text category and a text, and the initial model is selected from a preset set of pre-trained models.

[0013] In some embodiments, the above first generating unit includes: a first generating module configured to input the formal features into a pre-trained hyperparameter generation model to obtain a hyperparameter group corresponding to the formal features, where the hyperparameter group includes model hyperparameters and training hyperparameters; a selecting module configured to select a pre-trained model matching the model hyperparameters in the hyperparameter group from the set of pre-trained models as the initial model; a second generating module configured to train the initial model according to the training hyperparameters in the hyperparameter group to generate a text classification model.

[0014] In some embodiments, the above-mentioned first generation unit includes: an acquisition module configured to acquire an initial set of hyperparameter groups, where the initial hyperparameter groups include initial model hyperparameters and initial training hyperparameters; a determination module configured to select an initial hyperparameter group from the initial set of hyperparameter groups and perform the following determination steps: select a pre-trained model that matches the initial model hyperparameters in the selected initial hyperparameter group from a set of pre-trained models as the initial model; train the selected initial model according to the initial training hyperparameters in the selected initial hyperparameter group to generate a quasi-text classification model corresponding to the initial hyperparameter group; evaluate the generated quasi-text classification model based on a validation text set to generate an evaluation result; in response to determining that the generated evaluation result meets the hyperparameter group determination condition, determine a text classification model from the quasi-text classification models corresponding to the evaluation results that meet the hyperparameter group determination condition; an update module configured to, in response to determining that the generated evaluation result does not meet the hyperparameter group determination condition, update the initial hyperparameter groups in the initial set of hyperparameter groups; select an initial hyperparameter group from the updated initial set of hyperparameter groups and continue to perform the determination steps.

[0015] In some embodiments, the above-mentioned statistical features further include content features characterizing the text content; the determination module is further configured to: select a pre-trained model that matches the content features and the initial model hyperparameters in the selected initial hyperparameter group from a set of pre-trained models as the initial model, where the pre-trained model corresponds to a semantic label.

[0016] In some embodiments, the above-mentioned determination module further includes: a selection sub-module configured to select a training sample from a set of training samples and perform the following training steps: input the sample text of the selected training sample into the selected initial model to generate a text category; determine a difference value according to the generated text category and the sample category corresponding to the input sample text; determine whether the difference value meets the training completion condition, where the difference value and the training completion condition are determined based on the initial training hyperparameters in the selected hyperparameter group; in response to determining that the training completion condition is met, determine the selected initial model as the quasi-text classification model corresponding to the selected hyperparameter group; an adjustment sub-module configured to, in response to determining that the training completion condition is not met, adjust the relevant parameters of the selected initial model, select a training sample from the set of training samples, use the adjusted initial model as the selected initial model, and continue to perform the training steps.

[0017] In some embodiments, the obtaining unit includes: a receiving module configured to receive a set of labeled texts sent by a client, where the labeled text includes text and text category labeling information corresponding to the text; a third generating module configured to divide the set of labeled texts to generate a training sample set and a verification text set, where the training sample includes the text as the sample text and the text category labeling information as the sample category corresponding to the sample text.

[0018] In some embodiments, the apparatus further includes: a receiving unit configured to receive a set of texts to be classified sent by a client; a second generating unit configured to input the set of texts to be classified into a text classification model to generate category information corresponding to the texts to be classified in the set of texts to be classified, where the category information is used to represent the category to which the text to be classified belongs, and the category information matches the sample category; a sending unit configured to send the generated category information and the corresponding text information to be classified to the client, where the text information to be classified is used to identify the text to be classified in the set of texts to be classified.

[0019] In a third aspect, an embodiment of the present disclosure provides a server, which includes: one or more processors; a storage device on which one or more programs are stored; when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.

[0020] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.

[0021] The method and apparatus for generating a classification model provided by the embodiments of the present disclosure first obtain a training sample set. Among them, the training sample includes a sample text and a sample category corresponding to the sample text. Then, the statistical features of the training sample set are determined. Among them, the statistical features include formal features representing the text length. After that, based on the statistical features and the training of the initial model, a text classification model is generated. Among them, the text classification model is used to represent the corresponding relationship between the text category and the text. The above initial model is selected from a preset set of pre-trained models. Thus, the automatic generation of the text classification model can be achieved without manual parameter tuning. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present disclosure will become more apparent:

[0023] Figure 1 is an exemplary system architecture diagram to which an embodiment of the present disclosure can be applied;

[0024] Figure 2 is a flowchart of an embodiment of a method for generating a classification model according to the present disclosure;

[0025] Figure 3 is a schematic diagram of an application scenario of a method for generating a classification model according to an embodiment of the present disclosure;

[0026] Figure 4 is a flowchart of another embodiment of a method for generating a classification model according to the present disclosure;

[0027] Figure 5 is a schematic structural diagram of an embodiment of an apparatus for generating a classification model according to the present disclosure;

[0028] Figure 6 is a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. Detailed Embodiments

[0029] The present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the relevant invention and not for limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.

[0030] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The present disclosure will be described in detail below with reference to the drawings and embodiments.

[0031] Figure 1 Illustrates an exemplary architecture 100 to which the method for generating a classification model or the apparatus for generating a classification model according to the present disclosure can be applied.

[0032] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0033] The terminal devices 101, 102, 103 interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as web browser applications, search applications, instant messaging tools, email clients, social platform software, text editing applications, reading applications, etc.

[0034] The terminal devices 101, 102, and 103 can be either hardware or software. When the terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with a display screen and supporting text display, including but not limited to smartphones, tablet computers, e-book readers, laptop computers, desktop computers, and so on. When the terminal devices 101, 102, and 103 are software, they can be installed in the above-listed electronic devices. It can be implemented as multiple software or software modules (such as software or software modules for providing distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0035] The server 105 can be a server that provides various services, such as a background server that provides support for text classification applications on the terminal devices 101, 102, and 103. Optionally, the server 105 can also be a cloud server. The background server can train a text classification model based on the obtained training sample set, and can analyze and process the text sent by the terminal device, and feedback the processing result (such as text category information) to the terminal device.

[0036] It should be noted that the above training sample set can also be directly stored locally on the server 105. The server 105 can directly extract the training sample set stored locally and perform model training. At this time, the terminal devices 101, 102, and 103 and the network 104 may not exist.

[0037] It should be noted that the server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (such as software or software modules for providing distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0038] It should be noted that the method for generating a classification model provided by the embodiments of the present disclosure is generally executed by the server 105. Correspondingly, the device for generating a classification model is generally set in the server 105.

[0039] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0040] Continuing to refer to Figure 2 , a flowchart 200 of an embodiment of the method for generating a classification model according to the present disclosure is shown. The method for generating a classification model includes the following steps:

[0041] Step 201, obtain a training sample set.

[0042] In this embodiment, the execution subject of the method for generating a classification model (such as Figure 1 the server 105 shown) can obtain a training sample set through a wired connection method or a wireless connection method. Among them, the above training samples may include sample texts and corresponding sample categories. Specifically, the above execution subject can obtain a training sample set pre-stored locally, or obtain a training sample set sent by an electronic device (such as Figure 1 the terminal device shown) communicatively connected thereto.

[0043] In practice, the above training samples can be obtained in various ways. As an example, technicians can label the categories of each text in the obtained text set. The text is associated and stored with the labeled category, and finally a training sample is obtained. As another example, the information resources on the portal website can be processed. For example, the articles on the web page can be used as sample texts, and the columns to which the articles belong can be used as sample categories to form training samples. A large number of training samples are formed through a large amount of data, and then a training sample set is composed.

[0044] In some alternative implementation manners of this embodiment, the above execution subject can also obtain a training sample set according to the following steps:

[0045] The first step is to receive the labeled text set sent by the user terminal.

[0046] In these implementation manners, the above execution subject can receive the labeled text set uploaded by the user terminal (such as Figure 1 the terminal device shown). Among them, the above labeled text may include the text and the text category annotation information corresponding to the text.

[0047] The second step is to divide the labeled text set to generate a training sample set and a validation text set.

[0048] In these implementation manners, the above execution subject can divide the labeled text set received in the first step according to a certain ratio to obtain a training sample set and a validation text set. Among them, the above training samples may include the text as the sample text and the text category annotation information as the sample category corresponding to the sample text. Usually, the above ratio can be preset, such as 8:2 or 7:3. Optionally, the above ratio can also be set according to the user's selection.

[0049] Step 202, determine the statistical characteristics of the training sample set.

[0050] In this embodiment, the above-mentioned execution entity can determine the statistical features of the training sample set in various ways. Among them, the above-mentioned statistical features may include formal features representing the text length. The above-mentioned statistical features may include, but are not limited to, at least one of the following: the number of categories of sample categories, the maximum text length, the minimum text length, and the average text length.

[0051] In some alternative implementation manners of this embodiment, the above-mentioned statistical features may further include content features representing the text content. The above-mentioned statistical features may include, but are not limited to, at least one of the following: word frequency vectors, text feature vectors determined based on a vector space model (VSM).

[0052] Step 203: Generate a text classification model based on the statistical features and the training of the initial model.

[0053] In this embodiment, based on the statistical features and the training of the initial model, the above-mentioned execution entity can generate a text classification model in various ways. Among them, the above-mentioned text classification model can be used to represent the correspondence between text categories and texts. The above-mentioned initial model can be selected from a preset pre-trained model set. The pre-trained models in the above-mentioned pre-trained model set can be models pre-trained based on massive data sets in different fields (such as finance, law, technology, sports). The above-mentioned pre-trained models can be used to represent the correspondence between text categories and texts. The above-mentioned pre-trained models can contain sufficient basic semantics, reasoning, and other knowledge in related fields. It can be understood that the above-mentioned pre-trained models generated based on massive data sets in different fields can be regarded as the initial representations of model-agnostic meta-learning (MAML) in different fields respectively. Thus, after subsequent training and adjustment of the model results, the above-mentioned pre-trained models can be used for text classification in sub-fields of the fields to which the above-mentioned massive data sets belong.

[0054] In some alternative implementation manners of this embodiment, the above-mentioned execution entity can generate a text classification model according to the following steps:

[0055] The first step: Input the formal features into a pre-trained hyper-parameters generation model to obtain a set of hyper-parameters corresponding to the formal features.

[0056] In these alternative implementation manners, the above-mentioned execution entity may input the formal features determined in step 202 into a pre-trained hyperparameter generation model to obtain a hyperparameter group corresponding to the formal features. Among them, the above-mentioned hyperparameter group may include model hyperparameters and training hyperparameters. The above-mentioned model hyperparameters may be used to characterize the attributes of the model itself. For example, the above-mentioned model hyperparameters may include, but are not limited to, at least one of the following: the number of layers of the neural network, the number of nodes in each hidden layer, the dimension of the word vector (embedding). The above-mentioned training hyperparameters may be used to indicate the training process of the model. For example, the above-mentioned training hyperparameters may include, but are not limited to, at least one of the following: learning rate, batch size, clip c, dropout value (such as 0.5), L2 regularization value (such as 1.0).

[0057] In these alternative implementation manners, the above-mentioned hyperparameter generation model may be used to characterize the correspondence between the hyperparameter group and the formal features. As an example, the above-mentioned hyperparameter generation model may be a correspondence table pre-specified by a person skilled in the art based on the statistics of a large number of formal features and corresponding hyperparameter groups with good training effects, and storing the correspondence between multiple formal features and hyperparameter groups. As another example, the above-mentioned hyperparameter generation model may be a model trained and generated using a machine learning algorithm based on a large number of samples. Among them, the above-mentioned samples may be composed of formal features and corresponding hyperparameter groups with good training effects.

[0058] Second, select a pre-trained model that matches the model hyperparameters in the hyperparameter group from the pre-trained model set as the initial model.

[0059] In these implementation manners, the above-mentioned execution entity may select a pre-trained model whose model structure matches the model hyperparameters in the hyperparameter group obtained in the first step from the pre-trained model set as the initial model. Among them, the above-mentioned matching may include being the same or similar. For example, if the above-mentioned model hyperparameter is that the number of hidden layers is 2, then the above-mentioned initial model may be a neural network including 2 hidden layers.

[0060] Third, train the initial model according to the training hyperparameters in the hyperparameter group to generate a text classification model.

[0061] In these implementation manners, the above-mentioned execution entity may, according to the indication of the training hyperparameters in the hyperparameter group obtained in the first step, use various machine learning algorithms to train the initial model selected in the second step to generate a text classification model.

[0062] In some alternative implementation manners of this embodiment, the above-mentioned execution entity may also generate a text classification model according to the following steps:

[0063] Step 1: Obtain the initial set of hyperparameter groups.

[0064] In these implementation manners, the above-mentioned execution entity can first obtain the initial set of hyperparameter groups through a wired connection method or a wireless connection method. Among them, the above-mentioned initial hyperparameter group may include initial model hyperparameters and initial training hyperparameters. The relevant descriptions of the initial model hyperparameters and initial training hyperparameters in the above-mentioned initial hyperparameter group may be consistent with the model hyperparameters and training hyperparameters in the foregoing hyperparameter group, and will not be elaborated here.

[0065] In these implementation manners, the above-mentioned execution entity can obtain the initial set of hyperparameter groups pre-stored locally, or can obtain the initial set of hyperparameter groups sent by an electronic device (such as a data server) communicatively connected thereto. Optionally, the above-mentioned execution entity can also obtain the initial set of hyperparameter groups that match the statistical features from the local or the above-mentioned electronic device. For example, there may be a matching relationship between the average text length in the statistical features and the training times threshold in the initial hyperparameter group. Optionally, the above-mentioned execution entity can also randomly generate the initial values corresponding to each initial hyperparameter group in the initial hyperparameter group, and then generate the initial set of hyperparameter groups.

[0066] Step 2: Select an initial hyperparameter group from the initial set of hyperparameter groups, and perform the following determination steps. The above-mentioned determination steps may include:

[0067] S1. Select a pre-trained model that matches the initial model hyperparameters in the selected initial hyperparameter group from the set of pre-trained models as the initial model.

[0068] In these implementation manners, the above-mentioned execution entity can select one or more pre-trained models that match the initial model hyperparameters in the selected initial hyperparameter group from the set of pre-trained models as the initial model. Optionally, the above-mentioned execution entity can also select one or more pre-trained models that match the statistical features as the initial model. For example, there may be a matching relationship between the maximum text length in the statistical features and the number of hidden layer nodes in the initial hyperparameter group.

[0069] Optionally, based on the content features in the above-mentioned statistical features, the above-mentioned execution entity can also select one or more pre-trained models that match the content features and the initial model hyperparameters in the selected initial hyperparameter group from the set of pre-trained models as the initial model. Among them, the pre-trained model may correspond to a semantic label. Generally, the above-mentioned semantic label may be consistent with the domain of the dataset on which the pre-trained model in the above-mentioned set of pre-trained models is based during the pre-training process. The above-mentioned execution entity can determine the matching degree between the above-mentioned content features and the semantic label in various ways. As an example, the above-mentioned execution entity can calculate the similarity between the word frequency vector of the above-mentioned training sample set and the above-mentioned semantic label.

[0070] S2. Train the selected initial model according to the initial training hyperparameters in the selected initial hyperparameter group to generate a quasi-text classification model corresponding to the initial hyperparameter group.

[0071] Based on the above optional implementation manners, the above execution entity can, according to the indication of the initial training hyperparameters in the initial hyperparameter group selected in the above second step, use various machine learning algorithms to train the initial model selected in the above step S1 to generate a quasi-text classification model corresponding to the selected initial hyperparameter group.

[0072] Optionally, the above execution entity can also generate a quasi-text classification model corresponding to the selected initial hyperparameter group according to the following steps:

[0073] Step 1. Select training samples from the training sample set and perform the following training steps. The above training steps may include:

[0074] S21. Input the sample text of the selected training sample into the selected initial model to generate a text category.

[0075] S22. Determine a difference value according to the generated text category and the sample category corresponding to the input sample text.

[0076] S23. Determine whether the difference value meets the training completion condition.

[0077] Based on the above optional implementation manners, the above difference value and training completion condition can be determined based on the initial training hyperparameters in the selected hyperparameter group. As an example, the above initial training hyperparameters may include a loss function and a training completion condition threshold. The above difference value can be determined according to the loss function. The above training completion condition threshold may include at least one of the following: a training duration threshold, a training times threshold, a difference value threshold, and an accuracy threshold of a validation text set.

[0078] S24. In response to determining that the training completion condition is met, determine the selected initial model as the quasi-text classification model corresponding to the selected hyperparameter group.

[0079] Step 2. In response to determining that the training completion condition is not met, adjust the relevant parameters of the selected initial model, select training samples from the training sample set, use the adjusted initial model as the selected initial model, and continue to perform the above training steps.

[0080] Based on the above optional implementation manners, in response to determining that the training completion condition is not satisfied, the above-mentioned execution entity may adopt various methods to adjust the relevant parameters of the selected initial model, and select training samples from the training sample set, use the adjusted initial model as the selected initial model, and continue to execute the above-mentioned training steps. It should be noted that according to whether the number of selected samples is one or more, the methods for adjusting the relevant parameters may include, but are not limited to, at least one of the following: batch gradient descent (BGD), stochastic gradient descent (SGD), and mini-batch gradient descent (MBGD).

[0081] S3. Evaluate the generated quasi-text classification model based on the validation text set to generate an evaluation result.

[0082] Based on the above optional implementation manners, the above-mentioned execution entity may use the validation text set to evaluate the generated quasi-text classification model, and then generate an evaluation result. Among them, the above-mentioned validation text set may be preset. The above-mentioned validation text may include the text and the validation text annotation information corresponding to the text. Optionally, the above-mentioned validation text set may also be generated based on the division of the annotation text set sent by the client received.

[0083] S4. In response to determining that the generated evaluation result meets the hyperparameter group determination condition, determine a text classification model from the quasi-text classification models corresponding to the evaluation results that meet the hyperparameter group determination condition.

[0084] Based on the above optional implementation manners, in response to determining that the generated evaluation result meets the hyperparameter group determination condition, the above-mentioned execution entity may determine a text classification model from the quasi-text classification models corresponding to the evaluation results that meet the hyperparameter group determination condition according to the manner of selecting the initial hyperparameter group from the initial hyperparameter group set. Among them, the above-mentioned hyperparameter group determination condition may include, but is not limited to, at least one of the following: the accuracy rate of the validation text set exceeds a preset accuracy rate threshold, the number of iterations of the initial hyperparameter group exceeds a preset number threshold, and the difference between the accuracy rates of the validation text sets corresponding to the initial hyperparameter groups in two adjacent iterations is less than a preset difference threshold.

[0085] As an example, the above-mentioned execution entity can select one initial hyperparameter group from the initial hyperparameter group set each time. Then, the above-mentioned execution entity can determine the quasi-text classification model corresponding to the evaluation result that meets the hyperparameter group determination condition as the text classification model. As another example, the above-mentioned execution entity can select multiple initial hyperparameter groups from the initial hyperparameter group set each time. Then, there can also be multiple quasi-text classification models corresponding to the evaluation results that meet the hyperparameter group determination condition. The above-mentioned execution entity usually can select the quasi-text classification model with the best evaluation result as the text classification model. For example, the best evaluation result mentioned above can be the highest accuracy rate on the validation text set.

[0086] In the third step, in response to determining that the generated evaluation result does not meet the hyperparameter group determination condition, update the hyperparameter groups in the initial hyperparameter group set; and select initial hyperparameter groups from the updated initial hyperparameter group set, and continue to execute the above-mentioned determination step.

[0087] In these implementation manners, in response to determining that the generated evaluation result does not meet the hyperparameter group determination condition, the above-mentioned execution entity can update the hyperparameter groups in the initial hyperparameter group set in various ways; and the above-mentioned execution entity can also select initial hyperparameter groups from the updated initial hyperparameter group set and continue to execute the above-mentioned determination step.

[0088] It should be noted that according to whether the number of selected initial hyperparameter groups is one or more, the above-mentioned update method can include but is not limited to at least one of the following: genetic algorithm (GA), simulated annealing algorithm (SA), ant colony algorithm (ACA), Bayesian Optimization.

[0089] In some optional implementation manners of this embodiment, the above-mentioned execution entity can also send information indicating that the training of the text classification model is completed to the target terminal.

[0090] Continue to refer to Figure 3 , Figure 3 is a schematic diagram of an application scenario of a method for generating a classification model according to an embodiment of the present disclosure. In Figure 3In the application scenario, user 301 uses the terminal device 302 to upload the annotated investment text set 303 to the background server 304. The background server 304 can determine that the average text length of the annotated investment text set 303 is 3000 words. Then, the background server 304 can generate a corresponding hyperparameter set 3031 according to the determined average text length. Next, the background server 304 can select a pre-trained model trained based on the financial domain dataset from the preset pre-trained model set 3032 as the initial model 3033, and train the above initial model 3033 according to the generated hyperparameter set 3031 to generate a text classification model 305. Optionally, the background server 304 can also send a prompt message 306 indicating that the model training is completed to the terminal device 302.

[0091] Currently, one of the existing technologies usually obtains feature word vectors through complex feature engineering techniques or trains an initial model according to preset hyperparameters. Since the design of feature engineering and the selection of hyperparameters often require users to have rich modeling experience, the technical threshold for training text classification models is relatively high. However, the method provided in the above embodiment of the present disclosure determines the statistical features of the training sample set, selects a pre-trained model according to the statistical features, and trains it in a manner matching the statistical features. It realizes the generation of a text classification model through the fine-tuning of the pre-trained model and a small-scale sample. Moreover, no manual parameter tuning is required during the training process, greatly reducing the user's usage threshold. Furthermore, it enables users to quickly implement text classification applications suitable for their own needs with low-cost investment.

[0092] Further referring to Figure 4 which shows a flowchart 400 of another embodiment of the method for generating a classification model. The flowchart 400 of the method for generating a classification model includes the following steps:

[0093] Step 401, obtaining a training sample set.

[0094] Step 402, determining the statistical features of the training sample set.

[0095] Step 403, generating a text classification model based on the statistical features and the training of the initial model.

[0096] The above steps 401, 402, and 403 are respectively the same as steps 201, 202, and 203 in the foregoing embodiment, and the descriptions of steps 201, 202, and 203 above also apply to steps 401, 402, and 403, and will not be repeated here.

[0097] Step 404, receiving the text set to be classified sent by the user terminal.

[0098] In this embodiment, the execution subject of the method for generating a classification model (such as Figure 1 the server 105 shown) can receive the text set to be classified sent by the user terminal (such as Figure 1 the terminal device shown) through a wired connection or a wireless connection.

[0099] It should be noted that the above step 404 and step 401 can be executed substantially in parallel, or step 404 can be executed first and then step 401.

[0100] Step 405: Input the text set to be classified into the text classification model to generate the category information corresponding to the text to be classified in the text set to be classified.

[0101] In this embodiment, the above execution subject can input the text set to be classified received in step 405 into the text classification model generated in step 403 to generate the category information corresponding to the text to be classified in the text set to be classified. Among them, the above category information can be used to represent the category to which the text to be classified belongs. It can be understood that since the above text classification model is trained based on the training sample set obtained in the above step 401, the generated category information corresponding to the text to be classified can match the sample category of the above training sample.

[0102] Step 406: Send the generated category information and the corresponding text information to be classified to the user terminal.

[0103] In this embodiment, the above execution subject can send the generated category information and the corresponding text information to be classified to the user terminal through a wired connection or a wireless connection. Among them, the above text information to be classified can be used to identify the text to be classified in the above text set to be classified.

[0104] From Figure 4 it can be seen that the process 400 of the method for generating a classification model in this embodiment reflects the steps of classifying the text to be classified uploaded by the user by using the generated text classification model. Thus, the solution described in this embodiment can train a text classification model with the labeled text uploaded by the user, and then use the trained text classification model to classify the unlabeled text uploaded, so as to realize the targeted training and application of the required text classification model by using the existing samples without manual parameter adjustment. Furthermore, the training and use thresholds of the text classification model are reduced.

[0105] Further referring to Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an apparatus for generating a classification model. This apparatus embodiment is related to Figure 2The method embodiments shown correspond to this, and the device can be specifically applied to various electronic devices.

[0106] As Figure 5 As shown, the device 500 for generating a classification model provided in this embodiment includes an acquisition unit 501, a determination unit 502, and a first generation unit 503. Among them, the acquisition unit 501 is configured to acquire a training sample set, where the training samples include sample texts and corresponding sample categories; the determination unit 502 is configured to determine the statistical features of the training sample set, where the statistical features include formal features representing the text length; the first generation unit 503 is configured to generate a text classification model based on the statistical features and the training of an initial model, where the text classification model is used to represent the correspondence between text categories and texts, and the initial model is selected from a preset set of pre-trained models.

[0107] In this embodiment, in the device 500 for generating a classification model: the specific processing of the acquisition unit 501, the determination unit 502, and the first generation unit 503 and the technical effects brought by them can respectively refer to Figure 2 the relevant descriptions of steps 201, step 202, and step 203 in the corresponding embodiments, which will not be elaborated here.

[0108] In some optional implementation manners of this embodiment, the above first generation unit 503 may include a first generation module (not shown in the figure), a selection module (not shown in the figure), and a second generation module (not shown in the figure). Among them, the above first generation module can be configured to input the formal features into a pre-trained hyperparameter generation model to obtain a hyperparameter group corresponding to the formal features. Among them, the above hyperparameter group may include model hyperparameters and training hyperparameters. The above selection module can be configured to select a pre-trained model that matches the model hyperparameters in the hyperparameter group from the set of pre-trained models as the initial model. The above second generation module can be configured to train the initial model according to the training hyperparameters in the hyperparameter group to generate a text classification model.

[0109] In some alternative implementation manners of this embodiment, the foregoing first generation unit 503 may include an obtaining module (not shown in the figure), a determining module (not shown in the figure), and an updating module (not shown in the figure). Among them, the foregoing obtaining module may be configured to obtain a set of initial hyperparameter groups. Among them, the foregoing initial hyperparameter group may include initial model hyperparameters and initial training hyperparameters. The foregoing determining module may be configured to select an initial hyperparameter group from the set of initial hyperparameter groups and perform the following determining steps: select a pre-trained model that matches the initial model hyperparameters in the selected initial hyperparameter group from the set of pre-trained models as the initial model; train the selected initial model according to the initial training hyperparameters in the selected initial hyperparameter group to generate a quasi-text classification model corresponding to the initial hyperparameter group; evaluate the generated quasi-text classification model based on the validation text set to generate an evaluation result; in response to determining that the generated evaluation result meets the hyperparameter group determination condition, determine a text classification model from the quasi-text classification models corresponding to the evaluation results that meet the hyperparameter group determination condition. The foregoing updating module may be configured to update the initial hyperparameter groups in the set of initial hyperparameter groups in response to determining that the generated evaluation result does not meet the hyperparameter group determination condition; select an initial hyperparameter group from the updated set of initial hyperparameter groups and continue to perform the determining steps.

[0110] In some alternative implementation manners of this embodiment, the foregoing statistical feature may further include a content feature characterizing the text content. The foregoing determining module may be further configured to: select a pre-trained model that matches the content feature and the initial model hyperparameters in the selected initial hyperparameter group from the set of pre-trained models as the initial model. Among them, the foregoing pre-trained model may correspond to a semantic label.

[0111] In some alternative implementation manners of this embodiment, the foregoing determining module may further include: a selection sub-module (not shown in the figure), an adjustment sub-module (not shown in the figure). Among them, the foregoing selection sub-module may be configured to select a training sample from the set of training samples and perform the following training steps: input the sample text of the selected training sample into the selected initial model to generate a text category; determine a difference value according to the generated text category and the sample category corresponding to the input sample text; determine whether the difference value meets the training completion condition, where the difference value and the training completion condition are determined based on the initial training hyperparameters in the selected hyperparameter group; in response to determining that the training completion condition is met, determine the selected initial model as the quasi-text classification model corresponding to the selected hyperparameter group. The foregoing adjustment sub-module may be configured to, in response to determining that the training completion condition is not met, adjust the relevant parameters of the selected initial model, select a training sample from the set of training samples, use the adjusted initial model as the selected initial model, and continue to perform the training steps.

[0112] In some alternative implementation manners of this embodiment, the above-mentioned obtaining unit 501 may include: a receiving module (not shown in the figure), and a third generating module (not shown in the figure). Among them, the above-mentioned receiving module may be configured to receive the labeled text set sent by the client. Among them, the above-mentioned labeled text may include text and text category labeling information corresponding to the text. The above-mentioned generating module may be configured to divide the labeled text set to generate a training sample set and a verification text set. Among them, the above-mentioned training sample may include text as a sample text and text category labeling information as a sample category corresponding to the sample text.

[0113] In some alternative implementation manners of this embodiment, the above-mentioned apparatus 500 for generating a classification model may include: a receiving unit (not shown in the figure), a second generating unit (not shown in the figure), and a sending unit (not shown in the figure). Among them, the above-mentioned receiving unit may be configured to receive the text set to be classified sent by the client. The above-mentioned second generating unit may be configured to input the text set to be classified into the text classification model to generate category information corresponding to the text to be classified in the text set to be classified. Among them, the above-mentioned category information may be used to represent the category to which the text to be classified belongs. The above-mentioned category information may match the sample category. The above-mentioned sending unit may be configured to send the generated category information and the corresponding text information to be classified to the client. Among them, the above-mentioned text information to be classified may be used to identify the text to be classified in the text set to be classified.

[0114] The apparatus provided in the above embodiment of the present disclosure obtains a training sample set through the obtaining unit 501. Among them, the training sample includes a sample text and a sample category corresponding to the sample text. Then, the determining unit 502 determines the statistical features of the training sample set. Among them, the statistical features include formal features characterizing the text length. After that, the first generating unit 503 generates a text classification model based on the statistical features and the training of the initial model. Among them, the text classification model is used to represent the corresponding relationship between the text category and the text. The initial model is selected from a preset pre-training model set. Thus, the automatic generation of the text classification model can be realized without manual parameter adjustment.

[0115] Next, refer to Figure 6 , which shows a schematic structural diagram of an electronic device (such as Figure 1 the server in Figure 6 shown) 600 suitable for implementing the embodiments of the present disclosure.

[0116] As Figure 6As shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0117] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 the electronic device 600 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had. Figure 6 Each block shown in may represent one device or, as needed, multiple devices.

[0118] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the method of the embodiment of the present disclosure are executed.

[0119] It should be noted that the computer-readable medium described in the embodiments of the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the embodiments of the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0120] The above computer-readable medium can be included in the above server; or it can exist separately without being assembled into the server. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the server, the server is caused to: obtain a training sample set, where the training sample includes a sample text and a sample category corresponding to the sample text; determine the statistical features of the training sample set, where the statistical features include formal features characterizing the text length; generate a text classification model based on the statistical features and the training of an initial model, where the text classification model is used to represent the correspondence between the text category and the text, and the initial model is selected from a preset set of pre-trained models.

[0121] Computer program code for performing the operations of the embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., connected through the Internet using an Internet service provider).

[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in an order different from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0123] The units involved in the embodiments described in the present disclosure may be implemented in software or in hardware. The described units may also be provided in a processor. For example, it may be described as a processor including an acquisition unit, a determination unit, and a first generation unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the acquisition unit may also be described as "the unit for acquiring a training sample set, where the training sample includes a sample text and a sample category corresponding to the sample text".

[0124] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the embodiments of the present disclosure that have similar functions.

Claims

1. A method for generating a classification model, comprising: Obtaining a training sample set, where the training samples include sample texts and sample categories corresponding to the sample texts; Determining statistical features of the training sample set, where the statistical features include formal features characterizing the text length; Obtaining an initial hyperparameter group set including a plurality of initial hyperparameter groups that match the statistical features; selecting an initial hyperparameter group from the initial hyperparameter group set, and performing the following determination steps on the selected initial hyperparameter group: selecting a pre-trained model that matches the initial model hyperparameters in the initial hyperparameter group from a preset pre-trained model set as an initial model; training the initial model according to the initial training hyperparameters in the initial hyperparameter group to generate a quasi-text classification model; evaluating the quasi-text classification model based on a validation text set to generate an evaluation result; determining a text classification model for characterizing the correspondence between text categories and texts from the quasi-text classification models corresponding to the evaluation results that meet the hyperparameter group determination conditions; in response to the generated evaluation result not meeting the hyperparameter group determination conditions, updating the initial hyperparameter groups in the initial hyperparameter group set; selecting an initial hyperparameter group from the updated initial hyperparameter group set, and continuing to perform the determination steps.

2. The method according to claim 1, wherein, The statistical features further include content features characterizing the text content; And The step of selecting a pre-trained model that matches the initial model hyperparameters in the initial hyperparameter group from a preset pre-trained model set as an initial model includes: Selecting a pre-trained model that matches the content features and the initial model hyperparameters in the initial hyperparameter group from the pre-trained model set as an initial model, where the pre-trained model corresponds to a semantic label.

3. The method according to claim 1, wherein The step of training the initial model according to the initial training hyperparameters in the initial hyperparameter group to generate a quasi-text classification model includes: Selecting training samples from the training sample set, and performing the following training steps: inputting the sample text of the selected training sample into the initial model to generate a text category; determining a difference value according to the generated text category and the sample category corresponding to the input sample text; determining whether the difference value meets the training completion condition, where the difference value and the training completion condition are determined based on the initial training hyperparameters in the selected hyperparameter group; in response to determining that the training completion condition is met, determining the initial model as the quasi-text classification model; In response to determining that the training completion condition is not met, adjusting the relevant parameters of the initial model, and selecting training samples from the training sample set, using the adjusted initial model as the initial model, and continuing to perform the training steps.

4. The method according to claim 1, wherein The step of obtaining the training sample set includes: Receiving an annotated text set sent by a client, where the annotated text includes a text and text category annotation information corresponding to the text; Dividing the annotated text set to generate the training sample set and the validation text set, where the training samples include the text as the sample text and the text category annotation information as the sample category corresponding to the sample text.

5. The method according to any one of claims 1-4, wherein, The method further includes: Receiving a set of texts to be classified sent by a client; Inputting the set of texts to be classified into the text classification model to generate category information corresponding to the texts to be classified in the set of texts to be classified, where the category information is used to represent the category to which the text to be classified belongs, and the category information matches the sample category; Sending the generated category information and the corresponding text information to be classified to the client, where the text information to be classified is used to identify the texts to be classified in the set of texts to be classified.

6. An apparatus for generating a classification model, including: An acquisition unit configured to acquire a set of training samples, where the training samples include sample texts and sample categories corresponding to the sample texts; A determination unit configured to determine statistical features of the set of training samples, where the statistical features include formal features representing the text length; An acquisition module configured to acquire an initial hyperparameter set including a plurality of initial hyperparameter groups that match the statistical features; A determination module configured to select an initial hyperparameter group from the initial hyperparameter set and perform the following determination steps on the selected initial hyperparameter group: selecting a pre-trained model that matches the initial model hyperparameters in the initial hyperparameter group from a preset set of pre-trained models as an initial model; training the initial model according to the initial training hyperparameters in the initial hyperparameter group to generate a quasi-text classification model; evaluating the quasi-text classification model based on a set of validation texts to generate an evaluation result; determining a text classification model for representing the correspondence between text categories and texts from the quasi-text classification models corresponding to the evaluation results that meet the hyperparameter group determination conditions; An update module configured to, in response to the generated evaluation result not meeting the hyperparameter group determination conditions, update the initial hyperparameter groups in the initial hyperparameter set; selecting an initial hyperparameter group from the updated initial hyperparameter set and continuing to perform the determination steps.

7. The apparatus according to claim 6, wherein The statistical features further include content features representing the text content; the determination module is further configured to: Select a pre-trained model that matches the content features and the initial model hyperparameters in the selected initial hyperparameter group from a preset set of pre-trained models as an initial model, where the pre-trained model corresponds to a semantic label.

8. The apparatus according to claim 6, wherein The determination module further includes: A selection sub-module configured to select training samples from the set of training samples and perform the following training steps: inputting the sample texts of the selected training samples into the selected initial model to generate text categories; determining a difference value according to the generated text categories and the sample categories corresponding to the input sample texts; determining whether the difference value meets the training completion condition, where the difference value and the training completion condition are determined based on the initial training hyperparameters in the selected hyperparameter group; in response to determining that the training completion condition is met, determining the selected initial model as the quasi-text classification model corresponding to the selected hyperparameter group; An adjustment sub-module, configured to adjust relevant parameters of the selected initial model in response to determining that the training completion condition is not met, and select training samples from the training sample set, use the adjusted initial model as the selected initial model, and continue to execute the training step.

9. The apparatus according to claim 6, wherein, The obtaining unit includes: A receiving module, configured to receive an annotated text set sent by a client, where the annotated text includes text and text category annotation information corresponding to the text; A third generating module, configured to divide the annotated text set to generate the training sample set and the verification text set, where the training sample includes text as a sample text and text category annotation information as a sample category corresponding to the sample text.

10. The device according to any one of claims 6-9, wherein, The device further includes: A receiving unit, configured to receive a text set to be classified sent by a client; A second generating unit, configured to input the text set to be classified into the text classification model to generate category information corresponding to the text to be classified in the text set to be classified, where the category information is used to represent the category to which the text to be classified belongs, and the category information matches the sample category; A sending unit, configured to send the generated category information and the corresponding text information to be classified to the client, where the text information to be classified is used to identify the text to be classified in the text set to be classified.

11. A server, including: One or more processors; A storage device, on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-5.

12. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by a processor, it implements the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Method and system for performing machine learning process for text classification

    CN108875045A