Training Method, Device, Electronic Device, and Storage Medium for Text Classification Model

By constructing a text representation network with similar semantics and different text representation networks, the problem that text classification models in the prior art cannot fully train samples with different semantics is solved, and the accuracy of the model is improved.

CN114637851BActive Publication Date: 2025-08-05BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210288560.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-23
Publication Date
2025-08-05
Estimated Expiration
2042-03-23

AI Technical Summary

Technical Problem

In the prior art, text classification models only consider feature representations with similar semantics during training, and cannot fully train samples with different semantics, resulting in poor model accuracy.

Method used

The first text representation network with similar semantics and the second text representation network, and the third text representation network with different semantics are constructed, and by training these three networks, the semantics and different samples are fully considered.

Benefits of technology

The learning effect of text classification model is improved, so that it can better handle samples with different semantics, and improve the accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114637851B_ABST
    Figure CN114637851B_ABST
Patent Text Reader

Abstract

This application discloses a training method, apparatus, electronic device, and storage medium for a text classification model, belonging to the field of model training technology. The text classification model training method includes: constructing a first text representation network, a second text representation network, and a third text representation network, wherein the first text representation network and the second text representation network are semantically similar representation networks, and the second text representation network and the third text representation network are semantically different representation networks; and inputting a text dataset into the first, second, and third text representation networks to train the text classification model, thereby obtaining a trained text classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of model training technology, and specifically relates to a training method, device, electronic device and storage medium for a text classification model. Background Art

[0002] Text classification is one of the most basic and important tasks in natural language processing. After training the model through machine learning or deep learning methods, the trained model can be used to label text units such as phrases, sentences, paragraphs, and even articles.

[0003] In the existing technology, semantically similar feature representations are constructed during the model training process. Only semantically similar samples constructed by the feature representation network with the same thrown value are considered, and samples with different semantics cannot be trained. As a result, the semantically different samples cannot be fully trained and learned during the learning process of the text classification model, resulting in poor accuracy of the trained text classification model. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a text classification model training method, a text classification model training device, an electronic device and a readable storage medium, which ensure that semantically different samples can be fully trained and learned during the text classification model learning process, so that the trained text classification model has higher accuracy.

[0005] In a first aspect, an embodiment of the present application provides a training method for a text classification model, comprising: constructing a first text representation network, a second text representation network, and a third text representation network, wherein the first text representation network and the second text representation network are representation networks with similar semantics, and the second text representation network and the third text representation network are representation networks with different semantics; inputting a text data set into the first text representation network, the second text representation network, and the third text representation network to train the text classification model to obtain a trained text classification model.

[0006] In the second aspect, an embodiment of the present application provides a training device for a text classification model, including: a construction module for constructing a first text representation network, a second text representation network and a third text representation network, the first text representation network and the second text representation network are representation networks with similar semantics, and the second text representation network and the third text representation network are representation networks with different semantics; a training module for inputting a text data set into the first text representation network, the second text representation network and the third text representation network to train the text classification model to obtain a trained text classification model.

[0007] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method of the first aspect.

[0008] In a fourth aspect, an embodiment of the present application provides a readable storage medium, which stores a program or instruction. When the program or instruction is executed by a processor, the steps of the training method of the text classification model in the first aspect are implemented.

[0009] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the training method of the text classification model as in the first aspect.

[0010] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the training method of the text classification model as in the first aspect.

[0011] In an embodiment of the present application, before training the model, a first text representation network and a second text representation network with similar semantics are constructed, and a third text representation network with different semantics is constructed. By training the first text representation network, the second text representation network and the third text representation, samples with similar semantics and samples with different semantics are fully taken into consideration, ensuring that samples with different semantics can be fully trained and learned during the learning process of the text classification model, so that the accuracy of the trained text classification model is higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 One of the flow charts of the training method of the text classification model provided in the embodiment of the present application is shown;

[0013] Figure 2 The second flowchart of the text classification model training method provided in the embodiment of the present application is shown;

[0014] Figure 3 The third flowchart of the training method of the text classification model provided in the embodiment of the present application is shown;

[0015] Figure 4 FIG4 shows a fourth flow chart of the training method of the text classification model provided in an embodiment of the present application;

[0016] Figure 5 FIG5 shows a fifth flow chart of the training method of the text classification model provided in an embodiment of the present application;

[0017] Figure 6One of the schematic diagrams showing the structure of a text classification model after training according to the text classification model training method provided in an embodiment of the present application;

[0018] Figure 7 A second schematic diagram showing the structure of a text classification model after training according to the text classification model training method provided in an embodiment of the present application;

[0019] Figure 8 A third schematic diagram showing the structure of a text classification model after training according to the text classification model training method provided in an embodiment of the present application;

[0020] Figure 9 A structural block diagram of a text classification model training device provided in an embodiment of the present application is shown;

[0021] Figure 10 A structural block diagram of an electronic device provided in an embodiment of the present application is shown;

[0022] Figure 11 A schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0023] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0024] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0025] The following is combined with Figures 1 to 11 , through specific embodiments and their application scenarios, the training method of the text classification model, the training device of the text classification model, the electronic device and the readable storage medium provided in the embodiments of the present application are described in detail.

[0026] In an embodiment of the present application, a text classification model training method is provided for a virtual reality device. Figure 1FIG. 1 shows one of the flow charts of the training method of the text classification model provided in the embodiment of the present application, such as Figure 1 As shown in Figure 2, the training method for the text classification model includes:

[0027] Step 102: construct a first text representation network, a second text representation network, and a third text representation network;

[0028] The first text representation network and the second text representation network are representation networks with similar semantics, and the second text representation network and the third text representation network are representation networks with different semantics.

[0029] Step 104 : Input the text dataset into the first text representation network, the second text representation network, and the third text representation network to train the text classification model, so as to obtain a trained text classification model.

[0030] In an embodiment of the present application, before starting to train the model, a text feature representation network is constructed. In the process of constructing the text feature representation network, the present application constructs three representation networks: a first text representation network, a second text representation network, and a third text representation network. Among them, the first text representation network and the second text representation network are used as network representation structures with similar semantics, and the second text representation network and the third text representation network, as well as the first text representation network and the third text representation network, are used as network representation structures with different semantics. After constructing the text feature representation network, a text data set is obtained. In the training process of the model, the collected text data set is input into the first text representation network, the second text representation network, and the third text representation network to train the model. In the training process, the first text representation network, the second text representation network, and the third text representation network respectively output semantic vectors. The semantic vectors output by the first text representation network and the second text representation network are similar semantic vectors, and the semantic vector output by the third text representation network is different from the semantic vector output by the second text representation network. The model with the highest accuracy in the training process is determined based on the similar semantic vectors and the different semantic vectors, and the model with the highest progress is used as the text classification model after training.

[0031] Specifically, the first text representation network, the second text representation network, and the third text representation network can be constructed as an LSTM feature representation network and a Bert feature representation network. The original network structures of the first text representation network, the second text representation network, and the third text representation network are different, wherein the first text representation network has the same dropout value as the second text representation network and a different dropout value from the third text representation network, and the dropout value of the first text representation network is set to be smaller than the dropout value of the second text representation network, so that the constructed first text representation network and the second text representation network have similar semantic network representation structures, and the first text representation network and the third text representation network have different semantic network representation structures.

[0032] In related technologies, feature representation networks with identical throw values are constructed before model training to create semantically similar feature representations. During text classification model training, only semantically similar samples constructed using feature representation networks with identical throw values are considered, and training on semantically distinct samples is not performed. Consequently, the text classification model cannot fully train and learn semantically distinct samples.

[0033] In an embodiment of the present application, before training the model, a first text representation network and a second text representation network with similar semantics are constructed, and a third text representation network with different semantics is constructed. By training the first text representation network, the second text representation network and the third text representation, samples with similar semantics and samples with different semantics are fully taken into consideration, ensuring that samples with different semantics can be fully trained and learned during the learning process of the text classification model, so that the accuracy of the trained text classification model is higher.

[0034] In some embodiments of the present application, Figure 2 The second flow chart of the training method of the text classification model provided in the embodiment of the present application is shown as follows: Figure 2 As shown, constructing a first text representation network, a second text representation network, and a third text representation network includes:

[0035] Step 202, obtaining a first thrown value and a second thrown value;

[0036] Step 204: construct a first text representation network and a second text representation network according to the first thrown value;

[0037] Step 206: construct a third text representation network based on the second thrown value;

[0038] The first thrown value is smaller than the second thrown value.

[0039] In the embodiment of the present application, in the process of constructing the first text representation network, the second text representation network, and the third text representation network, it is necessary to first construct corresponding initial representation networks, and the initial representation networks corresponding to the above three text representation networks are identical representation networks. After constructing the three identical initial representation networks, the initial representation networks are subjected to a throw process according to the first throw value to obtain the first text representation network, the initial representation network is subjected to a throw process according to the second throw value to obtain the second text representation network, and the initial representation network is subjected to a throw process according to the third throw value.

[0040] It is worth noting that the dropout value is the dropout value during the dropout process of the text representation network. The dropout value can reflect the number of neurons randomly dropped in the text representation network. During the dropout process of the text representation network, the smaller the dropout value, the fewer neurons are randomly reduced, and the larger the dropout value, the more neurons are randomly reduced. Therefore, the first text representation network and the second text representation network obtained by the same dropout value processing have the same number of reduced neurons, but the reduced neurons are not exactly the same. Therefore, the first and second text representation networks obtained are semantically similar representation networks, and the first and third text representation networks are semantically different representation networks.

[0041] In this embodiment of the present application, a first text representation network and a second text representation network are constructed using a first thrown value, thereby ensuring semantic similarity between the first and second text representation networks. Furthermore, a third text representation network is constructed using a second thrown value that is greater than the first thrown value, thereby ensuring semantic differences between the third text representation network and the first text representation network.

[0042] In some embodiments of the present application, the value range of the first throw value is greater than 0.00001 and less than 0.5; the value range of the second throw value is greater than 0.50001 and less than 1.

[0043] In the embodiment of the present application, the range of the first throw value is set to 0.00001 to 0.5, which can ensure that the first text representation network and the second text representation network constructed based on the first throw value are semantically similar network representation structures. The range of the second throw value is set to 0.50001 to 1, which can ensure that the third text representation network structure is a network representation structure with semantically different semantics from the first text representation network and the second text representation network.

[0044] In some possible implementations, the first throw value is selected as 0.2, and the second throw value is selected as 0.8.

[0045] In some embodiments of the present application, before inputting the text dataset into the first text representation network, the second text representation network and the third text representation network to train the text classification model to obtain the trained text classification model, it also includes: obtaining the text dataset.

[0046] The specific method for obtaining a text data set includes: obtaining text sample data; and performing word segmentation processing on the text sample data according to preset rules to obtain a text data set.

[0047] In the embodiment of the present application, the text sample data is a text with complete semantics, including semantically complete terms, semantically complete sentences, and related articles, etc. After obtaining the text sample data, the text sample data is segmented according to a preset trajectory to obtain a text dataset.

[0048] Specifically, word segmentation can be performed on text sample data according to specific semantics, or according to characters. For example, if the text sample data is "Today's weather is cloudy and rainy", the text dataset obtained by semantic segmentation includes "today", "of", "weather", "for", and "cloudy and rainy day", while the text dataset obtained by character segmentation includes "today", "day", "of", "day", "air", "for", "cloudy", "rain", and "day".

[0049] It is worth noting that different text sample data are selected according to the specific application scenario of the text classification model. For example, if the text classification model is a classification model applied to the news field, news articles are used as text sample data.

[0050] In the embodiment of the present application, before training the text classification model, the text sample data is segmented according to a preset trajectory to obtain a corresponding text dataset. The text dataset can be configured according to different usage requirements to ensure the accuracy of the trained text classification model.

[0051] In some embodiments of the present application, Figure 3 The third flow chart of the training method of the text classification model provided in the embodiment of the present application is shown as follows: Figure 3 As shown, the text dataset is input into the first text representation network, the second text representation network, and the third text representation network for training to obtain a trained text classification model, including:

[0052] Step 302: Obtain a preset number of training times;

[0053] The preset number of training times is the number of times set in advance, that is, the number of cycles for training the text classification model.

[0054] Step 304: training multiple text classification models based on the text dataset according to a preset number of training times;

[0055] Step 306: Obtain a loss function corresponding to each text classification model in the multiple text classification models;

[0056] It's worth noting that the text classification model is trained for a preset number of times, resulting in a new text classification model each time. The new text classification model and its corresponding loss function are recorded. After training is complete, the most accurate trained text classification model is found based on the loss function corresponding to each text classification model.

[0057] Step 308: Determine a trained text classification model from the multiple text classification models based on the multiple loss functions.

[0058] In an embodiment of the present application, before starting model training, a preset number of training times for the model is determined. After the preset number of training times is determined, the text classification model is cyclically trained according to the preset number of training times. The results of each cyclic training are recorded, and the new text classification model obtained through training and its loss function are recorded. After the preset number of training times is reached, the most accurate trained text classification model can be selected based on the recorded text classification models and loss functions.

[0059] Specifically, during the training process, not only is the text classification model obtained during each training cycle recorded, but the corresponding loss function for each text classification model is also recorded. By recording the loss function, the convergence curve of the text classification model is obtained. The text classification model with the highest accuracy can be found in the convergence curve and used as the trained text classification model.

[0060] In an embodiment of the present application, during the cyclic training of the text classification model, the text classification model after each training is recorded, and the trained text classification model with the highest accuracy is searched among them, ensuring that the trained text classification model finally obtained is the model with the highest accuracy.

[0061] In some embodiments of the present application, Figure 4 The fourth flow chart of the training method of the text classification model provided in the embodiment of the present application is shown as follows: Figure 4 As shown, obtaining the loss function corresponding to each text classification model in multiple text classification models also includes:

[0062] Step 402: Obtain a first semantic vector, a second semantic vector, and a third semantic vector of each text classification model;

[0063] Among them, the first semantic vector corresponds to the first text representation network, the second semantic vector corresponds to the second text representation network, and the second semantic vector corresponds to the first text representation network.

[0064] Step 404: Determine a model loss function of the text classification model based on the first semantic vector, the second semantic vector, and the third semantic vector.

[0065] In an embodiment of the present application, the first text representation network, the second text representation network, and the third text representation network in each trained text classification model will output a semantic vector after inputting a text data set. Among them, the first text representation network outputs a first semantic vector, the second text representation network outputs a second semantic vector, and the third text representation network outputs a third semantic vector. The loss function of the corresponding text classification model can be calculated through the first semantic vector and the second semantic vector. After recording multiple trained text classification models, a convergence curve can be constructed according to the loss function corresponding to each trained text classification model, and the trained text classification model with the highest accuracy can be found according to the convergence curve to determine the trained text classification model.

[0066] It can be understood that the first semantic vector and the second semantic vector are semantic vectors with similar semantics, and the third semantic vector is a semantic vector with different semantics from the first semantic vector and the second semantic vector.

[0067] In an embodiment of the present application, a loss function is calculated based on the first semantic vector, the second semantic vector, and the third semantic vector, and a trained text classification model among multiple text classification models in the training process is screened based on the calculated loss function.

[0068] In some embodiments of the present application, Figure 5 The fifth flow chart of the training method of the text classification model provided in the embodiment of the present application is shown as follows: Figure 5 As shown, according to the first semantic vector, the second semantic vector and the third semantic vector, the model loss function of the text classification model is determined, including:

[0069] Step 502: determining a first loss function based on the first semantic vector and the preset vector;

[0070] Step 504: determining a second loss function based on the first semantic vector, the second semantic vector, and the third semantic vector;

[0071] Step 506: Determine a model loss function based on the first loss function and the second loss function.

[0072] In an embodiment of the present application, a first loss function can be determined based on the first semantic vector output by the first text representation network and a preset vector, where the preset vector is a vector of a label sample preset in advance by the user. Specifically, the first text representation network performs supervised training, i.e., a cross-entropy loss function is constructed for the first text representation network and the label sample to perform supervised learning, and the position where the probability of the first semantic vector output by the first text representation network is the highest is controlled to be the corresponding position of the label sample.

[0073] A second loss function is determined based on the first semantic vector, the second semantic vector, and the third semantic vector. Specifically, a second loss function based on distance sorting is constructed for the first text representation network and the second text representation network with similar semantics, and the second text representation network and the third text representation network with different semantics. During the model training process, the first semantic vector and the second semantic vector outputted by the first text representation network and the second text representation network are controlled to be close to each other, and the second semantic vector and the third semantic vector outputted by the second text representation network and the third text representation network are controlled to be far away from each other.

[0074] By training the text model, the first semantic vector output by the first text representation network of supervised learning is close to the label sample, the second semantic vector output by the second text representation network of unsupervised learning is close to the first semantic vector output by the first text representation network, and the third semantic vector output by the third text representation network of partially supervised learning is far away from the third semantic vector output by the second text representation network.

[0075] In an embodiment of the present application, the text classification model is jointly trained through the first loss function and the second loss function, so that the text classification model can be more fully trained, the representation distance of samples with similar semantics is closer, and the representation distance of samples with different semantics is farther, and the model representation effect is better.

[0076] In some embodiments of the present application, a second loss function is determined based on the first semantic vector, the second semantic vector, and the third semantic vector, including: calculating the first cosine distance between the first semantic vector and the second semantic vector; calculating the second cosine distance between the second semantic vector and the third semantic vector; and determining the second loss function based on the first cosine distance and the second cosine distance.

[0077] In the embodiment of the present application, in the process of determining the second loss function, it is necessary to calculate a first cosine distance between the first semantic vector and the second semantic vector, and a second pre-distance between the second semantic vector and the third semantic vector. The second loss function can be determined by the first cosine distance and the second cosine distance.

[0078] The second loss function is selected as the circle loss function. The circle loss function is used to consolidate the comparative learning of samples with different semantic representations, even if the second semantic vector is further away from the third semantic vector, thereby improving the representation effect of the trained text classification model.

[0079] Figure 6 One of the schematic diagrams showing the structure of the text classification model after training according to the training method of the text classification model provided in the embodiment of the present application is shown as follows: Figure 6 As shown, during the training process, the text dataset is input into the pre-trained text classification model and then the text dataset is embedded and mapped, and then the processed text dataset is input into the first text representation network, the second text representation network and the third text representation network respectively. The first text representation network of supervised learning and the label (label sample) are used to construct a cross entropy loss function for supervised learning, and the first semantic vector output by the first text representation network is controlled to correspond to the label sample. For the first text representation network and the second text representation network with similar semantics, and the second text representation network and the third text representation network with different semantics, a distance-sorting-based loss function (circle loss function) is constructed to control the first semantic vector output by the first text representation network to be close to the second semantic vector output by the second text representation network, and to control the second semantic vector output by the second text representation network to be far away from the third semantic vector output by the third text representation network.

[0080] In some embodiments of the present application, determining the model loss function based on the first loss function and the second loss function includes: determining the third loss function based on the second semantic vector and the preset vector; determining the model loss function based on the first loss function, the second loss function and the third loss function.

[0081] In an embodiment of the present application, a second loss function can be determined based on the second semantic vector output by the second text representation network and a preset vector, where the preset vector is a vector of a label sample preset in advance by the user. Specifically, the second text representation network performs supervised training, i.e., a cross-entropy loss function is constructed for the second text representation network and the label sample for supervised learning, and the position at which the probability of the second semantic vector output by the second text representation network is the highest is controlled to be the corresponding position of the label sample.

[0082] Figure 7 The second schematic diagram of the structure of the text classification model after training of the text classification model training method provided in the embodiment of the present application is shown as follows: Figure 7As shown, during the training process, the text dataset is input into the pre-trained text classification model and then the text dataset is embedded and mapped, and then the processed text dataset is input into the first text representation network, the second text representation network and the third text representation network respectively. The first text representation network of supervised learning and the label (label sample) are used to construct a cross-entropy loss function for supervised learning, and the first semantic vector output by the first text representation network is controlled to correspond to the label sample. The second text representation network of supervised learning and the label (label sample) are used to construct a cross-entropy loss function for supervised learning, and the second semantic vector output by the second text representation network is controlled to correspond to the label sample. For the first text representation network and the second text representation network with similar semantics, and the second text representation network and the third text representation network with different semantics, a distance-sorted pair loss function (circle loss function) is constructed to control the first semantic vector output by the first text representation network to be close to the second semantic vector output by the second text representation network, and to control the second semantic vector output by the second text representation network to be far away from the third semantic vector output by the third text representation network.

[0083] In the embodiment of the present application, not only the first text representation network and the preset label samples are supervised for learning, but also the second text representation network and the preset label samples are supervised for learning, thereby further improving the accuracy of the trained text classification model.

[0084] In some embodiments of the present application, a model loss function is determined based on the first loss function, the second loss function, and the third loss function, including: determining a fourth loss function based on the first semantic vector and the second semantic vector; determining the model loss function based on the first loss function, the second loss function, the third loss function, and the fourth loss function.

[0085] In an embodiment of the present application, a fourth loss function is constructed for the first text representation network and the second text representation network with similar semantics to perform unsupervised learning, so that the first semantic vector output by the first text representation network is close to the second semantic vector output by the second text representation network.

[0086] Specifically, the fourth loss function is the KL divergence (relative entropy divergence) loss function. The KL divergence loss function can be used to control the distance between the first semantic vector output by the first text representation network and the second semantic vector output by the second text representation network to be close, thereby further improving the accuracy of the model.

[0087] Figure 8 The third schematic diagram shows the structure of the text classification model after training according to the training method of the text classification model provided in the embodiment of the present application. Figure 8As shown, during the training process, the text dataset is input into the pre-trained text classification model and then the text dataset is embedded and mapped, and then the processed text dataset is input into the first text representation network, the second text representation network and the third text representation network respectively. The first text representation network of supervised learning is used to construct a cross-entropy loss function with the label (label sample) for supervised learning, and the first semantic vector output by the first text representation network is controlled to correspond to the label sample. The second text representation network of supervised learning is used to construct a cross-entropy loss function with the label (label sample) for supervised learning, and the second semantic vector output by the second text representation network is controlled to correspond to the label sample. For the first text representation network and the second text representation network with similar semantics, a KL divergence (relative entropy divergence) loss function is constructed for unsupervised learning, so that the first semantic vector output by the first text representation network is close to the second semantic vector output by the second text representation network. For the first text representation network and the second text representation network with similar semantics, and the second text representation network and the third text representation network with different semantics, a distance-sorting-based loss function (circle loss function) is constructed to control the first semantic vector output by the first text representation network to be close to the second semantic vector output by the second text representation network, and to control the second semantic vector output by the second text representation network to be far away from the third semantic vector output by the third text representation network.

[0088] The embodiment of the present application sets the first loss function, the second loss function, the third loss function and the fourth loss function in the model, so that the output of the first text representation network and the output of the second text representation network are closer to the sample label, the output of the first text representation network is closer to the output of the second text representation network, and the output of the second text representation network is further away from the output of the third text representation network. Through the joint training of multiple loss functions, the number of required preset label samples can be reduced, and the number of training steps can be reduced.

[0089] Specifically, the first and third loss functions are selected as cross entropy loss functions, the second loss function is selected as circle loss loss function (for loss function), and the third loss function is selected as KL divergence loss function. And weights are set for the cross entropy loss function, circle loss loss function, and KL divergence loss function. That is, the joint training method is as follows:

[0090] a×cross entropy loss function +b×circle loss function +c×KL divergence loss function;

[0091] Among them, a, b, and c are all constants.

[0092] The training method for a text classification model provided in the embodiments of the present application can be executed by a training device for a text classification model. In the embodiments of the present application, the training method for a text classification model performed by a training device for a text classification model is used as an example to illustrate the training device for a text classification model provided in the embodiments of the present application.

[0093] In some embodiments of the present application, a training device for a text classification model is provided. Figure 9 The structural block diagram of the training device of the text classification model provided in the embodiment of the present application is shown as follows: Figure 9 As shown, the training device 900 of the text classification model includes:

[0094] A construction module 902 is configured to construct a first text representation network, a second text representation network, and a third text representation network, wherein the first text representation network and the second text representation network are representation networks with similar semantics, and the second text representation network and the third text representation network are representation networks with different semantics;

[0095] The training module 904 is configured to input the text dataset into the first text representation network, the second text representation network, and the third text representation network to train the text classification model, so as to obtain a trained text classification model.

[0096] In an embodiment of the present application, before starting to train the model, a text feature representation network is constructed. In the process of constructing the text feature representation network, the present application constructs three representation networks: a first text representation network, a second text representation network, and a third text representation network. Among them, the first text representation network and the second text representation network are used as network representation structures with similar semantics, and the second text representation network and the third text representation network, as well as the first text representation network and the third text representation network, are used as network representation structures with different semantics. After constructing the text feature representation network, a text data set is obtained. In the training process of the model, the collected text data set is input into the first text representation network, the second text representation network, and the third text representation network to train the model. In the training process, the first text representation network, the second text representation network, and the third text representation network respectively output semantic vectors. The semantic vectors output by the first text representation network and the second text representation network are similar semantic vectors, and the semantic vector output by the third text representation network is different from the semantic vector output by the second text representation network. The model with the highest accuracy in the training process is determined based on the similar semantic vectors and the different semantic vectors, and the model with the highest progress is used as the text classification model after training.

[0097] Specifically, the first text representation network, the second text representation network, and the third text representation network can be constructed as an LSTM feature representation network and a Bert feature representation network. The original network structures of the first text representation network, the second text representation network, and the third text representation network are different, wherein the first text representation network has the same dropout value as the second text representation network and a different dropout value from the third text representation network, and the dropout value of the first text representation network is set to be smaller than the dropout value of the second text representation network, so that the constructed first text representation network and the second text representation network have similar semantic network representation structures, and the first text representation network and the third text representation network have different semantic network representation structures.

[0098] In related technologies, feature representation networks with identical throw values are constructed before model training to create semantically similar feature representations. During text classification model training, only semantically similar samples constructed using feature representation networks with identical throw values are considered, and training on semantically distinct samples is not performed. Consequently, the text classification model cannot fully train and learn semantically distinct samples.

[0099] In an embodiment of the present application, before training the model, a first text representation network and a second text representation network with similar semantics are constructed, and a third text representation network with different semantics is constructed. By training the first text representation network, the second text representation network and the third text representation, samples with similar semantics and samples with different semantics are fully taken into consideration, ensuring that samples with different semantics can be fully trained and learned during the learning process of the text classification model, so that the accuracy of the trained text classification model is higher.

[0100] In some embodiments of the present application, the text classification model training device 900 further includes:

[0101] An acquisition module, configured to acquire a first thrown value and a second thrown value;

[0102] The construction module 902 is further configured to construct a first text representation network and a second text representation network according to the first thrown value;

[0103] The construction module 902 is further configured to construct a third text representation network based on the second thrown value;

[0104] The first thrown value is smaller than the second thrown value.

[0105] In this embodiment of the present application, a first text representation network and a second text representation network are constructed using a first thrown value, thereby ensuring semantic similarity between the first and second text representation networks. Furthermore, a third text representation network is constructed using a second thrown value that is greater than the first thrown value, thereby ensuring semantic differences between the third text representation network and the first text representation network.

[0106] In some embodiments of the present application, the value range of the first throw value is greater than 0.00001 and less than 0.5; the value range of the second throw value is greater than 0.50001 and less than 1.

[0107] In the embodiment of the present application, the range of the first throw value is set to 0.00001 to 0.5, which can ensure that the first text representation network and the second text representation network constructed based on the first throw value are semantically similar network representation structures. The range of the second throw value is set to 0.50001 to 1, which can ensure that the third text representation network structure is a network representation structure with semantically different semantics from the first text representation network and the second text representation network.

[0108] In some possible implementations, the first throw value is selected as 0.2, and the second throw value is selected as 0.8.

[0109] In some embodiments of the present application, the acquisition module is further used to acquire text sample data;

[0110] The text classification model training device 900 further includes:

[0111] The word segmentation module is used to perform word segmentation processing on the text sample data according to preset rules to obtain a text dataset.

[0112] In the embodiment of the present application, before training the text classification model, the text sample data is segmented according to a preset trajectory to obtain a corresponding text dataset. The text dataset can be configured according to different usage requirements to ensure the accuracy of the trained text classification model.

[0113] In some embodiments of the present application, the acquisition module is further used to obtain a preset number of training times;

[0114] The preset number of training times is the number of times set in advance, that is, the number of cycles for training the text classification model.

[0115] The training module 904 is further configured to train multiple text classification models based on the text dataset according to a preset number of training times;

[0116] The acquisition module is also used to obtain the loss function corresponding to each text classification model in multiple text classification models;

[0117] The text classification model training device 900 further includes:

[0118] The determination module is used to determine a trained text classification model among multiple text classification models based on multiple loss functions.

[0119] In an embodiment of the present application, during the cyclic training of the text classification model, the text classification model after each training is recorded, and the trained text classification model with the highest accuracy is searched among them, ensuring that the trained text classification model finally obtained is the model with the highest accuracy.

[0120] In some embodiments of the present application, the acquisition module is further configured to acquire a first semantic vector, a second semantic vector, and a third semantic vector of each text classification model;

[0121] Among them, the first semantic vector corresponds to the first text representation network, the second semantic vector corresponds to the second text representation network, and the second semantic vector corresponds to the first text representation network.

[0122] The determination module is further used to determine the model loss function of the text classification model based on the first semantic vector, the second semantic vector and the third semantic vector.

[0123] In an embodiment of the present application, a loss function is calculated based on the first semantic vector, the second semantic vector, and the third semantic vector, and a trained text classification model among multiple text classification models in the training process is screened based on the calculated loss function.

[0124] In some embodiments of the present application, the determination module is further configured to determine a first loss function based on the first semantic vector and the preset vector;

[0125] The determination module is further configured to determine a second loss function based on the first semantic vector, the second semantic vector, and the third semantic vector;

[0126] The determination module is further used to determine the model loss function based on the first loss function and the second loss function.

[0127] In an embodiment of the present application, the text classification model is jointly trained through the first loss function and the second loss function, so that the text classification model can be more fully trained, the representation distance of samples with similar semantics is closer, and the representation distance of samples with different semantics is farther, and the model representation effect is better.

[0128] In some embodiments of the present application, the text classification model training device 900 further includes:

[0129] A calculation module, configured to calculate a first cosine distance between the first semantic vector and the second semantic vector;

[0130] The calculation module is further used to calculate the second cosine distance between the second semantic vector and the third semantic vector;

[0131] The determination module is further used to determine a second loss function based on the first cosine distance and the second cosine distance.

[0132] In the embodiment of the present application, in the process of determining the second loss function, it is necessary to calculate a first cosine distance between the first semantic vector and the second semantic vector, and a second pre-distance between the second semantic vector and the third semantic vector. The second loss function can be determined by the first cosine distance and the second cosine distance.

[0133] In some embodiments of the present application, the determination module is further configured to determine a third loss function based on the second semantic vector and the preset vector;

[0134] The determination module is further used to determine the model loss function based on the first loss function, the second loss function and the third loss function.

[0135] In an embodiment of the present application, a fourth loss function is constructed for the first text representation network and the second text representation network with similar semantics to perform unsupervised learning, so that the first semantic vector output by the first text representation network is close to the second semantic vector output by the second text representation network.

[0136] In some embodiments of the present application, the determining module is further configured to determine a fourth loss function based on the first semantic vector and the second semantic vector;

[0137] The determination module is also used to determine the model loss function based on the first loss function, the second loss function, the third loss function and the fourth loss function.

[0138] The embodiment of the present application sets the first loss function, the second loss function, the third loss function and the fourth loss function in the model, so that the output of the first text representation network and the output of the second text representation network are closer to the sample label, the output of the first text representation network is closer to the output of the second text representation network, and the output of the second text representation network is further away from the output of the third text representation network. Through the joint training of multiple loss functions, the number of required preset label samples can be reduced, and the number of training steps can be reduced.

[0139] The training device for the text classification model in the embodiments of the present application can be an electronic device or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., and the embodiments of the present application do not specifically limit it.

[0140] The training device for the text classification model in the embodiment of the present application can be a device having an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0141] The text classification model training device provided in the embodiment of the present application can implement each process implemented in the above method embodiment. To avoid repetition, it will not be described here.

[0142] Optionally, an embodiment of the present application further provides an electronic device, Figure 10 FIG. 1 shows a structural block diagram of an electronic device according to an embodiment of the present application. Figure 10 As shown, the electronic device 1000 includes a processor 1002 and a memory 1004. The memory 1004 stores programs or instructions that can be run on the processor 1002. When the program or instructions are executed by the processor 1002, the various steps of the above-mentioned method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, they are not described here.

[0143] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0144] Figure 11 A schematic diagram of the hardware structure of an electronic device implementing an embodiment of the present application.

[0145] The electronic device 1100 includes but is not limited to components such as a radio frequency unit 1101 , a network module 1102 , an audio output unit 1103 , an input unit 1104 , a sensor 1105 , a display unit 1106 , a user input unit 1107 , an interface unit 1108 , a memory 1109 , and a processor 1110 .

[0146] Those skilled in the art will understand that the electronic device 1100 may also include a power source (such as a battery) to power each component, and the power source may be logically connected to the processor 1110 through a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system. Figure 11 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be repeated here.

[0147] The processor 1110 is configured to construct a first text representation network, a second text representation network, and a third text representation network, wherein the first text representation network and the second text representation network are representation networks with similar semantics, and the second text representation network and the third text representation network are representation networks with different semantics;

[0148] The processor 1110 is configured to input the text dataset into the first text representation network, the second text representation network, and the third text representation network to train the text classification model, so as to obtain a trained text classification model.

[0149] In an embodiment of the present application, before training the model, a first text representation network and a second text representation network with similar semantics are constructed, and a third text representation network with different semantics is constructed. By training the first text representation network, the second text representation network and the third text representation, samples with similar semantics and samples with different semantics are fully taken into consideration, ensuring that samples with different semantics can be fully trained and learned during the learning process of the text classification model, so that the accuracy of the trained text classification model is higher.

[0150] Further, the processor 1110 is configured to obtain a first thrown value and a second thrown value;

[0151] The processor 1110 is further configured to construct a first text representation network and a second text representation network according to the first thrown value;

[0152] The processor 1110 is further configured to construct a third text representation network according to the second thrown value;

[0153] The first thrown value is smaller than the second thrown value.

[0154] In this embodiment of the present application, a first text representation network and a second text representation network are constructed using a first thrown value, thereby ensuring semantic similarity between the first and second text representation networks. Furthermore, a third text representation network is constructed using a second thrown value that is greater than the first thrown value, thereby ensuring semantic differences between the third text representation network and the first text representation network.

[0155] Furthermore, the value range of the first thrown value is greater than 0.00001 and less than 0.5; the value range of the second thrown value is greater than 0.50001 and less than 1.

[0156] In the embodiment of the present application, the range of the first throw value is set to 0.00001 to 0.5, which can ensure that the first text representation network and the second text representation network constructed based on the first throw value are semantically similar network representation structures. The range of the second throw value is set to 0.50001 to 1, which can ensure that the third text representation network structure is a network representation structure with semantically different semantics from the first text representation network and the second text representation network.

[0157] In some possible implementations, the first throw value is selected as 0.2, and the second throw value is selected as 0.8.

[0158] Furthermore, the processor 110 is further configured to obtain text sample data;

[0159] The processor 110 is further configured to perform word segmentation processing on the text sample data according to preset rules to obtain a text data set.

[0160] In the embodiment of the present application, before training the text classification model, the text sample data is segmented according to a preset trajectory to obtain a corresponding text dataset. The text dataset can be configured according to different usage requirements to ensure the accuracy of the trained text classification model.

[0161] Furthermore, the processor 1110 is further configured to obtain a preset number of training times;

[0162] The preset number of training times is the number of times set in advance, that is, the number of cycles for training the text classification model.

[0163] The processor 1110 is further configured to train multiple text classification models based on the text dataset according to a preset number of training times;

[0164] The processor 1110 is further configured to obtain a loss function corresponding to each of the multiple text classification models:

[0165] The processor 1110 is configured to determine a trained text classification model from among multiple text classification models based on multiple loss functions.

[0166] In an embodiment of the present application, during the cyclic training of the text classification model, the text classification model after each training is recorded, and the trained text classification model with the highest accuracy is searched among them, ensuring that the trained text classification model finally obtained is the model with the highest accuracy.

[0167] Furthermore, the processor 1110 is further configured to obtain a first semantic vector, a second semantic vector, and a third semantic vector of each text classification model;

[0168] Among them, the first semantic vector corresponds to the first text representation network, the second semantic vector corresponds to the second text representation network, and the second semantic vector corresponds to the first text representation network.

[0169] The processor 1110 is further configured to determine a model loss function of the text classification model based on the first semantic vector, the second semantic vector, and the third semantic vector.

[0170] In an embodiment of the present application, a loss function is calculated based on the first semantic vector, the second semantic vector, and the third semantic vector, and a trained text classification model among multiple text classification models in the training process is screened based on the calculated loss function.

[0171] Furthermore, the processor 1110 is further configured to determine a first loss function based on the first semantic vector and the preset vector;

[0172] The processor 1110 is further configured to determine a second loss function based on the first semantic vector, the second semantic vector, and the third semantic vector;

[0173] The processor 1110 is further configured to determine a model loss function based on the first loss function and the second loss function.

[0174] In an embodiment of the present application, the text classification model is jointly trained through the first loss function and the second loss function, so that the text classification model can be more fully trained, the representation distance of samples with similar semantics is closer, and the representation distance of samples with different semantics is farther, and the model representation effect is better.

[0175] Further, the processor 1110 is configured to calculate a first cosine distance between the first semantic vector and the second semantic vector;

[0176] The processor 1110 is further configured to calculate a second cosine distance between the second semantic vector and the third semantic vector;

[0177] The processor 1110 is further configured to determine a second loss function based on the first cosine distance and the second cosine distance.

[0178] In the embodiment of the present application, in the process of determining the second loss function, it is necessary to calculate a first cosine distance between the first semantic vector and the second semantic vector, and a second pre-distance between the second semantic vector and the third semantic vector. The second loss function can be determined by the first cosine distance and the second cosine distance.

[0179] Furthermore, the processor 1110 is further configured to determine a third loss function based on the second semantic vector and the preset vector;

[0180] The processor 1110 is further configured to determine a model loss function based on the first loss function, the second loss function, and the third loss function.

[0181] In an embodiment of the present application, a fourth loss function is constructed for the first text representation network and the second text representation network with similar semantics to perform unsupervised learning, so that the first semantic vector output by the first text representation network is close to the second semantic vector output by the second text representation network.

[0182] Furthermore, the processor 1110 is further configured to determine a fourth loss function based on the first semantic vector and the second semantic vector;

[0183] The processor 1110 is further configured to determine a model loss function based on the first loss function, the second loss function, the third loss function, and the fourth loss function.

[0184] The embodiment of the present application sets the first loss function, the second loss function, the third loss function and the fourth loss function in the model, so that the output of the first text representation network and the output of the second text representation network are closer to the sample label, the output of the first text representation network is closer to the output of the second text representation network, and the output of the second text representation network is further away from the output of the third text representation network. Through the joint training of multiple loss functions, the number of required preset label samples can be reduced, and the number of training steps can be reduced.

[0185] It should be understood that in an embodiment of the present application, the input unit 1104 may include a graphics processing unit (GPU) 11041 and a microphone 11042, and the graphics processor 11041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 1106 may include a display panel 11061, and the display panel 11061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1107 includes a touch panel 11071 and at least one of other input devices 11072. The touch panel 11071 is also called a touch screen. The touch panel 11071 may include two parts: a touch detection device and a touch controller. Other input devices 11072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and an operating stick, which will not be repeated here.

[0186] The memory 1109 can be used to store software programs and various data. The memory 1109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1109 may include a volatile memory or a non-volatile memory, or the memory 1109 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 1109 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0187] Processor 1110 may include one or more processing units. Optionally, processor 1110 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 1110.

[0188] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned text classification model training method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0189] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0190] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned text classification model training method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0191] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0192] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the training method embodiment of the above-mentioned text classification model, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0193] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0194] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0195] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A training method for a text classification model, characterized in that: include: Constructing a first text representation network, a second text representation network, and a third text representation network, wherein the first text representation network and the second text representation network are representation networks with similar semantics, and the second text representation network and the third text representation network are representation networks with different semantics; Inputting the text data set into the first text representation network, the second text representation network, and the third text representation network to train the text classification model to obtain a trained text classification model; The constructing of the first text representation network, the second text representation network, and the third text representation network includes: Get the first thrown value and the second thrown value; constructing the first text representation network and the second text representation network according to the first thrown value; constructing the third text representation network according to the second thrown value; Wherein, the first thrown value is smaller than the second thrown value.

2. The text classification model training method according to claim 1, characterized in that: The value range of the first thrown value is greater than 0.00001 and less than 0.5; The value range of the second thrown value is greater than 0.50001 and less than 1.

3. The text classification model training method according to claim 1 or 2, characterized in that: Before inputting the text dataset into the first text representation network, the second text representation network, and the third text representation network to train the text classification model to obtain the trained text classification model, the method further includes: Get text sample data; According to preset rules, the text sample data is segmented to obtain the text data set.

4. The text classification model training method according to claim 1 or 2, characterized in that: Inputting the text dataset into the first text representation network, the second text representation network, and the third text representation network for training to obtain a trained text classification model includes: Get the preset number of training times; Training a plurality of text classification models based on the text dataset according to the preset number of training times; Obtaining a loss function corresponding to each of the text classification models in the plurality of text classification models; A trained text classification model among the multiple text classification models is determined according to the multiple loss functions.

5. The text classification model training method according to claim 3, characterized in that: The acquiring of the loss function corresponding to each of the plurality of text classification models further includes: Obtaining a first semantic vector, a second semantic vector, and a third semantic vector for each of the text classification models; Determining a model loss function of the text classification model based on the first semantic vector, the second semantic vector, and the third semantic vector; Among them, the first semantic vector corresponds to the first text representation network, the second semantic vector corresponds to the second text representation network, and the second semantic vector corresponds to the first text representation network.

6. The text classification model training method according to claim 5, characterized in that: The determining of the model loss function of the text classification model according to the first semantic vector, the second semantic vector, and the third semantic vector includes: Determining a first loss function according to the first semantic vector and a preset vector; Determining a second loss function based on the first semantic vector, the second semantic vector, and the third semantic vector; Determine the model loss function based on the first loss function and the second loss function.

7. The text classification model training method according to claim 6, characterized in that: The determining a second loss function according to the first semantic vector, the second semantic vector, and the third semantic vector includes: Calculating a first cosine distance between the first semantic vector and the second semantic vector; Calculating a second cosine distance between the second semantic vector and the third semantic vector; A second loss function is determined according to the first cosine distance and the second cosine distance.

8. The text classification model training method according to claim 7, characterized in that: Determining the model loss function according to the first loss function and the second loss function includes: Determining a third loss function based on the second semantic vector and the preset vector; Determine the model loss function based on the first loss function, the second loss function and the third loss function.

9. The text classification model training method according to claim 8, characterized in that: The determining the model loss function according to the first loss function, the second loss function, and the third loss function includes: Determining a fourth loss function based on the first semantic vector and the second semantic vector; The model loss function is determined based on the first loss function, the second loss function, the third loss function and the fourth loss function.

10. A training device for a text classification model, characterized in that: include: A construction module is used to construct a first text representation network, a second text representation network, and a third text representation network, wherein the first text representation network and the second text representation network are representation networks with similar semantics, and the second text representation network and the third text representation network are representation networks with different semantics; A training module, configured to input a text dataset into the first text representation network, the second text representation network, and the third text representation network to train the text classification model, thereby obtaining a trained text classification model; An acquisition module, configured to acquire a first thrown value and a second thrown value; The construction module is further configured to construct the first text representation network and the second text representation network according to the first thrown value; and to construct the third text representation network according to the second thrown value; Wherein, the first thrown value is smaller than the second thrown value.

11. An electronic device, characterized in that: include: a memory on which programs or instructions are stored; A processor, configured to implement the steps of the text classification model training method according to any one of claims 1 to 9 when executing the program or instruction.

12. A readable storage medium having a program or instruction stored thereon, characterized in that: When the program or instruction is executed by a processor, the steps of the text classification model training method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Image Chinese subtitle generation method

    CN107909115A

  • Accelerating method of Chinese handwriting recognition based on convolution neural network

    CN109034281A

  • Chinese speech recognition method based on acoustic and language model training and joint optimization

    CN113808581A