Text classification method, device, equipment and storage medium based on artificial intelligence
By simultaneously performing mask training and text classification prediction training in the fine-tuning stage, the problem of low robustness of the Bert model is solved, improving the applicability of the model and the accuracy of text classification prediction.
Patent Information
- Application Number
- CN202210033719.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-01-12
AI Technical Summary
The pre-training stage of the Bert model uses the MLM training method and the fine-tuning training stage to obtain the output of the flag bits for classification prediction training. Due to the major changes in the training methods of the two stages, the model obtained by fine-tuning training is less robust.
A text classification method based on artificial intelligence is proposed. By obtaining multiple training samples, emotional word dictionary and synonym dictionary, mask training is carried out on the initial model, and combining the Bert model and classification prediction layer, mask training and text classification prediction training are carried out simultaneously in the fine-tuning stage.
It effectively alleviates the differences between the two training stages of the Bert model, increases the robustness of the trained model, makes the model more suitable for specific application scenarios, and improves the accuracy of text classification prediction.
Smart Images

Figure CN114416984B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an artificial intelligence-based text classification method, device, equipment and storage medium. Background Art
[0002] With the widespread application of the Bert (Bidirectional Encoder Representations from Transformers) model in the field of natural language processing (NLP), current NLP tasks are usually implemented based on the Bert model. Better results can be achieved by fine-tuning existing NLP tasks based on the pre-trained model.
[0003] The pre-training task of the Bert model is MLM (masked language model), which mainly masks some tokens (words) at the input stage of the model and predicts these masked tokens at the output stage of the model. In the fine-tuning training stage of the text classification task based on the Bert model, the MLM task is abandoned. The text is first input into the model, and then the final classification prediction is made by using the information of the [CLS] (marker) position at the output stage. Although this traditional fine-tuning training can effectively improve the accuracy of the final classification by using self-attention, there is a major change in the training methods of the two stages. The former predicts the masked tokens, and the latter obtains the information of [CLS] for classification prediction, resulting in a low robustness of the model obtained by fine-tuning training. Summary of the invention
[0004] The main purpose of this application is to provide a text classification method, device, equipment and storage medium based on artificial intelligence, aiming to solve the technical problem that the prior art uses the MLM training method in the Bert pre-training stage and obtains the output of the flag bit for classification prediction training in the fine-tuning training stage, and the model obtained by fine-tuning training has low robustness due to the major changes in the training methods of the two stages.
[0005] In order to achieve the above-mentioned invention object, the present application proposes a text classification method based on artificial intelligence, the method comprising:
[0006] Get the target text;
[0007] Inputting the target text into a preset text classification model for text classification prediction;
[0008] Obtaining a text classification result output by the text classification model as a target text classification result corresponding to the target text;
[0009] The text classification model is obtained through the following steps:
[0010] Using the acquired multiple training samples, sentiment word dictionary and synonym dictionary to perform mask training on the initial model, wherein the initial model is a model obtained based on the Bert model and the classification prediction layer, and the sentiment word dictionary and the synonym dictionary are used to replace words in the initial text samples in the training samples;
[0011] The initial model after training is used as the text classification model.
[0012] Furthermore, before the step of inputting the target text into a preset text classification model for text classification prediction, the method further includes:
[0013] Acquire a plurality of the training samples;
[0014] Sequentially acquiring the training samples from each of the training samples as samples to be trained;
[0015] Acquire words from the initial text sample of the sample to be trained according to a preset ratio to obtain a set of words to be replaced, and use the set of words to be replaced as a word calibration value;
[0016] Using the sentiment word dictionary to determine whether each word in the to-be-replaced word set is a sentiment word, and obtaining a sentiment word set and a non-sentiment word set;
[0017] According to a preset mask, the synonym dictionary, the to-be-replaced word set, the sentiment word set and the non-sentiment word set, the initial text sample of the to-be-trained sample is subjected to word replacement to obtain a target text sample;
[0018] Input the target text sample into the initial model to perform word prediction and text classification prediction at the mask position, respectively, to obtain a word prediction value and a text classification prediction value;
[0019] Training the initial model according to the text classification calibration value of the sample to be trained, the word calibration value, the word prediction value and the text classification prediction value;
[0020] Repeat the step of determining the samples to be trained until a preset training target is reached;
[0021] The initial model that achieves the training goal is used as the text classification model.
[0022] Further, the step of performing word replacement on the initial text sample of the sample to be trained according to the preset mask, the synonym dictionary, the set of words to be replaced, the emotional word set and the non-emotional word set to obtain the target text sample includes:
[0023] Using the mask symbol, replace each word corresponding to the non-emotional word set in the initial text sample of the sample to be trained, to obtain a text sample to be processed;
[0024] The synonym dictionary is used to replace each word corresponding to the sentiment word set in the text sample to be processed to obtain the target text sample.
[0025] Furthermore, the step of training the initial model according to the text classification calibration value of the sample to be trained, the word calibration value, the word prediction value and the text classification prediction value includes:
[0026] Inputting the text classification calibration value, the word calibration value, the word prediction value and the text classification prediction value of the sample to be trained into a preset target loss function to calculate the loss value, and obtain a target loss value;
[0027] The network parameters of the initial model are updated according to the target loss value.
[0028] Furthermore, the step of inputting the text classification calibration value, the word calibration value, the word prediction value and the text classification prediction value of the sample to be trained into a preset target loss function to calculate the loss value to obtain the target loss value includes:
[0029] Inputting the word calibration value and the word prediction value of the sample to be trained into a preset word prediction loss function to calculate the loss value, thereby obtaining a first loss value;
[0030] Inputting the text classification calibration value and the text classification prediction value of the sample to be trained into a preset text classification prediction loss function to calculate the loss value, thereby obtaining a second loss value;
[0031] Performing a weighted summation of the first loss value and the second loss value to obtain the target loss value;
[0032] Among them, the word prediction loss function and the text classification prediction loss function both adopt the cross entropy loss function.
[0033] Furthermore, before the step of obtaining the plurality of training samples, the step further includes:
[0034] Get multiple product review texts;
[0035] Deleting blank characters and repeated punctuation marks from each product review text to obtain a preprocessed text;
[0036] The training sample is generated according to each of the preprocessed texts.
[0037] Furthermore, the step of generating the training sample according to each of the preprocessed texts includes:
[0038] Obtaining positive and negative classification prediction results corresponding to each of the preprocessed texts;
[0039] The preprocessed text is used as an initial text sample of the training sample, and the positive and negative classification prediction results are used as text classification calibration values of the training sample.
[0040] The present application also proposes a text classification device based on artificial intelligence, the device comprising:
[0041] A text acquisition module is used to acquire the target text;
[0042] A text classification module, used for inputting the target text into a preset text classification model for text classification prediction;
[0043] A target text classification result determination module is used to obtain the text classification result output by the text classification model as the target text classification result corresponding to the target text;
[0044] The model training module is used to perform mask training on the initial model using the acquired multiple training samples, a preset sentiment word dictionary and a preset synonym dictionary, and use the initial model after the training as the text classification model, wherein the initial model includes: a Bert model and a text classification layer, and the sentiment word dictionary and the synonym dictionary are used to replace words in the initial text samples in the training samples.
[0045] The present application also proposes a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.
[0046] The present application also proposes a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0047] The present application discloses an artificial intelligence-based text classification method, apparatus, device and storage medium, wherein the method first obtains an initial model based on a Bert model and a classification prediction layer, then replaces words in the initial text sample in the training sample through the sentiment word dictionary and the synonym dictionary, and finally performs mask training on the initial model with the replaced training sample, thereby achieving simultaneous mask training and text classification prediction training in the fine-tuning stage, effectively alleviating the difference between the two training stages of the Bert model, increasing the robustness of the trained model, and making the trained model more suitable for specific application scenarios; by performing text classification prediction on the model obtained by simultaneously performing mask training and text classification prediction training in the fine-tuning stage, the accuracy of text classification prediction is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A flowchart of a text classification method based on artificial intelligence according to an embodiment of the present application;
[0049] Figure 2 This is a schematic block diagram of the structure of a text classification device based on artificial intelligence according to an embodiment of the present application;
[0050] Figure 3 A schematic block diagram of the structure of a computer device according to an embodiment of the present application.
[0051] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0053] Reference Figure 1 In an embodiment of the present application, a text classification method based on artificial intelligence is provided, and the method comprises:
[0054] S1: Get the target text;
[0055] S2: Inputting the target text into a preset text classification model for text classification prediction;
[0056] S3: Obtaining a text classification result output by the text classification model as a target text classification result corresponding to the target text;
[0057] The text classification model is obtained through the following steps:
[0058] Using the acquired multiple training samples, sentiment word dictionary and synonym dictionary to perform mask training on the initial model, wherein the initial model is a model obtained based on the Bert model and the classification prediction layer, and the sentiment word dictionary and the synonym dictionary are used to replace words in the initial text samples in the training samples;
[0059] The initial model after training is used as the text classification model.
[0060] This embodiment first obtains an initial model based on the Bert model and the classification prediction layer, then replaces words in the initial text sample in the training sample through the sentiment word dictionary and the synonym dictionary, and finally performs mask training on the initial model with the replaced training sample, so as to realize simultaneous mask training and text classification prediction training in the fine-tuning stage, effectively alleviate the difference between the two training stages of the Bert model, increase the robustness of the trained model, and make the trained model more suitable for specific application scenarios; by performing text classification prediction on the model obtained by simultaneously performing mask training and text classification prediction training in the fine-tuning stage, the accuracy of text classification prediction is improved.
[0061] For S1, the target text input by the user can be obtained, the target text sent by a third-party application system can be obtained, and the target text can also be obtained from a database.
[0062] The target text is the text that needs to be classified and predicted.
[0063] For S3, the text classification result output by the text classification model is obtained, and the obtained text classification result is used as the target text classification result corresponding to the target text.
[0064] Among them, the initial model is masked and trained using multiple training samples, sentiment word dictionaries, and synonym dictionaries, so that mask training and text classification prediction training can be performed simultaneously in the fine-tuning stage.
[0065] The training sample includes: an initial text sample and a text classification calibration value. The text classification calibration value is an accurate calibration result of the classification label of the initial text sample.
[0066] Optionally, the initial text sample is a text obtained by preprocessing the product review text. It is understandable that the initial text sample can also be generated based on other texts, which is not limited here.
[0067] The initial model is a model obtained based on the Bert model and the classification prediction layer, and the classification prediction layer is used to perform classification prediction on the information output by the Bert model for the flag bit.
[0068] The flag bit is expressed as: [CLS].
[0069] The sentiment word dictionary is used to determine whether the original word at the mask position is a sentiment word, and determine a replacement strategy for the mask position according to the determination result. The mask position in the initial text sample is replaced according to the replacement strategy, the synonym dictionary and the preset mask symbol.
[0070] The mask character is: [MASK].
[0071] In one embodiment, before the step of inputting the target text into a preset text classification model for text classification prediction, the method further includes:
[0072] S21: Acquire a plurality of the training samples;
[0073] S22: sequentially acquiring the training samples from each of the training samples as samples to be trained;
[0074] S23: Acquire words from the initial text sample of the sample to be trained according to a preset ratio to obtain a set of words to be replaced, and use the set of words to be replaced as a word calibration value;
[0075] S24: using the sentiment word dictionary to determine whether each word in the to-be-replaced word set is a sentiment word, and obtaining a sentiment word set and a non-sentiment word set;
[0076] S25: performing word replacement on the initial text sample of the sample to be trained according to a preset mask, the synonym dictionary, the set of words to be replaced, the set of sentiment words and the set of non-sentiment words to obtain a target text sample;
[0077] S26: Inputting the target text sample into the initial model to perform word prediction and text classification prediction at the mask position, respectively, to obtain a word prediction value and a text classification prediction value;
[0078] S27: training the initial model according to the text classification calibration value of the sample to be trained, the word calibration value, the word prediction value and the text classification prediction value;
[0079] S28: Repeat the step of determining the samples to be trained until a preset training target is reached;
[0080] S29: Using the initial model that achieves the training goal as the text classification model.
[0081] This embodiment first determines the set of words to be replaced and the word calibration values from the initial text sample, uses the sentiment word dictionary to judge whether the original word at the mask position is a sentiment word, and then replaces the mask position according to the result of the sentiment word judgment, the synonym dictionary and the preset mask symbol, and finally implements mask training and classification prediction training simultaneously according to the replaced text sample.
[0082] For S21, a plurality of the training samples input by the user may be obtained, a plurality of the training samples sent by a third-party application system may be obtained, or a plurality of the training samples may be obtained from a database.
[0083] For S22, any one of the training samples is obtained from the training samples, and the obtained training sample is used as a sample to be trained.
[0084] For S23, words are obtained from the initial text sample of the training sample according to a preset ratio, and each word obtained is used as the set of words to be replaced. That is, the position of each word in the set of words to be replaced corresponding to the initial text sample of the training sample is the mask position.
[0085] Optionally, the number of words in the word set to be replaced is multiplied by the replacement ratio to obtain a product result, and when the product result is an integer, the product result is used as the number of words in the word set to be replaced.
[0086] Optionally, the product result is obtained by multiplying the number of words in the word set to be replaced by the replacement ratio. When the product result is not an integer, the product result is rounded up and the number obtained by rounding up is used as the number of words in the word set to be replaced.
[0087] Optionally, the product result is obtained by multiplying the number of words in the word set to be replaced by the replacement ratio. When the product result is not an integer, the product result is rounded down and the number obtained by rounding down is used as the number of words in the word set to be replaced.
[0088] The replacement ratio can be set to any one of 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, and 15%.
[0089] The set of words to be replaced is the original words at each mask position, so the set of words to be replaced can be directly used as the word calibration value.
[0090] For S24, each word in the set of words to be replaced is searched in the sentiment word dictionary, and the sentiment word judgment result corresponding to the successfully searched word is determined as a sentiment word, and the sentiment word judgment result corresponding to the failed search word is determined as a non-sentiment word. In other words, the sentiment word judgment result corresponds to the words in the set of words to be replaced one by one.
[0091] The set of words to be replaced is divided into sets according to the sentiment word judgment result to obtain a sentiment word set and a non-sentiment word set. That is, the sentiment word set includes all the words in the set of words to be replaced whose sentiment word judgment result is sentiment, and the non-sentiment word set includes all the words in the set of words to be replaced whose sentiment word judgment result is non-sentiment.
[0092] For S25, determine the replacement strategy for each mask position corresponding to the sentiment word set and the non-sentiment word set, and perform word replacement on the initial text sample of the sample to be trained based on the replacement strategy, the preset mask symbol, the synonym dictionary, and the set of words to be replaced, and use the replaced initial text sample as the target text sample.
[0093] For S26, the target text sample is input into the initial model, the Bert model of the initial model performs word prediction for the mask position, and the result of the word prediction is used as the word prediction value, and the classification prediction layer of the initial model performs text classification prediction for the information corresponding to the flag position output by the Bert model of the initial model, and the data obtained by the text classification prediction is used as the text classification prediction value.
[0094] For S27, a target loss value is calculated according to the text classification calibration value of the sample to be trained, the word calibration value, the word prediction value and the text classification prediction value, and the network parameters of the initial model are updated according to the target loss value.
[0095] The method steps for updating the network parameters of the initial model according to the target loss value are not described in detail here.
[0096] For S28, the step of determining the samples to be trained is repeated, that is, steps S22 to S28 are repeated until a preset training target is reached.
[0097] The training objectives include: the target loss value reaches a first convergence condition or the number of iterations of the initial model reaches a second convergence condition.
[0098] The first convergence condition refers to that the magnitudes of the target loss values calculated twice consecutively satisfy the Lipschitz condition (Lipschitz continuity condition).
[0099] The number of iterations refers to the number of times the loss value of the initial model is calculated, that is, after being calculated once, the number of iterations increases by 1.
[0100] The second convergence condition is a specific value.
[0101] For S29, the initial model that achieves the training goal is a model that meets the expected requirements, so the initial model that achieves the training goal is directly used as the text classification model.
[0102] In one embodiment, the step of performing word replacement on the initial text sample of the sample to be trained according to the preset mask, the synonym dictionary, the set of words to be replaced, the sentiment word set and the non-sentiment word set to obtain the target text sample includes:
[0103] S251: using the mask symbol to replace each word corresponding to the non-emotional word set in the initial text sample of the sample to be trained, to obtain a text sample to be processed;
[0104] S252: Using the synonym dictionary, replace each word corresponding to the sentiment word set in the text sample to be processed to obtain the target text sample.
[0105] In this embodiment, the mask positions corresponding to the sentiment word set are replaced by the synonym dictionary, and the mask positions corresponding to the non-sentiment word set are replaced by mask symbols, thereby maintaining the consistency of the sentiment tendency between the initial text sample and the target text sample, improving the accuracy of model training, and improving the accuracy of text classification prediction.
[0106] For S251, each word corresponding to the non-emotional word set in the initial text sample of the sample to be trained is replaced with the mask symbol, and the initial text sample after replacement is used as the text sample to be processed.
[0107] For S252, any word in the sentiment word set is taken as a target word; the target word is matched with a synonym from the synonym dictionary, and the matched synonym is used to replace the word at the mask position corresponding to the target word in the text sample to be processed; the step of taking any word in the sentiment word set as a target word is repeated until the replacement of the mask position corresponding to each word in the sentiment word set is completed.
[0108] In one embodiment, the step of training the initial model according to the text classification calibration value, the word calibration value, the word prediction value and the text classification prediction value of the sample to be trained includes:
[0109] S271: Inputting the text classification calibration value, the word calibration value, the word prediction value and the text classification prediction value of the sample to be trained into a preset target loss function to calculate the loss value, and obtain a target loss value;
[0110] S272: Update the network parameters of the initial model according to the target loss value.
[0111] In this embodiment, the text classification calibration value, the word calibration value, the word prediction value and the text classification prediction value of the sample to be trained are input into a preset target loss function to calculate the loss value, thereby updating the network parameters of the initial model according to the loss of word prediction and the loss of text classification prediction in the fine-tuning stage, thereby realizing the simultaneous performance of mask training and text classification prediction training in the fine-tuning stage.
[0112] For S271, the text classification calibration value, the word calibration value, the word prediction value and the text classification prediction value of the sample to be trained are input into a preset target loss function to calculate the loss value, wherein the target loss function is a function obtained based on the cross entropy loss function.
[0113] For S272, the network parameters of the initial model are updated according to the target loss value, and the updated initial model is used for the next calculation of the word prediction value and the text classification prediction value, thereby realizing the iterative update of the network parameters of the initial model.
[0114] In one embodiment, the step of inputting the text classification calibration value, the word calibration value, the word prediction value and the text classification prediction value of the sample to be trained into a preset target loss function to calculate the loss value to obtain the target loss value includes:
[0115] S2711: Inputting the word calibration value and the word prediction value of the sample to be trained into a preset word prediction loss function to calculate the loss value, and obtaining a first loss value;
[0116] S2712: Inputting the text classification calibration value and the text classification prediction value of the sample to be trained into a preset text classification prediction loss function to calculate the loss value, and obtaining a second loss value;
[0117] S2713: Perform a weighted summation of the first loss value and the second loss value to obtain the target loss value;
[0118] Among them, the word prediction loss function and the text classification prediction loss function both adopt the cross entropy loss function.
[0119] This embodiment uses a cross entropy loss function to calculate the loss of word prediction, and uses a cross entropy loss function to calculate the loss of text classification prediction. The loss of word prediction and the loss of text classification prediction are weighted summed as the target loss function, thereby simultaneously updating the network parameters of the initial model according to the loss of word prediction and the loss of text classification prediction.
[0120] For S2711, the word calibration value and the word prediction value of the sample to be trained are input into a preset word prediction loss function to calculate the loss value, and the calculated loss value is used as the first loss value.
[0121] For S2712, the text classification calibration value and the text classification prediction value of the sample to be trained are input into a preset text classification prediction loss function to calculate the loss value, and the calculated loss value is used as the second loss value.
[0122] For S2713, the first loss value and the second loss value are weightedly summed, and the data obtained by the weighted sum is used as the target loss value.
[0123] Among them, when the first loss value and the second loss value are weightedly summed, the ratio of the second loss value to the target loss value ranges from 0% to 50%, which may include 0% and may also include 50%.
[0124] Optionally, the ratio of the second loss value to the target loss value is set to 30%.
[0125] In one embodiment, before the step of obtaining a plurality of training samples, the method further includes:
[0126] S211: Obtain multiple product review texts;
[0127] S212: performing a process of deleting blank characters and repeated punctuation marks on each of the product review texts to obtain a preprocessed text;
[0128] S213: Generate the training sample according to each of the preprocessed texts.
[0129] In this embodiment, blank characters and repeated punctuation marks are deleted from the product review text to obtain preprocessed text, and training samples are generated based on the preprocessed text, thereby improving the accuracy of the generated training samples.
[0130] For S211, multiple product review texts input by the user may be obtained, multiple product review texts sent by a third-party application system may be obtained, and multiple product review texts may be obtained from a database.
[0131] The product review text is the text of the user's review of the product.
[0132] For S212, a regular expression is used to perform whitespace character removal and repeated punctuation mark removal processing on the product review text, and the product review text after the whitespace character removal and repeated punctuation mark removal processing is completed is used as the preprocessed text.
[0133] For S213, a calibration value is determined according to the preprocessed text, and then a training sample is generated according to the preprocessed text and the determined calibration value. That is, the preprocessed text corresponds to the training sample one by one.
[0134] The calibration value can be the classification prediction result of emotion or the classification prediction result of product satisfaction.
[0135] In one embodiment, the step of generating the training sample according to each of the preprocessed texts includes:
[0136] S2131: Obtaining positive and negative classification prediction results corresponding to each of the preprocessed texts;
[0137] S2132: Using the preprocessed text as the initial text sample of the training sample, and using the positive and negative classification prediction results as the text classification calibration values of the training sample.
[0138] In this embodiment, the positive and negative classification prediction results are used as text classification calibration values, so that the model trained by the training samples is suitable for classification prediction of positive and negative emotions.
[0139] For S2131, the positive and negative classification prediction results corresponding to the preprocessed text sent by the user are obtained.
[0140] For S2132, the preprocessed text is used as the initial text sample of the training sample, and the positive and negative classification prediction results are used as the text classification calibration values of the training sample, thereby generating a training sample.
[0141] Reference Figure 2 , the present application also proposes a text classification device based on artificial intelligence, the device comprising:
[0142] A text acquisition module 100 is used to acquire a target text;
[0143] The text classification module 200 is used to input the target text into a preset text classification model for text classification prediction;
[0144] A target text classification result determination module 300 is used to obtain the text classification result output by the text classification model as the target text classification result corresponding to the target text;
[0145] The model training module 400 is used to perform mask training on the initial model using the acquired multiple training samples, a preset sentiment word dictionary and a preset synonym dictionary, and use the initial model after the training as the text classification model, wherein the initial model includes: a Bert model and a text classification layer, and the sentiment word dictionary and the synonym dictionary are used to replace words in the initial text samples in the training samples.
[0146] This embodiment first obtains an initial model based on the Bert model and the classification prediction layer, then replaces words in the initial text sample in the training sample through the sentiment word dictionary and the synonym dictionary, and finally performs mask training on the initial model with the replaced training sample, so as to realize simultaneous mask training and text classification prediction training in the fine-tuning stage, effectively alleviate the difference between the two training stages of the Bert model, increase the robustness of the trained model, and make the trained model more suitable for specific application scenarios; by performing text classification prediction on the model obtained by simultaneously performing mask training and text classification prediction training in the fine-tuning stage, the accuracy of text classification prediction is improved.
[0147] In one embodiment, the above-mentioned model training module 400 further includes: a training sample acquisition submodule, a mask submodule and a training submodule;
[0148] The training sample acquisition submodule is used to acquire a plurality of the training samples;
[0149] The mask submodule is used to obtain one of the training samples as a sample to be trained from each of the training samples, obtain words from the initial text sample of the sample to be trained, obtain a set of words to be replaced and a word calibration value, use the sentiment word dictionary to determine whether each word in the set of words to be replaced is a sentiment word, and obtain a sentiment word determination result, and perform word replacement on the initial text sample of the sample to be trained according to a preset mask, the synonym dictionary, the set of words to be replaced and each of the sentiment word determination results to obtain a target text sample;
[0150] The training submodule is used to input the target text sample into the initial model to perform word prediction and text classification prediction at the mask position respectively, obtain word prediction values and text classification prediction values, train the initial model according to the text classification calibration value of the sample to be trained, the word calibration value, the word prediction value and the text classification prediction value, repeat the step of determining the sample to be trained until a preset training goal is reached, and use the initial model that reaches the training goal as the text classification model.
[0151] In one embodiment, the mask submodule comprises: a to-be-replaced word set determination unit and a word calibration value determination unit;
[0152] The to-be-replaced word set determination unit is used to randomly acquire words from the initial text sample of the to-be-trained sample using a preset replacement ratio to obtain a to-be-replaced word set;
[0153] The word calibration value determination unit is used to use the to-be-replaced word set as the word calibration value.
[0154] In one embodiment, the mask submodule further includes: a set partitioning unit, a first masking unit, and a second masking unit;
[0155] The set division unit is used to divide the to-be-replaced word set into sets according to the emotion word judgment result to obtain an emotion word set and a non-emotion word set;
[0156] The first masking unit is used to replace each word corresponding to the non-emotional word set in the initial text sample of the sample to be trained by using the masking symbol to obtain a text sample to be processed;
[0157] The second masking unit is used to use the synonym dictionary to replace each word corresponding to the sentiment word set in the text sample to be processed to obtain the target text sample.
[0158] In one embodiment, the training submodule includes: a target loss value calculation unit and a network parameter updating unit;
[0159] The target loss value calculation unit is used to input the text classification calibration value of the sample to be trained, the word calibration value, the word prediction value and the text classification prediction value into a preset target loss function to calculate the loss value and obtain a target loss value;
[0160] The network parameter updating unit is used to update the network parameters of the initial model according to the target loss value, and use the updated initial model for the next calculation of the word prediction value and the text classification prediction value.
[0161] In one embodiment, the target loss value calculation unit includes: a first loss value calculation subunit, a second loss value calculation subunit and a weighted summation subunit;
[0162] The first loss value calculation subunit is used to input the word calibration value and the word prediction value of the sample to be trained into a preset word prediction loss function to calculate the loss value, so as to obtain a first loss value;
[0163] The second loss value calculation subunit is used to input the text classification calibration value and the text classification prediction value of the sample to be trained into a preset text classification prediction loss function to perform loss value calculation to obtain a second loss value;
[0164] The weighted summation subunit is used to perform weighted summation on the first loss value and the second loss value to obtain the target loss value;
[0165] Among them, the word prediction loss function and the text classification prediction loss function both adopt the cross entropy loss function.
[0166] In one embodiment, the training sample acquisition submodule includes: a product review text acquisition unit and a training sample generation unit;
[0167] The product review text acquisition unit is used to acquire multiple product review texts;
[0168] The training sample generation unit is used to delete blank characters and repeated punctuation marks from each of the product review texts to obtain preprocessed texts, obtain positive and negative classification prediction results corresponding to each of the preprocessed texts, use the preprocessed texts as the initial text samples of the training samples, and use the positive and negative classification prediction results as the text classification calibration values of the training samples.
[0169] Reference Figure 3 In the embodiment of the present application, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 3As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor designed by the computer is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as a text classification method based on artificial intelligence. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a text classification method based on artificial intelligence is implemented. The text classification method based on artificial intelligence includes: obtaining a target text; inputting the target text into a preset text classification model for text classification prediction; obtaining a text classification result output by the text classification model as a target text classification result corresponding to the target text; wherein the text classification model is obtained through the following steps: using a plurality of obtained training samples, a sentiment word dictionary and a synonym dictionary to perform mask training on an initial model, wherein the initial model is a model obtained based on a Bert model and a classification prediction layer, and the sentiment word dictionary and the synonym dictionary are used to perform word replacement on an initial text sample in the training sample; and using the initial model after training as the text classification model.
[0170] This embodiment first obtains an initial model based on the Bert model and the classification prediction layer, then replaces words in the initial text sample in the training sample through the sentiment word dictionary and the synonym dictionary, and finally performs mask training on the initial model with the replaced training sample, so as to realize simultaneous mask training and text classification prediction training in the fine-tuning stage, effectively alleviate the difference between the two training stages of the Bert model, increase the robustness of the trained model, and make the trained model more suitable for specific application scenarios; by performing text classification prediction on the model obtained by simultaneously performing mask training and text classification prediction training in the fine-tuning stage, the accuracy of text classification prediction is improved.
[0171] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, an artificial intelligence-based text classification method is implemented, comprising the steps of: obtaining a target text; inputting the target text into a preset text classification model for text classification prediction; obtaining a text classification result output by the text classification model as a target text classification result corresponding to the target text; wherein the text classification model is obtained by the following steps: performing mask training on an initial model using a plurality of acquired training samples, a sentiment word dictionary, and a synonym dictionary, wherein the initial model is a model obtained based on a Bert model and a classification prediction layer, and the sentiment word dictionary and the synonym dictionary are used to perform word replacement on an initial text sample in the training sample; and using the initial model after training as the text classification model.
[0172] The above-mentioned artificial intelligence-based text classification method first obtains an initial model based on the Bert model and the classification prediction layer, then replaces the words of the initial text sample in the training sample through the sentiment word dictionary and the synonym dictionary, and finally performs mask training on the initial model with the replaced training sample, thereby realizing the simultaneous performance of mask training and text classification prediction training in the fine-tuning stage, effectively alleviating the difference between the two training stages of the Bert model, increasing the robustness of the trained model, and making the trained model more suitable for specific application scenarios; by performing text classification prediction on the model obtained by simultaneously performing mask training and text classification prediction training in the fine-tuning stage, the accuracy of text classification prediction is improved.
[0173] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided in this application and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0174] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the existence of other identical elements in the process, device, article or method including the element.
[0175] The above description is only a preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A text classification method based on artificial intelligence, characterized in that: The method comprises: Get the target text; Inputting the target text into a preset text classification model for text classification prediction; Obtaining a text classification result output by the text classification model as a target text classification result corresponding to the target text; The text classification model is obtained through the following steps: Using the acquired multiple training samples, sentiment word dictionary and synonym dictionary to perform mask training on the initial model, wherein the initial model is a model obtained based on the Bert model and the classification prediction layer, and the sentiment word dictionary and the synonym dictionary are used to replace words in the initial text samples in the training samples; Using the initial model after training as the text classification model; Before the step of inputting the target text into a preset text classification model for text classification prediction, the method further includes: Acquire a plurality of the training samples; Sequentially acquiring the training samples from each of the training samples as samples to be trained; Acquire words from the initial text sample of the sample to be trained according to a preset ratio to obtain a set of words to be replaced, and use the set of words to be replaced as a word calibration value; Using the sentiment word dictionary to determine whether each word in the to-be-replaced word set is a sentiment word, and obtaining a sentiment word set and a non-sentiment word set; According to a preset mask, the synonym dictionary, the to-be-replaced word set, the sentiment word set and the non-sentiment word set, the initial text sample of the to-be-trained sample is subjected to word replacement to obtain a target text sample; Input the target text sample into the initial model to perform word prediction and text classification prediction at the mask position, respectively, to obtain a word prediction value and a text classification prediction value; Training the initial model according to the text classification calibration value of the sample to be trained, the word calibration value, the word prediction value and the text classification prediction value; Repeat the step of determining the samples to be trained until a preset training target is reached; Using the initial model that achieves the training goal as the text classification model; The step of performing word replacement on the initial text sample of the sample to be trained according to the preset mask, the synonym dictionary, the set of words to be replaced, the emotional word set and the non-emotional word set to obtain a target text sample comprises: Using the mask symbol, replace each word corresponding to the non-emotional word set in the initial text sample of the sample to be trained, to obtain a text sample to be processed; Using the synonym dictionary, each word corresponding to the sentiment word set in the text sample to be processed is replaced to obtain the target text sample; The step of training the initial model according to the text classification calibration value of the sample to be trained, the word calibration value, the word prediction value and the text classification prediction value comprises: Inputting the text classification calibration value, the word calibration value, the word prediction value and the text classification prediction value of the sample to be trained into a preset target loss function to calculate the loss value, and obtain a target loss value; Updating the network parameters of the initial model according to the target loss value; The step of inputting the text classification calibration value, the word calibration value, the word prediction value and the text classification prediction value of the sample to be trained into a preset target loss function to calculate the loss value to obtain the target loss value includes: Inputting the word calibration value and the word prediction value of the sample to be trained into a preset word prediction loss function to calculate the loss value, thereby obtaining a first loss value; Inputting the text classification calibration value and the text classification prediction value of the sample to be trained into a preset text classification prediction loss function to calculate the loss value, thereby obtaining a second loss value; Performing a weighted summation of the first loss value and the second loss value to obtain the target loss value; Among them, the word prediction loss function and the text classification prediction loss function both adopt the cross entropy loss function, and the weighted sum of the word prediction loss and the text classification prediction loss is used as the target loss function, and the network parameters of the initial model are simultaneously updated according to the word prediction loss and the text classification prediction loss.
2. The text classification method based on artificial intelligence according to claim 1, characterized in that: Before the step of obtaining a plurality of training samples, the method further includes: Get multiple product review texts; Deleting blank characters and repeated punctuation marks from each product review text to obtain a preprocessed text; The training sample is generated according to each of the preprocessed texts.
3. The text classification method based on artificial intelligence according to claim 2, characterized in that: The step of generating the training sample according to each of the preprocessed texts comprises: Obtaining positive and negative classification prediction results corresponding to each of the preprocessed texts; The preprocessed text is used as an initial text sample of the training sample, and the positive and negative classification prediction results are used as text classification calibration values of the training sample.
4. A text classification device based on artificial intelligence, characterized in that: The device for implementing the artificial intelligence-based text classification method according to any one of claims 1 to 3 comprises: A text acquisition module is used to acquire the target text; A text classification module, used for inputting the target text into a preset text classification model for text classification prediction; A target text classification result determination module is used to obtain the text classification result output by the text classification model as the target text classification result corresponding to the target text; The model training module is used to perform mask training on the initial model using the acquired multiple training samples, a preset sentiment word dictionary and a preset synonym dictionary, and use the initial model after the training as the text classification model, wherein the initial model includes: a Bert model and a text classification layer, and the sentiment word dictionary and the synonym dictionary are used to replace words in the initial text samples in the training samples.
5. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 3 are implemented.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Text classification prediction method and device, equipment and storage medium
CN113326379A
Data enhancement method and system based on sentiment analysis
CN113505202A