Text classification algorithm optimization method based on machine learning

By combining deep learning and traditional machine learning methods, unsupervised and semi-supervised learning and data augmentation technology, the problem of limited generalization capabilities of existing text classification algorithms for new fields or new tasks is solved, and more robust and efficient text classification performance is achieved.

CN120067321APending Publication Date: 2025-05-30NORTHWEST UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510279544.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing text classification algorithm optimization method has limited generalization capabilities for new fields or new tasks, and requires re-collecting a large amount of labeled data for training.

Method used

Combining deep learning and traditional machine learning methods, using the feature extraction ability of deep learning and the stability of traditional machine learning, use a large amount of label-free data to be pre-trained or self-trained through unsupervised and semi-supervised learning, and data augmentation is carried out through methods such as generating adversarial networks.

Benefits of technology

Improve the robustness and efficiency of the model, enhance the generalization ability of new fields or new tasks, and reduce the dependence on labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067321A_ABST
    Figure CN120067321A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of text classification algorithm optimization methods, in particular to a text classification algorithm optimization method based on machine learning. According to the technical scheme, the text classification algorithm optimization method based on machine learning comprises a data collection and preprocessing method; compared with a traditional text classification algorithm optimization method, the method generally depends on feature engineering and supervised learning, needs a large amount of marked data, is limited in generalization ability for a new field or a new task and a model, needs to collect a large amount of marked data again for training, and is high in efficiency. According to the text classification algorithm optimization method, through combination of deep learning and a traditional machine learning method, the feature extraction capability of deep learning and the stability of traditional machine learning are utilized to obtain more robust and efficient classification performance, and meanwhile, pre-training or self-training is performed through a large amount of label-free data, so that the classification efficiency is improved. The generalization ability of the model can be greatly improved through unsupervised and semi-supervised learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text classification algorithm optimization methods, and particularly to text classification algorithm optimization methods based on machine learning. Background Art

[0002] A text classification algorithm optimization method refers to a method that, for a text classification task, improves and enhances a classification algorithm through a series of technical means and strategies to improve the accuracy and efficiency of classification.

[0003] Existing text classification algorithm optimization methods usually rely on feature engineering and supervised learning, require a large amount of labeled data, and have limited generalization ability for new fields or new tasks. It is necessary to re-collect a large amount of labeled data for training.

[0004] Aiming at the problem that existing text classification algorithm optimization methods have limited generalization ability for new fields or new tasks and need to re-collect a large amount of labeled data for training, this text classification algorithm optimization method combines deep learning and traditional machine learning methods, utilizes the feature extraction ability of deep learning and the stability of traditional machine learning to obtain more robust and efficient classification performance. At the same time, through pre-training or self-training with a large amount of unlabeled data, unsupervised and semi-supervised learning can greatly improve the generalization ability of the model. Summary of the Invention

[0005] In order to overcome the problem that existing text classification algorithm optimization methods usually rely on feature engineering and supervised learning, require a large amount of labeled data, and have limited generalization ability for new fields or new tasks and need to re-collect a large amount of labeled data for training.

[0006] The technical solution of the present invention is: a text classification algorithm optimization method based on machine learning, including the following steps: S101: First, obtain text data from a public dataset, an internal data source, or a web crawler, and collect and process the data; S102: Select and extract features in the data through a training model; S103: Use a generative adversarial network or other generative models to synthesize new text data to enhance and expand the data; S104: Select and initialize a model according to the characteristics of the dataset and the requirements of the classification task; S105: Adjust the hyperparameters of the model according to experience or experiments, and train and optimize the model; S106: Train multiple different text classification models, use the prediction results of multiple base models as new features, and train a meta-model; S107: Select a model pre-trained on a large-scale dataset related to the target task, and fine-tune the pre-trained model on the target dataset to adapt to the text classification task in a specific domain; S108: Divide the dataset into a training set, a validation set, and a test set, and evaluate and tune the model using cross-validation and hyperparameter tuning; S109: Calculate the influence degree of different features on the model output, and use a word cloud or a heat map to display the classification results of the model to improve the interpretability of the model.

[0007] Preferably, when performing data collection and preprocessing, the following steps are included: S201: Determine the sources of text data, including public datasets, company internal databases, and web crawlers, and select text data containing the target categories according to the requirements of the classification task; S202: Download the data to local storage and check the integrity and availability of the data; S203: Use a data deduplication algorithm to delete exactly the same text records, and identify and delete or correct errors, irrelevant information, or outliers in the text; S204: Unify the encoding format, case, and punctuation of the text. For text containing missing values, fill in the missing parts according to the context or statistical methods, or delete samples with too many missing values; S205: Select a suitable word segmentation tool according to the language characteristics, and segment the text into words or phrases to form a list of segmented text.

[0008] Preferably, when performing feature selection and extraction, the following steps are included: S301: Use methods such as chi-square test or mutual information to calculate the correlation score between each feature and the category. According to the correlation score, select the features with higher scores as the model input and discard irrelevant or redundant features; S302: Represent the text as a set of words, and count the frequency of each word in the text to form a bag-of-words vector; S303: Consider the word frequency and inverse document frequency, and calculate the TF-IDF value of each word as the feature representation of the text; S304: Use the Word2Vec or GloVe embedding model to map words to continuous vectors to capture the semantic relationships between words.

[0009] Preferably, when performing data augmentation and expansion, the following steps are included: S401: Organize the synonyms of common words to form a synonym dictionary; S402: Randomly select words in the text and replace them with synonyms in the synonym dictionary to increase the diversity of the samples; S403: Parse the grammatical structure of the sentence using natural language processing techniques; S404: According to the grammatical structure, change the word order or components of the sentence to generate a new sentence while keeping the original meaning unchanged; S405: Select a suitable generation model according to the text characteristics, a generative adversarial network; S406: Use existing text data to train the generation model so that it can generate new text samples; S407: Utilize the trained generation model to synthesize new text data and expand the dataset.

[0010] Preferably, when performing model selection and initialization, it includes the following steps: S501: According to the specific requirements of the text classification task, select a suitable model type, including Naive Bayes, Support Vector Machine, Random Forest or Convolutional Neural Network; S502: On the same dataset, use the same evaluation metrics to conduct benchmark tests on different models to evaluate the performance characteristics of different models, select the model with better performance as a candidate, and compare the performance of different models according to the benchmark test results to select the model with better performance as a candidate; S503: Select the model most relevant to the target task from existing pre-trained models, regularly pay attention to the latest pre-trained models, update the model library, so as to select a pre-trained model with better performance for the deep learning model, and select the model pre-trained on a large-scale dataset as the initial model; S504: Use the API provided by the deep learning framework to load the weights of the pre-trained model into the target model. Load the weights of the pre-trained model into the model as the initial weights of the model. After loading the weights, check whether the input and output layers of the model match the task requirements and make adjustments if necessary; S505: According to the task requirements and model characteristics, select appropriate hyperparameters. According to the task requirements and model characteristics, set the hyperparameters of the model, and use methods such as grid search, random search or Bayesian optimization to search for the best combination within the specified hyperparameter range.

[0011] Preferably, when performing model training and optimization, it includes the following steps: S601: Based on experience or experiments, determine the value range of each hyperparameter, and select a set of hyperparameters from the value range as the initial settings of the model. For the learning rate, it can be set between 0.0001 and 0.1; for the batch size, it can be set between 16 and 256, and select a set of hyperparameters from the value range as the initial settings of the model; S602: Understand the characteristics of different optimizers, select a suitable optimizer for model training according to the model characteristics and training requirements. For complex and deep neural network models, select the Adam optimizer for adaptive learning rate adjustment; S603: Add L1 or L2 regularization terms to the model loss function to prevent model overfitting. L1 regularization helps sparsify the model weights, while L2 regularization smooths the weights. At the same time, set the early stopping condition, monitor the performance of the validation set, and stop training when the performance of the validation set does not improve for multiple consecutive epochs to prevent overfitting; S604: During the training process, monitor in real-time metrics such as the loss value and accuracy on the training set and validation set, and use TensorBoard or other visualization tools to plot the training curves to promptly detect and handle problems such as overfitting or underfitting; S605: During the training process, regularly save the model weights and training status so that training can be quickly resumed in case of training interruption or when fine-tuning is needed. At the same time, save the model with the best performance for subsequent evaluation and deployment. Add L1 or L2 regularization terms to the model loss function to prevent model overfitting. Monitor the performance of the validation set during the training process and stop training when the performance no longer improves.

[0012] Preferably, when performing the ensemble learning method, the following steps are included: S701: Use different algorithms, including Naive Bayes, Support Vector Machine, Random Forest, or Convolutional Neural Network, or hyperparameter settings to train multiple text classification base models. Use different algorithms or hyperparameter settings to train multiple text classification models; S702: Combine the prediction results of multiple base models using the voting method or the averaging method. For classification tasks, use the majority voting method; for regression tasks, use the averaging method. In addition, assign different weights according to the performance of the base models to improve the accuracy of the ensemble model. Combine the prediction results of multiple models using the voting method or the averaging method to improve the classification accuracy; S703: Evaluate the performance of each base model through methods such as cross-validation, and select multiple models with better performance as the final base models. Select multiple models with better performance as base models and train them separately; S704: Use the prediction results of the base models as new features and train a meta-model for the final prediction; S705: Perform the same evaluation and tuning process on the ensemble model as on a single base model. Use methods such as cross-validation and hyperparameter tuning to evaluate the performance of the ensemble model, and adjust the parameter settings of the base models and the meta-model according to the evaluation results.

[0013] Preferably, when performing transfer learning and domain adaptation, the following steps are included: S801: Analyze the domain characteristics of the target task, and select a domain that is relevant to it and has sufficient data as the source domain. Select a domain that is relevant to the target task and has sufficient data as the source domain; S802: Evaluate multiple pre-trained models on the source domain dataset, and select a model with better performance as the starting point for transfer learning. Select a pre-trained model with better performance on the source domain as the starting point for transfer learning, and formulate different fine-tuning strategies according to the differences between the source domain and the target domain; S803: According to the similarity between the source domain and the target domain, select to freeze some layers or all layers of the pre-trained model. During the fine-tuning process, freeze some layers of the pre-trained model and keep their weights unchanged; S804: Use the data of the target domain to train the remaining layers of the pre-trained model, and adjust its weights to adapt to the target task. Train the remaining layers of the pre-trained model to adapt to the text classification task of the target domain. During the fine-tuning process, the learning rate can be appropriately reduced to prevent the model from overfitting; S805: According to the training effect, gradually unfreeze and train more layers until the model performance no longer improves significantly. According to the training effect, gradually unfreeze and train more layers until the best performance is achieved. During the fine-tuning process, monitor the performance of the model on the validation set and adjust the training strategy in a timely manner.

[0014] Preferably, when performing model evaluation and optimization, the following steps are included: S901: Randomly divide the dataset into a training set, a validation set, and a test set to ensure that the data distributions among the sets are consistent; S902: Determine the proportions of the training set, the validation set, and the test set according to the dataset size and task requirements; S903: Select appropriate evaluation metrics according to the task requirements, including accuracy, precision, recall, or F1 score; S904: Use the evaluation metrics to calculate the performance of the model on the validation set and the test set; S905: Determine the K value of cross-validation and divide the dataset into K subsets; S906: Sequentially use each subset as the validation set, and use the remaining subsets as the training set for training, and calculate the evaluation metrics; S907: Calculate the average value of the evaluation metrics of K-fold cross-validation as the stable performance estimate of the model; S908: Select tuning methods such as grid search or random search according to the number of hyperparameters and computing resources; S909: Search for the best combination within the specified hyperparameter range and evaluate the performance of the model corresponding to each set of hyperparameters.

[0015] Preferably, when performing model interpretation and visualization, the following steps are included: S1001: Calculate the influence degree of each feature on the model output using the built-in methods of the model or external tools; S1002: Sort the features according to the feature importance scores to identify the features that contribute more to the model performance; S1003: Select a suitable visualization tool according to the model type and interpretation requirements, and use the visualization tool to generate a classification decision path diagram or heat map of the model for a single text sample; S1004: Select a suitable display method according to the task requirements and audience characteristics, use a chart-making tool to produce the display content of the model classification results, and present the display content to the audience in the form of a report, presentation, or web page, etc.

[0016] Advantages of the present invention: 1. Compared with traditional text classification algorithm optimization methods, which usually rely on feature engineering and supervised learning, require a large amount of labeled data, and have limited generalization ability for new fields or new tasks, and need to re-collect a large amount of labeled data for training. This text classification algorithm optimization method combines deep learning and traditional machine learning methods, utilizes the feature extraction ability of deep learning and the stability of traditional machine learning to obtain more robust and efficient classification performance. At the same time, through pre-training or self-training with a large amount of unlabeled data, unsupervised and semi-supervised learning can greatly improve the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Shown is a schematic flow framework diagram of the text classification algorithm optimization method based on machine learning of the present invention; Figure 2 Shown is a schematic diagram of the data enhancement and expansion work flow of the text classification algorithm optimization method based on machine learning of the present invention; Figure 3 Shown is a schematic diagram of the model selection and initialization process of the text classification algorithm optimization method based on machine learning of the present invention; Figure 4 Shown is a schematic diagram of the model evaluation and tuning process of the text classification algorithm optimization method based on machine learning of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The present invention will be further described below in conjunction with the drawings and embodiments.

[0019] Please refer to Figure 1-2 , the present invention provides an embodiment: a text classification algorithm optimization method based on machine learning, including the following steps: S101: First, obtain text data from public data sets, internal data sources, or web crawlers, and collect and process the data; S102: Select and extract features in the data through a trained model; S103: Use a generative adversarial network or other generative model to synthesize new text data to enhance and expand the data; S104: Select and initialize a model according to the characteristics of the dataset and the requirements of the classification task; S105: Adjust the hyperparameters of the model based on experience or experiments, and train and optimize the model; S106: Train multiple different text classification models, use the prediction results of multiple base models as new features, and train a meta-model; S107: Select a model pre-trained on a large-scale dataset related to the target task, and fine-tune the pre-trained model on the target dataset to adapt to the text classification task in a specific domain; S108: Divide the dataset into a training set, a validation set, and a test set, and use cross-validation and hyperparameter tuning to evaluate and optimize the model; S109: Calculate the influence degree of different features on the model output, and use a word cloud or a heat map to display the classification results of the model to improve the interpretability of the model.

[0020] Preferably, when collecting and preprocessing data, the following steps are included: S201: Determine the source of the text data, including public datasets, company internal databases, and web crawlers. According to the requirements of the classification task, select text data containing the target categories; S202: Download the data to local storage and check the integrity and availability of the data; S203: Use a data deduplication algorithm to delete exactly the same text records, and identify and delete or correct errors, irrelevant information, or outliers in the text; S204: Unify the encoding format, case, and punctuation of the text. For text containing missing values, fill in the missing parts according to the context or statistical methods, or delete samples with too many missing values; S205: Select a suitable word segmentation tool according to the language characteristics, and segment the text into words or phrases to form a list of segmented text.

[0021] Preferably, when selecting and extracting features, the following steps are included: S301: Use methods such as chi-square test or mutual information to calculate the correlation score between each feature and the category. According to the correlation score, select the features with higher scores as the model input, and discard irrelevant or redundant features; S302: Represent the text as a set of words, count the frequency of each word in the text, and form a bag-of-words vector; S303: Consider the word frequency and inverse document frequency, calculate the TF-IDF value of each word, and use it as the feature representation of the text; S304: Use the Word2Vec or GloVe embedding model to map words into continuous vectors and capture the semantic relationships between words.

[0022] Preferably, when performing data augmentation and expansion, the following steps are included: S401: Organize the synonyms of common words to form a thesaurus; S402: Randomly select words in the text and replace them with synonyms in the thesaurus to increase the diversity of samples; S403: Use natural language processing techniques to parse the grammatical structure of sentences; S404: According to the grammatical structure, change the word order or components of the sentence to generate new sentences while keeping the original meaning unchanged; S405: Select a suitable generation model according to the text characteristics, such as a generative adversarial network; S406: Use the existing text data to train the generation model so that it can generate new text samples; S407: Use the trained generation model to synthesize new text data and expand the dataset.

[0023] Preferably, when performing model selection and initialization, the following steps are included: S501: According to the specific requirements of the text classification task, select a suitable model type, including Naive Bayes, Support Vector Machine, Random Forest, or Convolutional Neural Network; S502: On the same dataset, perform benchmark tests on different models using the same evaluation metrics. According to the benchmark test results, compare the performance of different models and select the model with better performance as a candidate; S503: Select the model most relevant to the target task from the existing pre-trained models, regularly pay attention to the latest pre-trained models, and update the model library to select a pre-trained model with better performance; S504: Use the API provided by the deep learning framework to load the weights of the pre-trained model into the target model. After loading the weights, check whether the input and output layers of the model match the task requirements and make adjustments if necessary; S505: According to the task requirements and model characteristics, select appropriate hyperparameters, and use methods such as grid search, random search, or Bayesian optimization to search for the best combination within the specified hyperparameter range. S502: Evaluate the performance characteristics of different models and select the model with better performance as a candidate; S503: For deep learning models, select the model pre-trained on a large-scale dataset as the initial model; S504: Load the weights of the pre-trained model into the model as the initial weights of the model; S505: Set the hyperparameters of the model according to the task requirements and model characteristics.

[0024] Preferably, when performing model training and optimization, the following steps are included: S601: According to experience or experiments, determine the value range of each hyperparameter, and select a set of hyperparameters from the value range as the initial settings of the model. For the learning rate, it can be set between 0.0001 and 0.1; for the batch size, it can be set between 16 and 256, and select a set of hyperparameters from the value range as the initial settings of the model; S602: Understand the characteristics of different optimizers, and select a suitable optimizer for model training according to the model characteristics and training requirements. For complex and deep neural network models, select the Adam optimizer for adaptive learning rate adjustment; S603: Add L1 or L2 regularization terms to the model loss function to prevent model overfitting. L1 regularization helps to sparsify the model weights, while L2 regularization can smooth the weights. At the same time, set the early stopping condition, monitor the performance of the validation set, and stop training when the performance of the validation set does not improve for multiple consecutive epochs to prevent overfitting; S604: During the training process, monitor the loss values, accuracies and other metrics on the training set and validation set in real time, and use TensorBoard or other visualization tools to draw the training curves to detect and handle problems such as overfitting or underfitting in a timely manner; S605: During the training process, save the model weights and training status regularly so that it can be quickly restored in case of training interruption or when fine-tuning is needed. At the same time, save the model with the best performance for subsequent evaluation and deployment. S601: According to experience or experiments, determine the value range of each hyperparameter, and select a set of hyperparameters from the value range as the initial settings of the model; S602: Understand the characteristics of different optimizers, and select a suitable optimizer for model training according to the model characteristics and training requirements; S603: Add L1 or L2 regularization terms to the model loss function to prevent model overfitting, and monitor the performance of the validation set during the training process, and stop training when the performance no longer improves.

[0025] Preferably, when performing the ensemble learning method, the following steps are included: S701: Use different algorithms, including Naive Bayes, Support Vector Machine, Random Forest, or Convolutional Neural Network, or hyperparameter settings to train multiple text classification base models; S702: Combine the prediction results of multiple base models using voting or averaging methods. For classification tasks, use the majority voting method; for regression tasks, use the averaging method. Additionally, assign different weights according to the performance of the base models to improve the accuracy of the ensemble model; S703: Evaluate the performance of each base model through methods such as cross-validation, and select multiple models with better performance as the final base models; S704: Use the prediction results of the base models as new features to train a meta-model for the final prediction; S705: Perform the same evaluation and tuning process on the ensemble model as on a single base model. Use methods such as cross-validation and hyperparameter tuning to evaluate the performance of the ensemble model, and adjust the parameter settings of the base model and the meta-model according to the evaluation results. S701: Train multiple text classification models using different algorithms or hyperparameter settings; S702: Combine the prediction results of multiple models using voting or averaging methods to improve the accuracy of classification; S703: Select multiple models with better performance as base models and train them separately; S704: Use the prediction results of the base models as new features to train a meta-model for the final prediction.

[0026] Preferably, when performing transfer learning and domain adaptation, the following steps are included: S801: Analyze the domain characteristics of the target task and select a relevant domain with sufficient data as the source domain; S802: Evaluate multiple pre-trained models on the source domain dataset, select a model with better performance as the starting point for transfer learning, and formulate different fine-tuning strategies according to the differences between the source domain and the target domain; S803: Select to freeze some layers or all layers of the pre-trained model according to the similarity between the source domain and the target domain; S804: Use the data of the target domain to train the remaining layers of the pre-trained model and adjust its weights to adapt to the target task. During the fine-tuning process, the learning rate can be appropriately reduced to prevent the model from overfitting; S805: According to the training effect, gradually unfreeze and train more layers until the model performance no longer improves significantly. During the fine-tuning process, monitor the performance of the model on the validation set and adjust the training strategy in a timely manner. S801: Select a domain relevant to the target task and with sufficient data as the source domain; S802: Select a pre-trained model with better performance on the source domain as the starting point for transfer learning; S803: During the fine-tuning process, freeze some layers of the pre-trained model and keep their weights unchanged; S804: Train the remaining layers of the pre-trained model to adapt to the text classification task in the target domain; S805: Gradually unfreeze and train more layers according to the training effect until the best performance is achieved.

[0027] Preferably, when performing model evaluation and tuning, the following steps are included: S901: Randomly divide the dataset into a training set, a validation set, and a test set to ensure consistent data distribution among the sets; S902: Determine the proportions of the training set, the validation set, and the test set according to the dataset size and task requirements; S903: Select appropriate evaluation metrics according to the task requirements, including accuracy, precision, recall, or F1 score; S904: Calculate the performance of the model on the validation set and the test set using the evaluation metrics; S905: Determine the K value for cross-validation and divide the dataset into K subsets; S906: Successively use each subset as the validation set and the remaining subsets as the training set for training, and calculate the evaluation metrics; S907: Calculate the average value of the evaluation metrics for K-fold cross-validation as an estimate of the stable performance of the model; S908: Select a tuning method such as grid search or random search according to the number of hyperparameters and computing resources; S909: Search for the best combination within the specified hyperparameter range and evaluate the performance of the model corresponding to each set of hyperparameters.

[0028] Preferably, when performing model interpretation and visualization, the following steps are included: S1001: Use the built-in method of the model or an external tool to calculate the degree of influence of each feature on the model output; S1002: Sort the features according to the feature importance scores to identify the features that contribute more to the model performance; S1003: Select an appropriate visualization tool according to the model type and interpretation requirements, and use the visualization tool to generate a classification decision path diagram or a heat map of the model for a single text sample; S1004: Select an appropriate display method according to the task requirements and the characteristics of the audience, use a chart-making tool to produce the display content of the model classification results, and present the display content to the audience in the form of a report, a presentation, or a web page.

[0029] Example 1 Project Background: An e-commerce platform needs to automatically classify user reviews in order to quickly understand users' feedback on products or services. The review content covers multiple aspects such as product quality, service attitude, and logistics speed. It is necessary to build an efficient text classification model that can accurately classify user reviews into predefined categories.

[0030] Implementation Method S1101: Export the user review data of the past year from the e-commerce platform database, including the review text and the corresponding category labels. The data volume is 100,000 pieces, covering multiple product and service categories; S1102: Use a Python script to remove duplicate reviews to ensure that each review is unique, process noisy data, and then standardize the text format. Convert all reviews to lowercase and unify punctuation marks; S1103: Use the jieba word segmentation tool to segment Chinese reviews and remove stop words such as "de", "le", "zai", etc. to reduce invalid features; S1104: Use the chi-square test to calculate the correlation score between each word and the category, and select the top 5000 words with higher scores as the model input features; S1105: Build a bag-of-words model, count the frequency of each word in the review, calculate the TF-IDF value as the feature representation of the text, and enhance the model's ability to recognize important words; S1106: Build a synonym dictionary for common words. For example, synonyms of "hao" include "youxiu", "bang", etc. Randomly select words in the review for synonym replacement to increase the diversity of the samples; S1107: Use natural language processing technology to parse the grammatical structure of the review, reorganize the sentence components, such as swapping the subject and the predicate to generate new sentences while keeping the original meaning unchanged; S1108: Select the generative adversarial network (GAN) as the data synthesis model, use the existing review data to train the GAN model so that it can generate new review samples that conform to grammatical and semantic rules, and synthesize 5000 new reviews to expand the dataset; S1109: According to the task requirements, select the convolutional neural network as the text classification model. CNN performs well in text classification tasks and can capture local features and reduce the number of parameters; S1110: Select a CNN model pre-trained on a large-scale Chinese corpus as the initial model, load the weights of the pre-trained model as the initial weights of the model, and set model parameters such as the learning rate of 0.001 and the batch size of 64; S1111: Determine the range of hyperparameters, such as a learning rate between 0.0001 and 0.01, a batch size between 32 and 128, and select a set of hyperparameters as the initial settings of the model, such as a learning rate of 0.001 and a batch size of 64; S1112: Select Adam optimizer for model training because it has the ability to adjust the learning rate adaptively and can accelerate the convergence speed; S1113: Add L2 regularization term to the model loss function to prevent overfitting of the model, set early stopping condition, and stop training when the accuracy of the validation set does not improve for 5 consecutive epochs; S1114: Train multiple CNN models, use different hyperparameter settings or data augmentation methods, and use voting to combine the prediction results of multiple models to improve classification accuracy; S1115: Select a CNN model with better performance as the base model, train them separately, use the prediction results of the base model as new features, and train a random forest meta-model for final prediction; S1116: Randomly divide the data set into a training set (70%), a validation set (15%), and a test set (15%) to ensure that the data distribution between the sets is consistent and avoid data leakage; S1117: Select accuracy, precision, recall and F1 score as evaluation indicators, calculate the performance of the model on the validation set and test set, and comprehensively evaluate the model effect; S1118: Select 5-fold cross validation, divide the data set into 5 subsets, use each subset as a validation set in turn, and train the remaining subsets as training sets, and calculate the evaluation index, and calculate the average value of the evaluation index of the 5 cross validations as the stable performance estimate of the model; S1119: Use the grid search method to search for the best combination within the specified hyperparameter range, evaluate the model performance corresponding to each set of hyperparameters, and select the hyperparameter combination with the best performance as the final setting of the model; S1120: Use the SHAP tool to calculate the degree of influence of each feature on the model output, sort the features, and identify the features that contribute more to the model performance, such as words such as "quality", "service", and "speed"; S1121: Use the LIME tool to generate a classification decision path diagram of the model for a single review sample, and analyze how the model makes classification decisions based on features through visualization results to improve the interpretability of the model; S1122: Prepare a classification result presentation report to display key indicators such as the number of review samples in each category and classification accuracy, and present the presentation report in PDF format to project team members and e-commerce platform management to help them understand and evaluate the performance of the model.

[0031] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those skilled in the art.

Claims

1. A text classification algorithm optimization method based on machine learning; characterized in that: The following steps are included: S101: First, obtain text data from public data sets, internal data sources or web crawlers, and collect and process the data; S102: Select and extract features from the data through the training model; S103: Use a generative adversarial network or other generative models to synthesize new text data to enhance and expand the data; S104: Select and initialize the model according to the characteristics of the data set and the requirements of the classification task; S105: Adjust the hyperparameters of the model based on experience or experiments, and train and optimize the model; S106: training multiple different text classification models, using the prediction results of multiple base models as new features, and training a meta-model; S107: Select a model that is relevant to the target task and pre-trained on a large-scale dataset, and fine-tune the pre-trained model on the target dataset to adapt it to the text classification task in a specific field; S108: Divide the data set into training set, validation set and test set, and use cross-validation and hyperparameter tuning to evaluate and tune the model; S109: Calculate the influence of different features on the model output, use word cloud diagrams or heat maps to display the classification results of the model, and improve the interpretability of the model.

2. The text classification algorithm optimization method based on machine learning according to claim 1, characterized in that: The following steps are included in data collection and preprocessing: S201: Determine the source of text data, including public data sets, internal company databases, and web crawlers, and select text data containing the target category according to the requirements of the classification task; S202: Download the data to local storage and check the integrity and availability of the data; S203: using a data deduplication algorithm to delete identical text records, and identifying and deleting or correcting errors, irrelevant information or outliers in the text; S204: Unify the encoding format, capitalization, and punctuation of the text. For texts with missing values, fill in the missing parts based on the context or statistical methods, or delete samples with too many missing values; S205: Select an appropriate word segmentation tool according to the language characteristics, divide the text into words or phrases, and form a text list after word segmentation.

3. The text classification algorithm optimization method based on machine learning according to claim 1, characterized in that: When selecting and extracting features, the following steps are included: S301: Use methods such as chi-square test or mutual information to calculate the correlation score between each feature and the category. According to the correlation score, select the features with higher scores as model inputs and discard irrelevant or redundant features; S302: Represent the text as a collection of words, count the frequency of each word in the text, and form a bag-of-words vector; S303: Considering the word frequency and inverse document frequency, the TF-IDF value of each word is calculated as the feature representation of the text; S304: Use the Word2Vec or GloVe embedding model to map words into continuous vectors to capture the semantic relationship between words.

4. The text classification algorithm optimization method based on machine learning according to claim 1, characterized in that: When performing data enhancement and expansion, the following steps are included: S401: Arrange the synonyms of common words to form a synonym dictionary; S402: randomly select words in the text and replace them with synonyms in the synonym dictionary to increase the diversity of the sample; S403: Analyze the grammatical structure of the sentence using natural language processing technology; S404: according to the grammatical structure, the word order or components of the sentence are changed to generate a new sentence while keeping the original meaning unchanged; S405: Select an appropriate generation model based on the characteristics of the text to generate a confrontation network; S406: Use existing text data to train the generative model so that it can generate new text samples; S407: Use the trained generative model to synthesize new text data and expand the data set.

5. The text classification algorithm optimization method based on machine learning according to claim 1, characterized in that: When performing model selection and initialization, the following steps are included: S501: According to the specific requirements of the text classification task, select a suitable model type, including naive Bayes, support vector machine, random forest or convolutional neural network; S502: Benchmark different models using the same evaluation index on the same data set, compare the performance of different models based on the benchmark test results, and select the model with better performance as a candidate; S503: Select the model most relevant to the target task from the existing pre-trained models, regularly pay attention to the latest pre-trained models, and update the model library in order to select a pre-trained model with better performance; S504: Use the API provided by the deep learning framework to load the weights of the pre-trained model into the target model. After loading the weights, check whether the input and output layers of the model match the task requirements and make adjustments if necessary; S505: Select appropriate hyperparameters according to task requirements and model characteristics, and use grid search, random search, or Bayesian optimization to search for the best combination within the specified hyperparameter range. S502: Evaluate the performance characteristics of different models and select the model with better performance as a candidate; S503: For the deep learning model, select a model pre-trained on a large-scale dataset as the initial model; S504: Load the weights of the pre-trained model into the model as the initial weights of the model; S505: Set the model's hyperparameters according to task requirements and model characteristics.

6. The text classification algorithm optimization method based on machine learning according to claim 1, characterized in that: When training and optimizing the model, the following steps are included: S601: According to experience or experiments, determine the value range of each hyperparameter, select a set of hyperparameters from the value range as the initial settings of the model, the learning rate can be set between 0.0001 and 0.1; the batch size can be set between 16 and 256, and select a set of hyperparameters from the value range as the initial settings of the model; S602: Understand the characteristics of different optimizers, select a suitable optimizer for model training according to model characteristics and training requirements, and select the Adam optimizer for adaptive learning rate adjustment for complex and deep neural network models; S603: Add L1 or L2 regularization terms to the model loss function to prevent the model from overfitting. L1 regularization helps to sparse the model weights, while L2 regularization can smooth the weights. At the same time, set early stopping conditions and monitor the performance of the validation set. When the performance of the validation set does not improve for multiple consecutive epochs, stop training to prevent overfitting; S604: During the training process, real-time monitoring of indicators such as loss value and accuracy on the training set and validation set is performed, and the training curve is drawn using TensorBoard or other visualization tools to timely discover and handle problems such as overfitting or underfitting; S605: During the training process, the model weights and training status are saved regularly so that they can be quickly restored when the training is interrupted or needs fine-tuning, and the best performance model is saved for subsequent evaluation and deployment. S601: Based on experience or experiments, the value range of each hyperparameter is determined, and a set of hyperparameters is selected from the value range as the initial settings of the model; S602: Understand the characteristics of different optimizers, and select a suitable optimizer for model training according to model characteristics and training requirements; S603: Add L1 or L2 regularization terms to the model loss function to prevent the model from overfitting, monitor the performance of the validation set during training, and stop training when the performance no longer improves.

7. The text classification algorithm optimization method based on machine learning according to claim 1, characterized in that: When performing an ensemble learning method, the following steps are included: S701: Train multiple text classification base models using different algorithms including Naive Bayes, Support Vector Machine, Random Forest or Convolutional Neural Network or hyperparameter settings; S702: Use voting method or averaging method to combine the prediction results of multiple base models. For classification tasks, majority voting method is used; for regression tasks, averaging method is used. In addition, different weights are assigned according to the performance of the base models to improve the accuracy of the integrated model; S703: Evaluate the performance of each base model through cross-validation and other methods, and select multiple models with better performance as the final base models; S704: Using the prediction results of the base model as new features, training a meta-model for final prediction; S705: Perform the same evaluation and tuning process on the integrated model as on the single base model, use methods such as cross-validation and hyperparameter tuning to evaluate the performance of the integrated model, and adjust the parameter settings of the base model and meta-model according to the evaluation results. S701: Train multiple text classification models using different algorithms or hyperparameter settings; S702: Use voting method or averaging method to combine the prediction results of multiple models to improve the accuracy of classification; S703: Select multiple models with better performance as base models and train them separately; S704: Use the prediction results of the base model as new features and train a meta-model for final prediction.

8. The text classification algorithm optimization method based on machine learning according to claim 1, characterized in that: When performing transfer learning and domain adaptation, the following steps are included: S801: Analyze the domain characteristics of the target task and select a domain that is relevant and has sufficient data as the source domain; S802: Evaluate multiple pre-trained models on the source domain dataset, select the model with better performance as the starting point of transfer learning, and formulate different fine-tuning strategies according to the differences between the source domain and the target domain; S803: Select to freeze some or all layers of the pre-trained model according to the similarity between the source domain and the target domain; S804: Use the data from the target domain to train the remaining layers of the pre-trained model, and adjust their weights to adapt to the target task. During the fine-tuning process, the learning rate can be appropriately reduced to prevent the model from overfitting; S805: According to the training results, gradually unfreeze and train more layers until the model performance is no longer significantly improved. During the fine-tuning process, monitor the performance of the model on the validation set and adjust the training strategy in time. S801: Select a field that is relevant to the target task and has sufficient data as the source field; S802: Select a pre-trained model with better performance in the source domain as the starting point for transfer learning; S803: During the fine-tuning process, freeze some layers of the pre-trained model to keep their weights unchanged; S804: training the remaining layers of the pre-trained model to adapt it to the text classification task in the target domain; S805: According to the training effect, gradually unfreeze and train more layers until the best performance is achieved.

9. The text classification algorithm optimization method based on machine learning according to claim 1, characterized in that: When evaluating and tuning the model, the following steps are included: S901: Randomly divide the data set into training set, validation set and test set to ensure that the data distribution between the sets is consistent; S902: Determine the ratio of training set, validation set and test set according to the size of the data set and task requirements; S903: Select appropriate evaluation indicators according to task requirements, including accuracy, precision, recall or F1 score; S904: Use the evaluation index to calculate the performance of the model on the validation set and the test set; S905: Determine the K value for cross-validation and divide the data set into K subsets; S906: Take each subset as a validation set in turn, and train the remaining subsets as training sets, and calculate the evaluation index; S907: Calculate the average value of the evaluation index of K cross-validation as the stability performance estimate of the model; S908: Select a tuning method such as grid search or random search based on the number of hyperparameters and computing resources; S909: Search for the best combination within the specified hyperparameter range and evaluate the model performance corresponding to each set of hyperparameters.

10. The text classification algorithm optimization method based on machine learning according to claim 1, characterized in that: When interpreting and visualizing a model, the following steps are involved: S1001: Use the model's built-in methods or external tools to calculate the degree of influence of each feature on the model output; S1002: Sort the features according to their importance scores and identify the features that contribute most to the model performance; S1003: Select an appropriate visualization tool based on the model type and interpretation requirements, and use the visualization tool to generate a classification decision path diagram or heat map of the model for a single text sample; S1004: Select an appropriate presentation method based on task requirements and audience characteristics, use chart making tools to create presentation content of the model classification results, and present the presentation content to the audience in the form of a report, presentation, or web page.

Citation Information

Cited By

  • Method and system for automatically judging abnormal type of data link in power data

    CN120780532A

  • A method and system for automatically identifying power data platform data link exception types

    CN120780532B