Text Prompt Classification Method Based on Improved BERT-WWM

By improving the BERT-WWM model, introducing sparse attention and regularization technology, combining LAMB optimizer and meta-learning model to optimize hyperparameters, the classification accuracy and inefficiency in power system engineering document processing is solved, and high accuracy and efficient text classification is achieved.

CN119377413BActive Publication Date: 2025-05-30NANCHANG KECHEN ELECTRIC POWER TEST & RES CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411979450.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-30
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

When processing power system engineering documents, the prior art has problems such as limited text processing capabilities, insufficient model generalization capabilities, low parameter optimization efficiency and lack of intelligence in hyperparameter adjustment, resulting in low classification accuracy and low processing efficiency.

Method used

The text telephony classification method based on improved BERT-WWM is adopted to improve the generalization ability and processing efficiency of the model by introducing a sparse attention mechanism, adaptively adjusting the mask ratio in sparse attention, combining Dropout regularization and DropConnect regularization, and using LAMB optimizer and LSTM-based meta-learning model.

Benefits of technology

It significantly improves the classification accuracy and processing efficiency of power system engineering texts, can effectively identify and classify complex professional terms and long text information, and the classification accuracy reaches 95%, and reduces the overfitting phenomenon and training convergence time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119377413B_ABST
    Figure CN119377413B_ABST
Patent Text Reader

Abstract

The present invention discloses a text prompting classification method based on an improved BERT-WWM, which includes the following steps: collecting power infrastructure text data; constructing a BERT-WWM model, and introducing sparse attention to replace the multi-head self-attention in the BERT-WWM model, and at the same time adaptively adjusting the mask ratio in the sparse attention to obtain an improved BERT-WWM model; training the improved BERT-WWM model with a data set, and at the same time optimizing the parameters of the improved BERT-WWM model to obtain an optimized improved BERT-WWM model; inputting the power infrastructure text data to be classified into the optimized improved BERT-WWM model to obtain a classification result; through the uniquely designed optimized improved BERT-WWM model, the present invention significantly improves the classification accuracy and processing efficiency of power system engineering texts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text classification, and specifically to a text cue classification method based on an improved BERT-WWM. Background Art

[0002] With the increasing growth of power system engineering construction, project management plays a crucial role in ensuring the smooth implementation of projects, improving project quality, and controlling costs. The management of power system engineering information covers detailed information on numerous devices and facilities from power generation stations to substations, such as the technical parameters of transformers, line layouts, equipment capacities, etc. This information is usually saved in Word documents, containing a large amount of complex professional data and descriptive texts. In current technical means, project managers generally extract and manage information through manual methods, general text extraction tools, or basic regular expressions. Although these methods have a certain effect in preliminary data processing, there are still significant deficiencies.

[0003] From a large background perspective, traditional engineering information management methods are difficult to achieve the intelligence, rapidity, and efficiency of information, which poses a huge challenge to power engineering project management. In the case of a large amount of information and frequent data processing requirements, the power system engineering requires more efficient technical means to improve the accuracy of data extraction and management, reduce the uncertainty brought by human factors, and at the same time have flexible expansion capabilities to adapt to new technologies and new requirements; from a small background perspective, the current specific methods for power system engineering document processing still mainly rely on the following technologies: Manual entry method: Staff manually enter key information into Excel or a database by reading Word documents. This method has high requirements for personnel, the data entry work is heavy and inefficient, and human operations are often prone to errors. General text extraction tools: Common OCR or PDF conversion software can convert documents into text, but these general tools lack the ability to understand power system professional terms, cannot accurately identify or process data containing professional vocabulary, and are relatively sensitive to format changes, and the extraction results need secondary cleaning and sorting. Regular expression matching: Using regular expressions to identify and extract data in a specific format in the document, this method is effective when dealing with relatively simple text structures, but for complex formats and unstructured text content, the recognition accuracy is low and it is difficult to meet the requirements of power engineering information extraction.

[0004] Defects of the current existing technologies:

[0005] 1. Limited text processing ability: The existing text processing methods mainly rely on simple regular expression matching or basic text extraction tools, and cannot effectively handle long-distance text dependencies. Especially when dealing with power system engineering documents containing complex professional terms, problems such as inaccurate extraction or omission of important information often occur.

[0006] 2. Insufficient model generalization ability: Traditional text classification models are prone to overfitting when dealing with new samples. Especially in scenarios with strong professionalism such as the power system field, the model often has difficulty accurately understanding and classifying unseen professional terms and expressions.

[0007] 3. Inefficient parameter optimization: In the process of model training with existing methods, the parameter optimization strategy is relatively simple. Fixed learning rates and unified parameter update methods are often used, making it difficult to adaptively adjust according to the parameter characteristics of different levels, resulting in slow model convergence speed and unstable training effects.

[0008] 4. Lack of intelligence in hyperparameter tuning: Traditional hyperparameter tuning methods mainly rely on manual experience or simple grid search. This method not only consumes time and effort but also is difficult to find the optimal parameter combination. Especially when dealing with large-scale power system text data, repeated attempts are often required to obtain better results. Summary of the Invention

[0009] In view of the deficiencies of the prior art, the present invention provides a text keyword classification method based on improved BERT-WWM, aiming to solve the problems mentioned in the background art.

[0010] To achieve the above object, the present invention provides the following technical solutions: A text keyword classification method based on improved BERT-WWM, including the following steps:

[0011] Step S1: Collect power infrastructure text data for preprocessing and construct a dataset. The power infrastructure text data consists of sentences of different lengths;

[0012] Step S2: Construct a BERT-WWM model, and introduce sparse attention to replace the multi-head self-attention in the BERT-WWM model. At the same time, adaptively adjust the mask ratio in the sparse attention to obtain an improved BERT-WWM model;

[0013] Step S3: Use the dataset to train the improved BERT-WWM model, and at the same time optimize the parameters of the improved BERT-WWM model to obtain an optimized improved BERT-WWM model;

[0014] Step S4: Input the power infrastructure text data to be classified into the optimized improved BERT-WWM model to obtain the classification result;

[0015] The specific process of optimizing the parameters of the improved BERT-WWM model is as follows:

[0016] Combine Dropout regularization and DropConnect regularization to regularize the parameters of the improved BERT-WWM model;

[0017] Use the LAMB optimizer as the optimizer for the improved BERT-WWM model;

[0018] Use an LSTM-based meta-learning model to optimize the hyperparameters during the training of the improved BERT-WWM model.

[0019] Furthermore, adaptively adjust the mask ratio in sparse attention, expressed as:

[0020] ;

[0021] In the formula, represents the mask ratio when the sentence length in the power infrastructure text data is ; and are adjustable parameters.

[0022] Furthermore, the specific process of training the improved BERT-WWM model using the dataset is as follows: Divide the dataset into a training set and a validation set according to a ratio, use the training set to train the improved BERT-WWM model, and use the validation set to validate the improved BERT-WWM model.

[0023] Furthermore, the specific process of combining Dropout regularization and DropConnect regularization to regularize the parameters of the improved BERT-WWM model is as follows:

[0024] Through the weight factor weight the dropout probabilities of Dropout regularization and DropConnect regularization to obtain the joint regularization probability , expressed as:

[0025] ;

[0026] In the formula, and respectively represent the dropout probabilities of Dropout regularization and DropConnect regularization;

[0027] Dynamically adjust the dropout probability according to the gradient information of each layer of the improved BERT-WWM model, expressed as:

[0028] ;

[0029] In the formula, is the base dropout probability, To improve the gradient of the weights of the th layer of the BERT-WWM model, is the maximum value of the gradients of all layers of the BERT-WWM model; represents the dropout probability of the Dropout regularization of the th layer of the BERT-WWM model;

[0030] Calculate the regularization strength, expressed as:

[0031] ;

[0032] In the formula, represents the regularization strength at the th layer of the BERT-WWM model and the training epoch ; is the total number of layers of the BERT-WWM model; is the total number of training epochs; is a preset initial regularization strength; is a hyperparameter that controls the decay rate; represents the current training epoch;

[0033] Design the gradient update rule of the BERT-WWM model; in the case of joint regularization probability, the gradient update formula of the BERT-WWM model is:

[0034] ;

[0035] In the formula, is the weight of the th layer of the BERT-WWM model; is the learning rate; is the loss function with respect to the partial derivative of the weight of the th layer; and are the mask matrices of Dropout regularization and DropConnect regularization respectively;

[0036] Based on the performance of the BERT-WWM model on the validation set, dynamically adjust and :

[0037] ;

[0038] ;

[0039] In the formula, represents the change in the validation set loss; is the adjustment coefficient; is the loss of the current validation set; is at the time of the th training; is the at the time of the th training; is at the time of the th training; is at the time of the th training.

[0040] Furthermore, the specific process of optimizing the hyperparameters during the training of the improved BERT-WWM model using the LSTM-based meta-learning model is as follows:

[0041] Set the rd training task of the improved BERT-WWM model on the dataset as , and record the hyperparameters of and the performance data of the improved BERT-WWM model after completing the rd training task ;

[0042] Learn the hyperparameters and performance data of the historical training tasks, and generate the hyperparameter selection for the current training task, expressed as:

[0043] ;

[0044] ;

[0045] In the formula, represents the hyperparameter selection for the th training task generated by the network; represents the hyperparameter selection for the th training task generated by the network;

[0046] The process of the

[0047] network generating the hyperparameter selection is expressed as:

[0048] ;

[0049] In the formula, represents the hyperparameter selection strategy; and are the weight matrices of the hyperparameters and performance data, and are the activation functions of the LSTM network; denote the weight matrix in the network; and is the bias term;

[0050] According to , train the improved BERT-WWM model in the current training task, and introduce an adaptive mechanism to adjust the hyperparameters according to the performance data of the improved BERT-WWM model.

[0051] Furthermore, according to , the specific process of training the improved BERT-WWM model in the current training task and introducing an adaptive mechanism to adjust the hyperparameters according to the performance data of the improved BERT-WWM model is as follows:

[0052] Minimize the loss function of the training task by adjusting the hyperparameters , that is, select a set of optimal hyperparameter configurations in each training task to achieve the optimal performance of the improved BERT-WWM model; the loss function of the training task is expressed as:

[0053] ;

[0054] In the formula, denotes the output of the improved BERT-WWM model at ; is the loss corresponding to the training task ;

[0055] Optimize the loss functions of all training tasks , and find the optimal hyperparameter configuration, which is expressed as:

[0056] ;

[0057] In the formula, is the total number of training tasks;

[0058] Introduce the real-time feedback information of the training task, that is, the accuracy change or loss change of the improved BERT-WWM model, and dynamically adjust the selection of hyperparameters, which is expressed as:

[0059] ;

[0060] In the formula, is the current learning rate; is the adjustment factor; is the change amount of; is the current loss value;

[0061] Minimize the sum of the loss functions of all training tasks to obtain the hyperparameters of the improved BERT-WWM model optimized by the LSTM-based meta-learning model , which is expressed as:

[0062] ;

[0063] In the formula, represents minimizing the hyperparameters of the improved BERT-WWM model .

[0064] Furthermore, the LAMB optimizer is used as the optimizer for the improved BERT-WWM model, which is expressed as:

[0065] ;

[0066] In the formula, and respectively represent the estimated values of the mean and variance of the gradients of the improved BERT-WWM model; is a constant; represents hierarchical correction.

[0067] Furthermore, collect power infrastructure text data for preprocessing, and the preprocessing includes:

[0068] Text cleaning, using regular expressions to clean the power infrastructure text data;

[0069] Information extraction, extracting the key information of the power infrastructure text data;

[0070] Text vectorization, vectorizing the power infrastructure text data after text cleaning.

[0071] Compared with the existing technologies, the present invention has the following beneficial effects:

[0072] (1) Through the uniquely designed optimized improved BERT-WWM model, the present invention significantly improves the classification accuracy and processing efficiency of power system engineering texts. The test results show that the classification accuracy of this solution on the standard test set reaches 95%, and the processing speed is also significantly improved compared with the original BERT model, and it can effectively identify and classify complex professional terms and long text information.

[0073] (2) By introducing the sparse attention mechanism into the BERT-WWM model and adaptively adjusting the mask ratio in the sparse attention, the present invention effectively reduces the computational complexity, greatly improves the performance of the model when processing long texts, and at the same time enhances the ability to capture long-distance dependencies; by adopting joint regularization and the LAMB optimizer to optimize and improve the BERT-WWM model, the generalization ability of the model is significantly enhanced, reducing the overfitting phenomenon of the model, improving the training convergence speed by about, and making the training process more stable; by using the LSTM-based meta-learning model to optimize the hyperparameters during the training of the BERT-WWM model, the hyperparameter search time is greatly shortened, while ensuring the stability of the model performance.

[0074] (3) Compared with the prior art, the present invention breaks through the limitation of traditional text classification methods relying on simple feature extraction, and realizes the deep semantic understanding of power system professional texts through the designed and optimized improved BERT-WWM model; in practical applications, the present invention can automatically adapt to different types of engineering documents and complete high-quality classification tasks without manual intervention. Brief Description of the Drawings

[0075] Figure 1 It is the flowchart of the method of the present invention. Detailed Embodiments

[0076] As Figure 1 shown, the present invention provides a technical solution: a text keyword classification method based on an improved BERT-WWM, including the following steps:

[0077] Step S1: Collect power infrastructure text data for preprocessing and construct a data set. The power infrastructure text data consists of sentences of different lengths.

[0078] The collected power infrastructure text data may come from reports, documents or other records, and the content contains a large amount of unstructured information, so preprocessing operations are required. The preprocessing includes:

[0079] Text cleaning, using regular expressions to clean the power infrastructure text data, removing useless characters, spaces, special symbols, etc. in the power infrastructure text data to improve the data quality.

[0080] Information extraction, extracting the key information of the power infrastructure text data to ensure that the data retains meaningful content.

[0081] Text vectorization, vectorizing the power infrastructure text data after text cleaning. For example, the text is converted into processable numerical data through methods such as word embedding (such as Word2Vec, TF-IDF or BERT vectorization).

[0082] Step S2: Construct a BERT-WWM model, and introduce Sparse Attention to replace the multi-head self-attention in the BERT-WWM model. At the same time, adaptively adjust the mask ratio in the Sparse Attention to obtain an improved BERT-WWM model.

[0083] Among them, Sparse Attention reduces the computational complexity by restricting the number of matches between queries and keys, and is expressed as:

[0084] (1);

[0085] In the formula, represent queries, keys, and values respectively; represents the dimension of the key vector; represents the transpose of; represents the mask matrix, which is used to retain the calculations within the local window; represents the element-wise product; represents function; represents Sparse Attention.

[0086] Among them, adaptively adjusting the mask ratio in the Sparse Attention is expressed as:

[0087] (2);

[0088] In the formula, represents the mask ratio when the sentence length in the text data of the power infrastructure is ; and are adjustable parameters used for adaptive adjustment of the sentence length; in this way, the improved BERT-WWM model can more flexibly adjust the mask ratio according to the sentence length or complexity, so as to better capture the semantic relationships in the sentence.

[0089] Step S3: Use the dataset to train the improved BERT-WWM model, and at the same time optimize the parameters of the improved BERT-WWM model to obtain an optimized improved BERT-WWM model;

[0090] Step S3.1: Divide the dataset into a training set and a validation set according to a ratio. Use the training set to train the improved BERT-WWM model, and use the validation set to validate the improved BERT-WWM model.

[0091] Step S3.2: Combine Dropout regularization and DropConnect regularization to regularize the parameters of the improved BERT-WWM model.

[0092] Step S3.21: Joint probability initialization;

[0093] To simultaneously utilize the advantages of Dropout regularization and DropConnect regularization, the dropout probabilities of Dropout regularization and DropConnect regularization are weighted and combined through a weight factor to obtain a joint regularization probability , expressed as:

[0094] (3);

[0095] In the formula, and respectively represent the dropout probabilities of Dropout regularization and DropConnect regularization.

[0096] Step S3.22: Adaptive probability adjustment;

[0097] To better control the dropout probability of each layer of the improved BERT-WWM model, the dropout probability is dynamically adjusted according to the gradient information of each layer of the improved BERT-WWM model, expressed as:

[0098] (4);

[0099] In the formula, is the base dropout probability, is the gradient of the weights of the th layer of the improved BERT-WWM model, is the maximum value of the gradients of all layers of the improved BERT-WWM model; represents the dropout probability of Dropout regularization for the th layer of the improved BERT-WWM model.

[0100] Step S3.23: Regularization strength calculation;

[0101] The regularization strength is dynamically adjusted according to the depth of the layers of the improved BERT-WWM model and the current training epoch, expressed as:

[0102] (5);

[0103] In the formula, represents the regularization strength at the th layer of the improved BERT-WWM model and training epoch . The regularization strength controls the influence degree of the regularization term in the improved BERT-WWM model and is usually used to prevent overfitting; is the total number of layers of the improved BERT-WWM model; is the total number of training rounds; is a preset initial regularization strength, which is used to control the basic value of the regularization strength and determines the degree of regularization for improving the BERT-WWM model at the beginning; is a hyperparameter that controls the attenuation rate, which determines how the regularization strength decreases as the number of layers of the improved BERT-WWM model increases; represents the current training round.

[0104] Step S3.24: Design the gradient update rule for the improved BERT-WWM model;

[0105] In the case of joint regularization probability, the gradient update formula for the improved BERT-WWM model is:

[0106] (6);

[0107] In the formula, is the weight of the th layer of the improved BERT-WWM model; is the learning rate; is the loss function with respect to the partial derivative of the weight of the th layer, representing the direction of weight update; and are the mask matrices for Dropout regularization and DropConnect regularization, respectively.

[0108] Step S3.25: Adaptive adjustment mechanism;

[0109] Based on the performance of the improved BERT-WWM model on the validation set, dynamically adjust and :

[0110] (7);

[0111] (8);

[0112] In the formula, represents the change in the validation set loss; is the adjustment coefficient; is the current loss of the validation set; is at the th training; is the at the th training; is at the th training; is During the next training 。

[0113] Step S3.26: Regularization effect evaluation;

[0114] Finally, calculate the output of the improved BERT-WWM model after regularization:

[0115] (9);

[0116] In the formula, is the activation function, represents the element-wise multiplication operation, is the output of the th layer of the improved BERT-WWM model; is the output of the th layer of the improved BERT-WWM model.

[0117] Combining Dropout regularization and DropConnect regularization can further enhance the robustness of the model. Dropout regularization is used to discard neurons, while DropConnect regularization is used to discard weight connections. Through this combination strategy, the sparsity of the network can be increased, thus enabling the model to have better generalization ability. This combination usually performs well in large-scale deep neural networks, especially pre-trained models like BERT.

[0118] Step S3.3: In large-scale deep learning models, especially in large-scale pre-trained models like BERT, conventional optimizers such as Adam or SGD are prone to problems of gradient explosion or gradient vanishing during training, especially when the parameter updates of deep neural networks are uneven; adopt the LAMB optimizer (Layer-wise Adaptive Moments optimizer for BERT) as the optimizer for the improved BERT-WWM model; the LAMB optimizer effectively solves this problem by introducing a hierarchical adaptive adjustment strategy to adaptively adjust the learning rate of each layer according to its characteristics during the training of each layer.

[0119] The LAMB optimizer is expressed as:

[0120] (10);

[0121] In the formula, represents the hyperparameters of the improved BERT-WWM model, usually the weights or biases of the improved BERT-WWM model; and respectively represent the estimated values of the mean and variance of the gradient of the improved BERT-WWM model, is a small constant to prevent division by zero errors; represents layer-level correction; the LAMB optimizer is a specific adjustment for each layer, making the updates for different layers more reasonable.

[0122] Step S3.4: Traditional hyperparameter optimization methods such as grid search and random search usually have huge computational overheads and are not efficient enough when faced with complex high-dimensional hyperparameter spaces. A meta-learning model based on LSTM (Meta-Learner LSTM) is used to optimize the hyperparameters during the training of the improved BERT-WWM model. Specifically:

[0123] Step S3.41: Historical task data collection: Set the th training task of the improved BERT-WWM model on the dataset as , and record the hyperparameters and the performance data of the improved BERT-WWM model after completing the th training task .

[0124] Step S3.42: Generate hyperparameter adjustment strategies: Learn the hyperparameters and performance data of historical training tasks to generate the hyperparameter selection for the current training task, expressed as:

[0125] (11);

[0126] (12);

[0127] In the formula, represents the hyperparameter selection for the th training task generated by the network; represents the hyperparameter selection for the th training task generated by the network.

[0128] The process of the

[0129] network generating hyperparameter selection is expressed as:

[0130] (13);

[0131] In the formula, represents the hyperparameter selection strategy; and are the weight matrices of hyperparameters and performance data, and is the activation function of the LSTM network (e.g., the tanh function or the sigmoid function); represents the weight matrix in the network, which is used to weight the hyperparameter selection strategy and pass the weighted information to the activation function ; and is the bias term, which is used to control the activation value of the hyperparameter selection strategy i.e., it directly affects the generation process of hyperparameter selection; is the final adjustment term for the LSTM network to output the hyperparameter selection strategy, which directly affects the final output of the LSTM network. It enables the LSTM network to generate suitable hyperparameters according to the needs of the current training task by changing the activation value of the LSTM network output.

[0132] Step S3.43: Hyperparameter update: According to , train the improved BERT-WWM model in the current training task, and introduce an adaptive mechanism to adjust the hyperparameters according to the performance data of the improved BERT-WWM model.

[0133] Minimize the loss function of the training task by adjusting the hyperparameters , that is, select a set of optimal hyperparameter configurations in each training task to achieve the optimal performance of the improved BERT-WWM model; the loss function of the training task is expressed as:

[0134] (15);

[0135] In the formula, represents the output of the improved BERT-WWM model under ; is the loss corresponding to the training task .

[0136] The overall objective of the Meta-Learner LSTM based on LSTM is to optimize the loss functions of all training tasks , and find the optimal hyperparameter configuration, which is expressed as:

[0137] (16);

[0138] In the formula, is the total number of training tasks.

[0139] Among them, during the training process of the improved BERT-WWM model, the learning rate can be adjusted according to the training progress of the improved BERT-WWM model (the loss function Automatically decay through the change of

[0140] (17);

[0141] In the formula, is the current learning rate; is the adjustment factor; is the change amount of is the current loss value.

[0142] Minimize the sum of the loss functions of all training tasks to obtain the hyperparameters of the improved BERT-WWM model optimized by the LSTM-based meta-learning model , expressed as:

[0143] (18);

[0144] In the formula, represents minimizing the hyperparameters of the improved BERT-WWM model.

[0145] Through the LSTM-based meta-learning model, the automatic optimization of hyperparameters can be effectively achieved, enabling each new training task to quickly select the most appropriate hyperparameter configuration, which not only reduces the computational overhead of hyperparameter search but also reduces manual intervention through automation, improving the optimization efficiency and model performance.

[0146] Step S4: Input the power infrastructure text data to be classified into the optimized improved BERT-WWM model to obtain the classification result.

[0147] To prove the effectiveness of the optimized improved BERT-WWM model of the present invention, it is experimentally compared with traditional support vector machines (SVM), Naive Bayes classifiers, and BERT models on standard test sets (such as the AG News dataset, 20 Newsgroups dataset), and the comparison results are shown in Table 1.

[0148]

[0149] Table 1 Experimental comparison of different models

[0150] As can be seen from the results in Table 1, the classification accuracy of the optimized Improved BERT-WWM model of the present invention reaches 95%, which is significantly higher than that of traditional methods such as Support Vector Machine (SVM) and Naive Bayes classifier, proving the advantages of the optimized Improved BERT-WWM model of the present invention in understanding and classifying complex texts; the F1 score of the optimized Improved BERT-WWM model of the present invention is also better than that of other models, indicating that the optimized Improved BERT-WWM model of the present invention can achieve a better balance between precision and recall when dealing with complex and rare categories; Support Vector Machine (SVM) and Naive Bayes classifier are far superior to the optimized Improved BERT-WWM model of the present invention in terms of processing speed, but the processing speed of the optimized Improved BERT-WWM model of the present invention is still significantly higher than that of the BERT model; based on this result, the true effectiveness of the optimized Improved BERT-WWM model of the present invention is proved, and its competitiveness in text classification tasks is demonstrated.

[0151] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A text prompt classification method based on improved BERT-WWM, characterized in that: The steps include: Step S1: collecting power infrastructure text data for preprocessing and constructing a data set, where the power infrastructure text data consists of sentences of different lengths; Step S2: Build a BERT-WWM model, introduce sparse attention to replace the multi-head self-attention in the BERT-WWM model, and adaptively adjust the mask ratio in the sparse attention to obtain an improved BERT-WWM model; Step S3: using the data set to train the improved BERT-WWM model, and optimizing the parameters of the improved BERT-WWM model to obtain an optimized improved BERT-WWM model; Step S4: input the power infrastructure text data to be classified into the optimized improved BERT-WWM model to obtain the classification result; The specific process of optimizing and improving the parameters of the BERT-WWM model is as follows: Combine Dropout regularization and DropConnect regularization to regularize the parameters of the improved BERT-WWM model; The LAMB optimizer is used as the optimizer for improving the BERT-WWM model; An LSTM-based meta-learning model is used to optimize the hyperparameters of the improved BERT-WWM model training; Adaptively adjust the mask ratio in sparse attention, expressed as: ; In the formula, Indicates that the sentence length in the power infrastructure text data is The mask ratio when and It is an adjustable parameter.

2. The text prompting classification method based on improved BERT-WWM according to claim 1, characterized in that: The specific process of using the dataset to train the improved BERT-WWM model is: divide the dataset into a training set and a validation set in proportion, use the training set to train the improved BERT-WWM model, and use the validation set to validate the improved BERT-WWM model.

3. The text prompting classification method based on improved BERT-WWM according to claim 2, characterized in that: The specific process of combining Dropout regularization and DropConnect regularization to regularize the parameters of the improved BERT-WWM model is as follows: By weight factor The dropout probability of Dropout regularization and DropConnect regularization is weighted combined to obtain the joint regularization probability , expressed as: ; In the formula, and Represent the dropout probability of Dropout regularization and DropConnect regularization respectively; The drop probability is dynamically adjusted according to the gradient information of each layer of the improved BERT-WWM model, expressed as: ; In the formula, is the basic drop probability, To improve the BERT-WWM model The gradient of the layer weights, To improve the maximum value of all layer gradients of the BERT-WWM model; Indicates the improved BERT-WWM model The dropout regularization dropout probability of the layer; Calculate the regularization strength, expressed as: ; In the formula, Indicates that in improving the BERT-WWM model Layers, training rounds The regularization strength when ; To improve the total number of layers of the BERT-WWM model; is the total number of training rounds; is a preset initial regularization strength; is a hyperparameter that controls the decay rate; Indicates the current training round; Design an improved gradient update rule for the BERT-WWM model; in the case of joint regularization probability, the gradient update formula for the improved BERT-WWM model is: ; In the formula, To improve the BERT-WWM model Layer weights; is the learning rate; is the loss function For Partial derivatives of layer weights; and They are the mask matrices for Dropout regularization and DropConnect regularization respectively; Based on the improved performance of the BERT-WWM model on the validation set, dynamically adjust and : ; ; In the formula, Indicates the change in validation set loss; is the adjustment factor; is the loss of the current validation set; for During training ; For the During training ; for During training ; for During training .

4. The text prompting classification method based on improved BERT-WWM according to claim 3, characterized in that: The specific process of optimizing the hyperparameters of the improved BERT-WWM model training using the LSTM-based meta-learning model is as follows: Set the improved BERT-WWM model on the dataset The training task is ,Record The hyperparameters and improved BERT-WWM model completed the Performance data after training tasks ; Learn the hyperparameters and performance data of historical training tasks and generate the hyperparameter selection of the current training task, expressed as: ; ; In the formula, Indicates that The network generated Hyperparameter selection for each training task; Indicates that The network generated Hyperparameter selection for each training task; The process of network generation hyperparameter selection is expressed as: ; ; In the formula, represents the hyperparameter selection strategy; and is the weight matrix of hyperparameters and performance data, and is the activation function of the LSTM network; express The weight matrix in the network; and is the bias term; according to , train the improved BERT-WWM model in the current training task, and introduce an adaptive mechanism to adjust the hyperparameters according to the performance data of the improved BERT-WWM model.

5. The text prompting classification method based on improved BERT-WWM according to claim 4, characterized in that: according to , train the improved BERT-WWM model in the current training task, and introduce an adaptive mechanism to adjust the hyperparameters according to the performance data of the improved BERT-WWM model. The specific process is as follows: Minimize the loss function of the training task by adjusting the hyperparameters , that is, selecting a set of optimal hyperparameter configurations in each training task to achieve the best improved BERT-WWM model performance; the loss function of the training task It is expressed as: ; In the formula, Indicates that the improved BERT-WWM model is The output of the following; It is a training task corresponding losses; Optimize the loss function for all training tasks , find the optimal hyperparameter configuration, expressed as: ; In the formula, is the total number of training tasks; Introducing real-time feedback information of training tasks improves the accuracy change or loss change of the BERT-WWM model and dynamically adjusts the selection of hyperparameters, which can be expressed as: ; In the formula, is the current learning rate; is the adjustment factor; yes The amount of change; is the current loss value; Minimize the sum of the loss functions of all training tasks and obtain the hyperparameters of the improved BERT-WWM model after optimization of the LSTM-based meta-learning model , expressed as: ; In the formula, Represents the hyperparameters of minimizing the improved BERT-WWM model .

6. The text prompting classification method based on improved BERT-WWM according to claim 5, characterized in that: The LAMB optimizer is used as the optimizer for the improved BERT-WWM model, expressed as: ; In the formula, and Respectively represent the estimated values ​​of the mean and variance of the improved BERT-WWM model gradient; is a constant; Indicates level correction.

7. The text prompting classification method based on improved BERT-WWM according to claim 6, characterized in that: Collect power infrastructure text data for preprocessing, including: Text cleaning, using regular expressions to clean power infrastructure text data; Information extraction, extracting key information from text data of power infrastructure; Text vectorization: vectorize the power infrastructure text data after text cleaning.

Citation Information

Patent Citations

  • Multi-modal fine-grained thesis classification method and system based on regularization ensemble learning

    CN116956214A

  • Electric power defect entity identification method based on deep learning

    CN117829138A