Power customer service customer demand classification prediction method based on BERT and BiLSTM fusion technology

Through the integration of BERT and BiLSTM technology, the automated classification of work order data of power customer service is realized, the problem of low efficiency of traditional manual classification is solved, the efficiency and accuracy of appeal processing is improved, and efficient data governance means are provided for power customer service.

CN120296514APending Publication Date: 2025-07-11SOUTH BRANCH OF CUSTOMER SERVICE CENT OF STATE GRID CORP OF CHINA
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510393463.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional power customer service demand processing methods rely on manual classification, which is inefficient and difficult to meet the timeliness needs of the big data era. The existing technology lacks efficient data governance methods.

Method used

The power customer service customer demand classification prediction method based on BERT and BiLSTM fusion technology is adopted, and the automatic classification and identification of power customer service work order data is realized through data preprocessing, deep learning model training and semantic enhancement.

Benefits of technology

It significantly improves the efficiency and accuracy of appeal processing, provides data analysis foundation for multiple links, provides technical support for active services, and improves the accuracy of model in effectiveness judgment, appeal monitoring and business scenario classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296514A_ABST
    Figure CN120296514A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a BERT and BiLSTM fusion technology-based power customer service customer appeal classification prediction method. The prediction method comprises the following steps: model training; the method specifically comprises the following steps: data preprocessing; labeling the data; judging validity; semantic enhancement; classifying business scenes; and model application: outputting a business scene classification result. According to the method, deep analysis of data of a plurality of links such as 95598 customer appeals and electricity utilization is realized, key risk points are identified, and a data basis and technical support are provided for subsequent active services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a classification prediction method for power customer service customer demands based on the fusion technology of BERT and BiLSTM. Background Technique

[0002] With the continuous development of the business of power enterprises, grid customer service is facing increasingly complex customer demands. These demands are affected by various factors such as region, season, hot events, and emergency situations, resulting in significant differences in the daily demand distribution. At present, the traditional demand processing method mainly relies on manual classification of incoming call demands, manually collecting information to analyze the daily demand distribution, and finally feeding it back to the business department for process optimization. However, this method not only has a large processing pressure and low efficiency, but also is difficult to meet the high requirements for timeliness in the big data era. Therefore, how to efficiently integrate, identify, extract, mine, and apply customer demands has become a key link in the current State Grid customer service center to carry out customer demand data governance.

[0003] In such a background, the traditional demand processing method obviously lacks the ability of sustainable development. With the rapid development of artificial intelligence technology, natural language processing (NLP), as an important research direction in the field of artificial intelligence, has shown strong application potential in many fields, such as machine translation, public opinion monitoring, automatic summarization, opinion extraction, and text classification, providing new ideas and methods for solving the problems of customer demand processing faced by the State Grid customer service center. NLP technology can automatically learn the complex structures and rules of language through deep learning models, such as recurrent neural networks (RNN), long short-term memory networks (LSTM), and pre-trained models based on the Transformer architecture, such as BERT, to achieve efficient text representation and understanding. Summary of the Invention

[0004] The present invention disassembles the governance problem of customer demands as a multi-link classification task, and significantly improves the efficiency and accuracy of demand processing through the application of NLP technology. The specific technical solutions are as follows:

[0005] A classification prediction method for power customer service customer demands based on the fusion technology of BERT and BiLSTM includes the following processes:

[0006] S100: Model training; specifically includes the following processes:

[0007] S110: Data preprocessing; the specific process is as follows:

[0008] S111: Analyze the power customer service work order data based on partial speech-to-text conversion. Each original work order record only retains the desensitized data content, and each data format is <work order number, business type, speech-to-text content>;

[0009] S112: Filter non-Chinese characters, special symbols, and redundant spaces using regular expressions, and manually complete the missing records during the speech-to-text process in combination with the accepted work orders.

[0010] S113: Perform word segmentation and part-of-speech tagging on the data, combine it with common stop words to build a stop word list to filter out irrelevant words and function words; jointly define a domain dictionary with experts to replace power industry terms and accurately identify and remove stop words in professional scenarios, and split all demand data into granularity according to the format of consecutive questions and answers between customers and customer service for later data annotation and corpus construction.

[0011] S120: Data annotation; based on the preprocessed data, carry out demand validity annotation, demand detection annotation, and business demand annotation; the specific process is as follows:

[0012] S121: Validity annotation, two domain experts judge and annotate valid demands based on the fine-grained dialogue demands of customers and customer service.

[0013] S122: Demand detection annotation, jointly annotated by three domain experts to identify whether the dialogue is reasonable, whether it is normal dialogue content, and whether there is a possibility of demand.

[0014] S123: Business classification annotation, referring to the list of 8 major theme businesses sorted out by business experts, assign eight business experts to conduct annotation of business subcategories. There are about 15 business subcategories on average for each major category, and about 1000 annotation quantities for each type.

[0015] S130: Validity judgment, based on the design method of a binary classification model in deep learning, use CNN and BiLSTM to learn the data sets labeled as "valid" respectively; the prediction results of the data set validity can be obtained.

[0016] S140: Semantic enhancement, by loading the BERT pre-trained model and combining with the input corpus to calculate dynamic semantic features to mine deep semantic information, and at the same time introduce a custom position weight factor to adjust the influence of the continuous dialogue length on the semantic vector. Finally, calculate the quantization similarity based on the vector cosine to extract related texts with higher similarity to achieve semantic enhancement.

[0017] S150: Business scenario classification; the specific process is as follows:

[0018] S151: Obtain the initial BERT semantic feature vectors based on the dataset, and through the training of multiple layers of bidirectional Transformer encoders, convert the initial BERT feature vectors into dynamic semantic feature vectors in combination with the context; then, in combination with the deep semantic vectors obtained from the training of the BERT model, further perform feature extraction and sentence representation construction through the bidirectional long short-term memory network BiLSTM;

[0019] S152: Use the BERT vectors that capture the deep semantic information of the text as the input to the BiLSTM layer;

[0020] S153: Process using the forward propagation module LSTML and the backward propagation module LSTMR respectively to capture the context information in the text;

[0021] S154: Apply the Softmax function to normalize the output vectors, thereby obtaining the probability distribution matrix of multiple classifications of business demands; this process ensures that the prediction results of each label are based on a probability framework that combines rich features and highly discriminative representations, providing accurate probability estimates for each demand classification; through model training and execution, this module can obtain the corresponding business subclass classifications under different themes of the dataset;

[0022] S200: Model application; Input the voice-to-text power customer service work order data into the model determined by S100, and output the business scenario classification results.

[0023] Preferably, in S130, the BiLSTM embedding layer is a 300-dimensional Word2vec pre-trained model; the CNN convolution kernel size filter_sizes is set to 2, 3, 4, and the number of neurons is 256; the number of neurons in the BiLSTM hidden layer is set to 128, and the number of layers is 3.

[0024] This patent realizes in-depth analysis of data in multiple links such as 95598 customer demands and electricity consumption, identifies key risk points, and provides a data basis and technical support for subsequent proactive services. Brief Description of the Drawings

[0025] Figure 1 It is a schematic diagram of the BERT model framework structure.

[0026] Figure 2 It is a flowchart of the model training process of the classification prediction method for power customer service customer demands based on the BERT and BiLSTM fusion technology of the present invention. Detailed Embodiments

[0027] BERT is a bidirectional encoding representation model of the Transformer architecture, with powerful feature extraction capabilities. This model realizes context-aware feature extraction through multi-level semantic modeling. In terms of principle design, its core innovations are reflected in three aspects:

[0028] (1) Adopting a deeply stacked Transformer encoding layer, global token associations are established through the cooperation of the self-attention mechanism and the feed-forward neural network. This architecture not only retains the advantages of the Transformer in dealing with long-range dependencies but also avoids the computational redundancy brought by the decoder, achieving a remarkable balance between semantic representation efficiency and effect. Secondly, the masked language modeling (MLM) pre-training task is designed to construct multi-dimensional semantic reasoning capabilities using a random masking strategy. Finally, the CLS aggregation mechanism is introduced to converge the sequence encoding results to a specific classification token, thus establishing an end-to-end semantic representation paradigm. In terms of model application, on the one hand, BERT forms a deep bidirectional encoder through multiple layers of Transformer stacking, which can effectively capture long-distance context dependencies. At the same time, it adopts a dynamic semantic representation mechanism to construct a multi-dimensional semantic space based on position encoding (Position Embedding) and segment encoding (Segment Embedding), fully obtaining the context information of the corpus and generating dynamic semantic feature vectors according to the actual text context.

[0029] (2) BERT can establish a pre-training-fine-tuning transfer paradigm. Not only can the parameterized feature representations be directly adapted to various downstream tasks, but in addition, BERT itself can also be used as a complete network structure to penetrate into different model architectures, giving full play to the advantages of semantic mining.

[0030] The bidirectional long short-term memory network (BiLSTM), as a classic architecture for sequence modeling, solves the gradient decay problem of the traditional recurrent neural network (RNN) through a bidirectional temporal information fusion mechanism. Its core innovation lies in constructing forward-backward dual-channel memory units to achieve context-sensitive temporal feature extraction. Compared with the unidirectional LSTM, BiLSTM concatenates the two independently trained direction vectors in the hidden layer dimension through a parameter sharing mechanism to form a composite representation with spatio-temporal awareness. Specifically, the key points of this architecture include three aspects:

[0031] (1) Each gated memory unit consists of an input gate, a forget gate, and an output gate to form a non-linear regulation system, and through the cooperation of the sigmoid function and the tanh activation function, the dynamic update of the memory cell state (Cell State) is realized.

[0032] (2) It has a two-way propagation path strategy. The Forward Layer in the forward propagation captures historical dependence features, and the Backward Layer in the backward propagation extracts future context clues. The two share the same word embedding matrix but independently maintain hidden states.

[0033] (3) A gradient flow optimization algorithm is designed. The identity mapping shortcut is used to design the memory cell update formula, which effectively alleviates the vanishing gradient phenomenon in the training of deep networks and can effectively improve the model training efficiency.

[0034] Control Figure 2 , A classification prediction method for power customer service customer demands based on the BERT and BiLSTM fusion technology includes the following processes:

[0035] S100: Model training; The specific process is as follows:

[0036] S110: Data preprocessing; The specific process is as follows:

[0037] S111: Analyze the power customer service work order data based on partial speech-to-text conversion. Each original work order record only retains the desensitized data content, and each data format is <work order number, business type, speech-to-text content>;

[0038] S112: Use regular expressions to filter non-Chinese characters, special symbols, and redundant spaces, and manually complete the records lost during the speech-to-text conversion in combination with the received work orders;

[0039] S113: Perform word segmentation and part-of-speech tagging on the data, combine it with common stop words to construct a stop word list to filter out irrelevant words and function words; jointly define a domain dictionary with experts to replace power industry terms and accurately identify stop words in professional scenarios and remove them. Split all demand data in the format of continuous questions and answers between customers and customer service representatives for later data annotation and corpus construction;

[0040] S120: Data annotation; Based on the preprocessed data, carry out demand validity annotation, demand detection annotation, and business demand annotation; The specific process is as follows:

[0041] S121: Validity annotation. Two domain experts judge and annotate valid demands based on the fine-grained dialogue demands between customers and customer service representatives; A total of 84,000 pieces of data are annotated, and the ratio of valid to invalid data is approximately 6:4;

[0042] S122: Appeal Detection and Annotation. Since the main purpose of validity judgment is to identify whether a dialogue is reasonable, whether it is normal dialogue content, and whether there is a possibility of an appeal, it is necessary to perform secondary annotation on the valid data to explore whether there are more specific business appeals. This part of the data is jointly annotated by three domain experts to identify whether the dialogue is reasonable, whether it is normal dialogue content, and whether there is a possibility of an appeal. A total of 123,000 pieces of data are annotated, and the ratio of data with appeals to data without appeals is approximately 7:3.

[0043] S123: Business Classification Annotation. Referring to the list of 8 major theme businesses sorted out by business experts, eight business experts are assigned for the annotation of business subcategories. On average, there are about 15 business subcategories for each major category, and the annotation quantity for each category is about 1,000 pieces.

[0044] S130: Validity Judgment. Based on the design method of a binary classification model in deep learning, the datasets labeled as "valid" are learned through CNN and BiLSTM respectively; the prediction results of the dataset validity can be obtained. The embedding layer of BiLSTM is a 300-dimensional Word2vec pre-trained model. The filter sizes of the CNN convolutional kernels are set to 2, 3, and 4, and the number of neurons is 256. The number of neurons in the hidden layer of BiLSTM is set to 128.

[0045] The number of layers is 3.

[0046] S140: Semantic Enhancement. By loading the BERT pre-trained model and combining it with the input corpus to calculate dynamic semantic features, deep semantic information is mined. At the same time, a custom position weight factor is introduced to adjust the influence of the continuous dialogue length on the semantic vector. Finally, the cosine similarity of vectors is calculated to quantify the similarity, and texts with higher similarity are extracted jointly to achieve semantic enhancement.

[0047] S150: Business Scenario Classification; The specific process is as follows:

[0048] Based on the dataset, the initial BERT semantic feature vectors are obtained and trained through a multi-layer bidirectional Transformer encoder. Combining the context, the initial BERT feature vectors are converted into dynamic semantic feature vectors. Then, combining the deep semantic vectors obtained from the BERT model training, further feature extraction and sentence representation construction are performed through the bidirectional long short-term memory network BiLSTM.

[0049] The BERT vectors that capture the deep semantic information of the text are used as the input to the BiLSTM layer.

[0050] The forward propagation module LSTML and the backward propagation module LSTMR are used to process respectively to capture the context information in the text.

[0051] Apply the Softmax function to normalize the output vector, thereby obtaining a probability distribution matrix for multiple classifications of business requirements; this process ensures that the prediction results for each label are based on a probability framework that synthesizes rich features and highly discriminative representations, providing accurate probability estimates for each requirement classification; after model training and execution, this module can obtain the business sub-classifications corresponding to different themes in the dataset.

[0052] S200: Model application; Input the power customer service work order data converted from speech to text into the model determined by S100, and output the business scenario classification result.

[0053] The experimental environment configuration is shown in Table 1.

[0054] Table 1 Experimental environment

[0055] Experimental environment Configuration Python 3.12.4 Tensorflow 2.17.0 Kears 3.6.0 Operating system Window11 CPU Inteli7-12700 GPU RTX2070Super

[0056] Evaluation metrics: To objectively evaluate the classification effect of each link sub-model, this experiment uniformly uses precision P, recall R, and the harmonic mean F1-score of precision and recall to comprehensively evaluate the model effect. Considering that the semantic enhancement module is an unsupervised model that requires business personnel to jointly collaborate to set thresholds, this experiment mainly evaluates the three links of the effectiveness judgment model, the requirement monitoring module, and the business scenario classification module. Among them, TP represents the actual positive sample and the predicted positive sample, FP represents the actual negative sample but the predicted positive sample, and FN represents the actual positive sample but the predicted negative sample. The specific calculation formulas are as follows:

[0057]

[0058] In this experiment, the text dataset of each module is divided into a training set, a test set, and a validation set according to the ratio of 6:2:2, and five-fold cross-validation is combined to improve the generalization ability of the model. In addition, since the business scenario classification model involves eight business themes, in order to comprehensively consider the model performance, this experiment conducts a weighted calculation of the results of the eight themes. The specific results are shown in Table 2.

[0059] Table 2 Main link model result analysis

[0060]

[0061] According to the experimental results, it can be found that this experiment designs a multi-stage joint data governance framework, which has always achieved high effects in multiple data processing, prediction, and transfer processes, and the F1-Score has remained above 88%. This indicates that the framework has good synergy in all aspects of data governance and can effectively improve the overall performance.

[0062] In the validity judgment module, the BiLSTM model outperforms the CNN model. The F1-Score of the BiLSTM model is 90.00%, which is 1.51 percentage points higher than that of the CNN model at 88.49%. This indicates that when dealing with the validity judgment task, the BiLSTM model can better capture the context information in the text and effectively identify invalid expressions in the conversation. In contrast, although the CNN model performs well in local feature extraction, when the dataset length is long, its ability to capture context information is weak, resulting in slightly worse performance.

[0063] Similarly, in the appeal monitoring module, the BiLSTM model also outperforms the CNN model. The F1-Score of the BiLSTM model is 2.27 percentage points higher than that of the CNN model, indicating that when dealing with long text data, the BiLSTM model can more effectively extract features, thus improving the prediction performance of the model. Although the number of convolutional kernels and the length of the sliding window of the CNN model have been increased in the design of this module, when dealing with long text data, its ability to capture long-distance dependencies is still weak, resulting in performance inferior to that of the BiLSTM model.

[0064] In the business scenario classification module, it is not difficult to see that the BETR-BiLSTM model performs best, with its F1-Score being 7.06 percentage points higher than that of the BETR-CNN model and 12.64 percentage points higher than that of the CNN model. Further analysis reveals that in the same CNN task, based on the fine-grained semantic representations provided by the BERT pre-trained vectors, better results can be achieved compared to solving the traditional Word2Vec pre-trained model. In addition, due to the lack of temporal modeling ability, the BETR-CNN model performs poorly, while the BETR-BiLSTM can perform context encoding on the output dynamic semantic vectors and capture the combined features in the semantic sequence, thus significantly improving the classification accuracy and recall rate.

Claims

1. A classification prediction method for customer demands of power customer service based on the fusion technology of BERT and BiLSTM, characterized in that The process includes the following: S100: Model training; specifically includes the following process: S110: Data preprocessing; the specific process is as follows: S111: Conduct analysis based on partial voice-to-text power customer service work order data, where only the desensitized data content is retained for each original work order record, and the format of each data record is <work order number, business type, voice-to-text content>; S112: Use regular expressions to filter non-Chinese characters, special symbols, and redundant spaces, and manually complete the records lost during the speech-to-text conversion process in conjunction with the accepted work order; S113: Perform word segmentation and part-of-speech tagging on the data, and combine it with common stop words to build a stop word list to filter out irrelevant words and function words; jointly define a domain dictionary with experts to replace power industry terms and accurately identify and remove stop words in professional scenarios, and split all demand data into granularity according to the format of continuous questions and answers between customers and customer service for later data annotation and corpus construction; S120: Data labeling: Based on the pre-processed data, labeling of claim validity, claim detection and business claim is performed. The specific process is as follows: S121: Validity labeling: two domain experts judge and label valid demands based on the fine-grained conversation demands between customers and customer service staff; S122: Appeal detection and annotation, which is jointly annotated by three domain experts to identify whether the conversation is reasonable, whether it is normal conversation content, and whether there is a possibility of appeal; S123: Business classification annotation, referring to the list of 8 major subject businesses sorted out by business experts, eight business experts were assigned to annotate the business subcategories. There are an average of about 15 business subcategories in each major category, and the number of annotations for each category is about 1,000; S130: Validity judgment, based on the binary classification model design method of deep learning, using CNN and BiLSTM to learn the data sets marked as "valid" respectively; the prediction results of the validity of the data sets can be obtained; S140: Semantic enhancement, by loading the BERT pre-trained model and calculating dynamic semantic features based on the input corpus to mine deep semantic information, while introducing custom position weight factors to adjust the impact of continuous conversation length on the semantic vector, and finally calculating the quantified similarity based on vector cosine to extract texts with high similarity and achieve semantic enhancement; S150: Business scenario classification; the specific process is as follows: S151: Obtain the BERT initial semantic feature vector based on the dataset, and after training with a multi-layer bidirectional Transformer encoder, convert the initial BERT feature vector into a dynamic semantic feature vector in combination with the context; Combined with the deep semantic vector obtained by BERT model training, further perform feature extraction and sentence representation construction through a bidirectional long short-term memory network (BiLSTM); S152: The BERT vector that captures the deep semantic information of the text is used as the input of the BiLSTM layer; S153: using the forward transfer module LSTML and the backward transfer module LSTMR to process respectively to capture the context information in the text; S154: Apply the Softmax function to normalize the output vector, thereby obtaining a probability distribution matrix for multiple classifications of business requirements; this process ensures that the prediction results for each label are based on a probability framework that combines rich features and highly discriminative representations, providing accurate probability estimates for each requirement classification; after model training and execution, this module can obtain the corresponding business subclass classifications under different themes of the dataset. S200: Model application; Input the voice-to-text power customer service work order data into the model determined by S100, and output the business scenario classification results.

2. The classification prediction method for power customer service customer demands based on the BERT and BiLSTM fusion technology according to claim 1, wherein, In the said S130, the BiLSTM embedding layer is a 300-dimensional Word2vec pre-trained model; the CNN convolution kernel size filter_sizes is set to 2, 3, and 4, and the number of neurons is 256; the number of neurons in the BiLSTM hidden layer is set to 128, and the number of layers is 3.

Citation Information

Cited By

  • Conference management method and system based on natural language processing and retrieval enhancement

    CN120634500A

  • A conference management method and system based on natural language processing and retrieval enhancement

    CN120634500B