Hot line platform label model training method based on deep neural network
By training the hotline platform's label model using a deep neural network, the problems of missed and false judgments in work order classification by traditional rule engines have been solved. This has enabled efficient and adaptive automatic work order classification and label generation, improving the accuracy and real-time performance of government hotline services.
Patent Information
- Application Number
- CN202511283412.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-12-23
AI Technical Summary
Traditional rule-based work order classification and dispatch methods lack a deep understanding of text semantics and context, leading to missed judgments and misjudgments. They are difficult to handle diverse expressions, have high maintenance costs, cannot quickly adapt to emergencies or new demands, and affect real-time response capabilities.
The hotline platform label model is trained using deep neural networks. Through data preprocessing, pre-trained BERT models, gating attention mechanisms, graph neural networks, and meta-learning frameworks, combined with an online feedback closed-loop mechanism, automatic work order classification and label generation are achieved, improving classification accuracy and adaptability.
It significantly improved the accuracy of work order classification, reduced the error rate of cross-departmental work order dispatch and the delay in responding to new events, reduced operation and maintenance costs, and met the needs of refined management and service quality.
Smart Images

Figure CN121189385A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of model training technology, and in particular to a method for training a label model for a hotline platform based on a deep neural network. Background Technology
[0002] Traditional rule-based work order classification and dispatch methods mainly rely on keyword matching or manually defined regular expression rules. They lack a deep understanding of text semantics and context, making it difficult to handle diverse expressions of the same event. This can easily lead to missed judgments and misjudgments, limiting the system's performance in large-scale, highly complex data processing scenarios.
[0003] Furthermore, rule engine systems are costly to maintain in practice, requiring continuous manual updates and optimization of the rule base to adapt to new event types or policy changes. New rules may conflict with existing rules, increasing system logic complexity and impacting real-time response capabilities. Traditional methods cannot quickly adapt to sudden events or new demands, leading to processing delays and sluggish responses, making it difficult to meet the needs of refined management and improved service quality. Summary of the Invention
[0004] To address the above shortcomings, this invention provides a hotline platform tag model training method based on deep neural networks, aiming to construct an intelligent, scalable, and adaptive data tagging system. By utilizing deep learning and semantic understanding technologies, it achieves automatic work order classification and tag generation, thereby improving classification accuracy, dispatch efficiency, and the overall service capabilities of the platform.
[0005] In a first aspect, the present invention provides the following technical solution: a hotline platform label model training method based on a deep neural network, comprising:
[0006] The input hotline work order text is preprocessed, and the structured metadata of the work order is extracted.
[0007] A pre-trained BERT model is used to semantically encode the pre-processed work order text to obtain text feature vectors. The structured metadata is converted into metadata vectors, and the text feature vectors and metadata feature vectors are dynamically fused through a gating attention mechanism.
[0008] The fused feature vectors are input into the hierarchical encoder, and the second-stage pre-training of the MLM task is performed using government hotline work order data to learn government terminology and citizens' spoken language expression patterns, and output domain-enhanced semantic representation vectors.
[0009] The co-occurrence probability of labels in historical work orders is statistically analyzed, a label association graph adjacency matrix is constructed, the domain-enhanced semantic representation vector is input into the graph neural network decoder, and the dependencies between labels are learned through the message passing mechanism to achieve multi-label collaborative prediction and output the label probability distribution.
[0010] The problem of imbalanced samples is solved by using label probability distribution and focus loss function, and the model is trained and optimized by combining meta-learning framework, so that the model can adaptively optimize new event type labels based on a small number of samples.
[0011] The trained teacher model is compressed into a student model through knowledge distillation, and the student model is deployed to the NPU hardware platform to achieve lightweight deployment. An online feedback closed-loop mechanism is established to collect manual correction results for incremental training.
[0012] Preferably, the data preprocessing step includes:
[0013] Based on a dialect dictionary, the work order text is normalized to convert local colloquial expressions into standard language expressions.
[0014] Named entity recognition technology is used to identify sensitive information such as ID card numbers, phone numbers, and detailed addresses in work order texts, and these are replaced with special mask tags for entity desensitization.
[0015] Sequence labeling models are used to detect and correct typos in work order texts, thereby improving text quality.
[0016] Preferably, the dynamic fusion step of the gating attention mechanism includes:
[0017] The joint feature vector is obtained by concatenating the text feature vector and the metadata vector.
[0018] The joint feature vector is linearly transformed by a fully connected layer, and the gating weights are calculated using the sigmoid activation function.
[0019] The text feature vector and metadata vector are weighted and combined based on gating weights to generate a fused feature vector.
[0020] Preferably, the layered encoder includes a bottom-level architecture and a top-level architecture. The bottom-level architecture is a general BERT model composed of 12 Transformers, and the top-level architecture is a government affairs domain enhancement layer composed of 4 Transformers.
[0021] Preferably, the step of outputting the domain-enhanced semantic representation vector includes:
[0022] The underlying encoder is used to capture general semantic features, and the top-level encoder is used to perform masked language model pre-training tasks based on government hotline work order corpus.
[0023] During the pre-training process, specific vocabularies and terminology databases from the government sector are introduced to semantically enhance colloquial expressions used by citizens, policy and regulatory terminology, and entities of government agencies.
[0024] Preferably, the output label probability distribution step includes:
[0025] Statistical analysis was performed on the multi-label annotation results of historical hotline work orders, and the co-occurrence probability between each label pair was calculated.
[0026] Each label is used as a node in the label association graph neural network, and an adjacency matrix of the label association graph neural network is constructed based on the co-occurrence probability.
[0027] The domain-enhanced semantic representation vector is input into the label association graph neural network, and the feature propagation and update between label nodes are performed by message passing mechanism to obtain the representation vector of each label node.
[0028] The updated label node representation vectors are classified, the probability distribution of each label is calculated, and the final multi-label prediction result is output based on the set probability threshold.
[0029] Preferably, the adaptive optimization step includes:
[0030] In multi-label classification training, the focus loss function is used as the task loss, and the class weights are adaptively set by combining the label frequency in the historical work order or meta-task support set.
[0031] The training is carried out using a model-independent meta-learning framework. In each meta-task, a small number of samples are divided into a support set and a validation set. The model parameters are first updated with gradients on the support set based on the task loss to obtain task-specific adaptation parameters.
[0032] Continue to calculate the loss value based on the task loss on the validation set, and backpropagate the gradient to update the global initialization parameters;
[0033] Through multi-round meta-task training, global initialization parameters with fast convergence capability are obtained.
[0034] Preferably, the lightweight deployment step includes:
[0035] The trained teacher model is set as a pre-trained BERT model, and a lightweight student model is trained as a BiLSTM network using the knowledge distillation method, so that the student model can reduce the number of model parameters and computational complexity while maintaining the predictive ability of the teacher model.
[0036] The trained student model is loaded onto the NPU hardware platform, and inference speed and computational efficiency are further improved through quantization, model pruning, or other lightweight optimization techniques.
[0037] Preferably, the online feedback closed-loop mechanism includes the following steps:
[0038] Collect manually corrected hotline work order annotations and automatically add them to the training dataset;
[0039] Incremental fine-tuning of the teacher model enables it to learn new label patterns or correct erroneous predictions;
[0040] The updated model continues to be deployed in the inference system to achieve continuous adaptation to new event types or minority class labels and improve prediction accuracy;
[0041] By repeatedly executing the above steps, an online feedback loop for model self-evolution is formed.
[0042] The present invention has the following beneficial effects:
[0043] 1. This invention introduces structured metadata (such as channel, region, time, etc.) from hotline work orders into an attention gating mechanism to achieve dynamic weighted fusion of textual features. In this way, the model can not only capture semantic information in the work order text but also highlight sensitive or important scenario features by combining structured information. For example, in special situations such as epidemic prevention and control, and emergencies, the fused features can significantly improve the model's accuracy and response speed in identifying relevant requests, thereby achieving more accurate work order classification and dispatch, and providing intelligent and efficient service support for government hotlines.
[0044] 2. To address the issue of traditional multi-label classification methods neglecting the correlation between labels, this invention constructs a label co-occurrence probability matrix as the adjacency matrix of a graph neural network, treating each label as a node for feature propagation and updating. Through the message passing mechanism of the GNN, the model can learn the dependencies between labels, achieving multi-label collaborative prediction. This not only improves the recognition accuracy of rare or combined labels but also enhances the system's classification capabilities in complex, multi-dimensional work order scenarios, providing a stable and reliable multi-label prediction solution for government hotlines.
[0045] 3. This invention, based on the general BERT model, introduces a second-stage domain-adaptive pre-training strategy, utilizing a massive amount of unlabeled government hotline work orders to perform a masked language modeling task. This strategy can capture the semantic features of citizens' colloquial expressions, policy terminology, and government agency entities, significantly improving the model's ability to represent government-related texts. Combined with fine-tuning training, the system demonstrates higher semantic understanding and classification accuracy in multi-label work order classification tasks, thus effectively supporting precise work order dispatch, trend analysis, and intelligent decision-making. Attached Figure Description
[0046] Figure 1 This is a flowchart of the hotline platform label model training method based on deep neural networks proposed in this invention.
[0047] Figure 2 This is a simplified diagram illustrating the implementation of the hotline platform label model training method based on deep neural networks proposed in this invention.
[0048] Figure 3 This is the corresponding figure for the improvement of the technical effect of the present invention. Specific implementation manners
[0049] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] Embodiment 1
[0051] In the first embodiment of the present invention, the present invention provides a method for training a label model of a hotline platform based on a deep neural network, as Figure 1 shown, including the following steps:
[0052] S100: Perform data preprocessing on the input hotline work order text, and extract the structured metadata of the work order at the same time;
[0053] Preferably, the data preprocessing step includes:
[0054] Perform dialect normalization on the work order text based on a dialect dictionary, and convert local spoken language expressions into standard language expressions;
[0055] Use named entity recognition technology to identify sensitive information such as ID numbers, phone numbers, and detailed addresses in the work order text, and replace them with special mask marks for entity desensitization;
[0056] Use a sequence labeling model to detect and correct typos in the work order text to improve the text quality.
[0057] Specifically, for local dialects or colloquial expressions that appear in the hotline work order, the system first calls the dialect dictionary to map local expressions to standard Chinese expressions. For example, the expression "it's raining and the house is leaking" that appears in a work order from the Guangdong region will be mapped to "it's raining and the house is leaking", and the Cantonese colloquialism "traffic jam" will be normalized to "traffic congestion".
[0058] To protect the privacy of citizens, the system uses named entity recognition technology to automatically identify sensitive information such as ID numbers, phone numbers, and detailed addresses in the work order text, and replaces them with special mask marks. For example, "ID number (consisting of 18 digits)" will be replaced with "[ID]", and the detailed address "No. 123 XX Road" will be replaced with "XX Road ** No.".
[0059] Work order texts often have spelling mistakes or non-standard inputs. The system uses a sequence labeling model, such as BiLSTM-CRF, to detect and correct spelling mistakes in work order texts. For example, the text that was misrecognized as "wood meter running very fast" due to OCR error in "water meter running very fast" is automatically corrected to the standard expression.
[0060] Through the above data preprocessing steps, the quality of hotline work order texts has been significantly improved, dialects and oral expressions have been unified, sensitive information has been effectively desensitized, and spelling mistakes have been corrected.
[0061] S200: Use the pre-trained BERT model to perform semantic encoding on the preprocessed work order text to obtain text feature vectors, convert the structured metadata into metadata vectors, and dynamically fuse the text feature vectors and metadata feature vectors through a gated attention mechanism;
[0062] Preferably, the step of dynamically fusing the gated attention mechanism includes:
[0063] Perform a concatenation operation on the text feature vector and the metadata vector to obtain a joint feature vector;
[0064] Perform a linear transformation on the joint feature vector through a fully connected layer, and use the sigmoid activation function to calculate the gating weights;
[0065] Based on the gating weights, perform a weighted combination of the text feature vector and the metadata vector to generate a fused feature vector.
[0066] Specifically, first, the text feature vector and the metadata vector are concatenated to obtain a joint feature vector [[ID=, where t and m represent the dimensions of the two vectors respectively.
[0067] In order to dynamically control the contribution ratio of text features and metadata features in the fusion, perform a linear transformation on the joint feature vector through a fully connected layer, and use the sigmoid activation function to generate a gating weight vector. This step can be expressed as: where, W g represents the weight matrix of the fully connected layer; b g represents the bias vector, σ represents the sigmoid activation function, making each gating value between 0 and 1. g represents the retention ratio of text features in each dimension. The larger the value, the greater the contribution of text features, and the smaller the value, the more the contribution of metadata features.
[0068] According to the gating weight vector g, perform a weighted combination of the text feature vector and the metadata vector to obtain the final fused feature vector F. This step can be expressed as F = g ⊙ T + (1 - g) ⊙ M ' , where, ⊙ represents element-wise multiplication, M' This means that a linear transformation maps the metadata vector M to a vector with the same dimension as the text features.
[0069] Through the aforementioned gating attention mechanism, the model can dynamically adjust the weights of text and metadata based on the specific work order scenario. When the work order text contains vague or colloquial expressions, the weight of metadata will be increased accordingly, enhancing the reliance on regional or channel features; when the work order text is complete and clear, the weight of text features will dominate, improving the accuracy of semantic understanding.
[0070] S300: Input the fused feature vector into the hierarchical encoder, use government hotline work order data for the second stage of pre-training of the masked language model (MLM) task, learn government terminology and citizens' spoken expression patterns, and output domain-enhanced semantic representation vectors;
[0071] Preferably, the layered encoder includes a bottom-level architecture and a top-level architecture. The bottom-level architecture is a general BERT model composed of 12 Transformers, and the top-level architecture is a government affairs domain enhancement layer composed of 4 Transformers.
[0072] Specifically, the hierarchical encoder comprises a bottom-level architecture and a top-level architecture, which are connected in series. The bottom-level architecture consists of a general BERT model composed of 12 layers of Transformers, used to extract general semantic features of the text. Each Transformer layer includes a multi-head self-attention mechanism and a feedforward network, capable of global context modeling of words in the input text sequence and capturing long-distance dependencies. The top-level architecture consists of 4 layers of Transformers, used for semantic enhancement of government hotline work orders. The top-level architecture incorporates a government-specific terminology list and an organization entity database through two-stage pre-training (masked language model task) on the government work order corpus.
[0073] The underlying BERT provides general semantic understanding capabilities, capable of handling complex syntax and long texts; the top-level government affairs enhancement layer strengthens the ability to recognize policy terminology, colloquial expressions, and institutional entities through two-stage pre-training. The hierarchical encoder significantly improves the overall accuracy of multi-label classification, especially in recognizing dialects, colloquial expressions, and work orders related to new policies, which is significantly better than the single-layer general BERT.
[0074] Preferably, the step of outputting the domain-enhanced semantic representation vector includes:
[0075] The underlying encoder is used to capture general semantic features, and the top-level encoder is used to perform masked language model pre-training tasks based on government hotline work order corpus.
[0076] During the pre-training process, specific vocabularies and terminology databases from the government sector are introduced to semantically enhance colloquial expressions used by citizens, policy and regulatory terminology, and entities of government agencies.
[0077] Specifically, the input to the bottom-level encoder is the feature vector F after gating and attention fusion. The output, obtained through the Transformer's self-attention mechanism, captures the global dependencies between words in the sequence, achieving general semantic modeling. The output of each Transformer layer is passed to the next layer through residual connections and layer normalization; this step can be represented as H. (l) =LayerNorm(H (l-1) +Attention(H (l-1) )), where H (0) =F, l=1,…,12, Attention represents self-attention, and LayerNorm represents layer normalization.
[0078] The top-level encoder consists of four Transformer layers connected in series, directly receiving the output H from the bottom-level encoder. (12) The top-level encoder undergoes two-stage MLM pre-training on the government hotline work order corpus to enhance the model's ability to understand government-specific expressions. The MLM training objective formula is:
[0079]
[0080] in, x represents the set of word indices to be masked. i H represents actual words. top This represents the sequence features output by the top-level encoder.
[0081] During pre-training, a domain-specific vocabulary and terminology database are introduced, including policy and regulatory terms, government agency entities, and common colloquial expressions. The vector output by the top-level encoder is the domain-enhanced semantic representation vector.
[0082] S400: Calculate the co-occurrence probability of tags in historical work orders, construct the adjacency matrix of the tag association graph, input the domain-enhanced semantic representation vector into the graph neural network decoder, learn the dependencies between tags through the message passing mechanism, realize multi-tag collaborative prediction and output the tag probability distribution;
[0083] Preferably, the output label probability distribution step includes:
[0084] Statistical analysis was performed on the multi-label annotation results of historical hotline work orders, and the co-occurrence probability between each label pair was calculated.
[0085] Each label is used as a node in the label association graph neural network, and an adjacency matrix of the label association graph neural network is constructed based on the co-occurrence probability.
[0086] The domain-enhanced semantic representation vector is input into the label association graph neural network, and the feature propagation and update between label nodes are performed by message passing mechanism to obtain the representation vector of each label node.
[0087] The updated label node representation vectors are classified, the probability distribution of each label is calculated, and the final multi-label prediction result is output based on the set probability threshold.
[0088] Specifically, the first step is to perform statistical analysis on the multi-label annotation results of historical hotline work orders. The label set is defined as L = {l1, l2, ..., l...}. n}, where n is the total number of tags. For any two tags l i and l j The co-occurrence probability of these work orders is calculated by traversing the historical work order dataset: P(l i ,l j ) = count(l i ∩l j ) / count(l i ∪l j ), where count(l i ∩l j ) indicates that it contains the tag l i and l j The number of work orders, count(l i ∪l j ) indicates that it contains the tag l i or l j The total number of work orders.
[0089] Based on the calculated co-occurrence probabilities, an adjacency matrix is constructed for the label association graph. in
[0090]
[0091] Where θ is a preset threshold (usually set to 0.1), used to filter weak correlations and reduce the impact of noise.
[0092] Domain-enhanced semantic representation vector H u As the initial node features of the graph neural network, message passing is performed through K layers of graph convolution: in, Let N(v) represent the hidden state representation of node v at layer k, N(v) represent the set of neighbor nodes of node v, MLP represents multilayer perceptron, which is used to transform the features of neighbor nodes, and GRU represents gated recurrent unit, which is used to fuse the current state and neighbor information.
[0093] After k layers of graph convolution, each label node obtains a final representation vector that incorporates neighbor information. The probability distribution of each label is calculated using a fully connected layer and a softmax function:
[0094]
[0095] Wherein, it represents the label l i The predicted probability, W out Let b represent the output layer weight matrix. out This represents the output layer bias vector. A probability threshold τ (typically 0.5) is set for each label l. i If the probability is greater than the probability threshold τ, the prediction result is 1; otherwise, it is 0. The final output multi-label prediction result is the set of all labels whose probabilities exceed the threshold.
[0096] This method, which constructs an adjacency matrix based on historical co-occurrence probabilities and uses GRU message passing, effectively models the dependencies between labels and improves the synergy and accuracy of multi-label prediction.
[0097] S500: It solves the sample imbalance problem based on label probability distribution and focus loss function, and combines the meta-learning framework to optimize model training, enabling the model to adaptively optimize new event type labels based on a small number of samples.
[0098] Preferably, the adaptive optimization step includes:
[0099] In multi-label classification training, the focus loss function is used as the task loss, and the class weights are adaptively set by combining the label frequency in the historical work order or meta-task support set.
[0100] The training is carried out using a model-independent meta-learning framework. In each meta-task, a small number of samples are divided into a support set and a validation set. The model parameters are first updated with gradients on the support set based on the task loss to obtain task-specific adaptation parameters.
[0101] Continue to calculate the loss value based on the task loss on the validation set, and backpropagate the gradient to update the global initialization parameters;
[0102] Through multi-round meta-task training, global initialization parameters with fast convergence capability are obtained.
[0103] Specifically, in multi-label classification tasks, the class distribution is extremely imbalanced, with common labels having a much larger sample size than rare labels. To avoid the model becoming overly biased towards the majority class, this invention employs a focus loss function as the training task loss. The focus loss is defined as: FL(p) = -α t (1-p) γlog(p), where p represents the model's predicted probability of the true class, γ is a modulating factor, typically 1–3, used to reduce the loss contribution of easily classified samples and highlight difficult samples, and α t This is a category balancing factor used to adjust the importance of different categories.
[0104] During the training phase, a Model-Independent Meta-Learning (MAML) framework is employed to enable the model to quickly adapt to new label scenarios with only a limited number of samples. Within this framework, the work order task is first divided into multiple meta-tasks, each containing a small number of samples, further subdivided into support and validation sets. On the support set, the task loss is calculated using the aforementioned focus loss function, and the model parameters are updated one or more times using gradients to obtain task-specific adaptation parameters. Where θ represents the current global model parameters, and α represents the learning rate. This represents the gradient based on the focus loss function on the support set. On the validation set, it utilizes the updated parameters θ. ' Continue calculating the loss based on the focus loss function. The gradient of this loss is then backpropagated to the initial parameters θ, thereby updating the globally initialized parameters. Where β represents the meta-learning rate. By repeatedly performing the above process on multiple meta-tasks, the global initialization parameter θ* with fast convergence capability is finally obtained.
[0105] Through the above implementation methods, the focus loss function enhances the learning of difficult samples and minority class labels. The MAML framework enables the model to quickly optimize with only a small number of samples when encountering new label types, avoiding large-scale retraining. By combining globally initialized parameters and label frequency adaptive weights, the model significantly improves performance in minority class and new label predictions.
[0106] S600: It compresses the trained teacher model into a student model through knowledge distillation, deploys the student model to the NPU hardware platform to achieve lightweight deployment, and establishes an online feedback closed-loop mechanism to collect manual correction results for incremental training.
[0107] Preferably, the lightweight deployment step includes:
[0108] The trained teacher model is set as a pre-trained BERT model, and a lightweight student model is trained as a BiLSTM network using the knowledge distillation method, so that the student model can reduce the number of model parameters and computational complexity while maintaining the predictive ability of the teacher model.
[0109] The trained student model is loaded onto the NPU hardware platform, and inference speed and computational efficiency are further improved through quantization, model pruning, or other lightweight optimization techniques.
[0110] Specifically, firstly, the pre-trained BERT model is used as the teacher model. While this model possesses strong text semantic representation and label prediction capabilities, its large parameter count and high computational complexity mean that direct deployment on edge hardware platforms like NPUs leads to inference latency, making it difficult to meet real-time requirements. To reduce model complexity while maintaining predictive power, a knowledge distillation method is used to compress the teacher model into a lightweight student model. The student model employs a BiLSTM structure, with a significantly smaller parameter size than BERT. During distillation, the teacher model's output serves as the "soft label," combined with the ground truth "hard label," to guide the student model's training. The distillation loss function consists of two parts: the difference loss between the teacher model's predicted probability distribution and the student model's output distribution, and the cross-entropy loss between the student model's predictions and the ground truth labels. By comprehensively optimizing these two parts, the student model can effectively learn the knowledge from the teacher model while exhibiting strong generalization ability.
[0111] The trained BiLSTM student model is loaded onto the NPU hardware platform for online real-time inference. To further improve operational efficiency, the following lightweight optimization techniques can be introduced during deployment: model quantization, which compresses model parameters from floating-point representation to low-bit integers, significantly reducing storage space and computational overhead with almost no loss of accuracy; model pruning, which prunes redundant neurons and connections, retaining parameters that have a significant impact on prediction results, further compressing the model size; and compilation optimization, which combines the instruction set and memory characteristics of the NPU to perform compiler-level optimizations such as operator fusion and memory optimization on the model inference computation graph to accelerate inference speed.
[0112] Through the lightweight deployment process described above, this invention enables real-time inference on the NPU hardware platform, significantly reducing latency and meeting the real-time requirements of hotline ticket processing. Simultaneously, compared to directly deploying the BERT model, storage and computation costs are significantly reduced.
[0113] Preferably, the online feedback closed-loop mechanism includes the following steps:
[0114] Collect manually corrected hotline work order annotations and automatically add them to the training dataset;
[0115] Incremental fine-tuning of the teacher model enables it to learn new label patterns or correct erroneous predictions;
[0116] The updated model continues to be deployed in the inference system to achieve continuous adaptation to new event types or minority class labels and improve prediction accuracy;
[0117] By repeatedly executing the above steps, an online feedback loop for model self-evolution is formed.
[0118] Specifically, after the model is deployed to the actual hotline platform, the system will output the label prediction results for each work order. If the prediction results are inaccurate, manual reviewers will make corrections. The system automatically collects and stores these manually corrected annotation results to form new high-quality annotation data.
[0119] The manually corrected results collected are directly added to the original training dataset to form an incremental dataset containing the latest annotation information. This process does not require completely relabeling historical data, but rather gradually expands the dataset so that the training corpus continuously covers more scenarios and new types of events.
[0120] Use the incremental dataset to perform incremental fine-tuning on the teacher model, keeping the existing model parameters as the initial weights to avoid the model forgetting historical knowledge; perform short-cycle fine-tuning training on the new data so that the model can quickly learn new label patterns or correct incorrect predictions; introduce regularization constraints during the training process to prevent the model from overfitting to a small amount of new data.
[0121] Replace the incrementally fine-tuned model into the inference system and continue to run on the hotline platform. In this way, the model can immediately reflect the latest learned label patterns, thereby improving the prediction accuracy for new event types or minority class labels.
[0122] Embodiment 2:
[0123] This embodiment provides a method for training a label model of a hotline platform based on a deep neural network, which is applied to an intelligent label system for government hotline work orders to achieve automated and multi-label accurate classification prediction of hotline work orders submitted by citizens. Its flow diagram is as Figure 2 shown, and the specific steps are as follows:
[0124] After a citizen submits a work order by phone or mini-program, the system first performs data preprocessing on the input work order text. Through a dialect dictionary, local colloquial expressions (such as "repair it quickly") are converted into standardized written language ("repair as soon as possible"); use named entity recognition technology to detect and desensitize sensitive information such as ID numbers, phone numbers, and detailed addresses in the work order, and replace them with special mask tokens; use a sequence annotation model to detect and correct typos in the work order text to ensure the quality of the input text. At the same time, structured metadata such as the channel (such as phone, website, WeChat), region (such as city district, street), and time (submission time, work order occurrence time) of the work order are extracted as additional input features.
[0125] A pre-trained Chinese BERT model is used to semantically encode the work order text, obtaining text feature vectors. The extracted metadata is converted into numerical vectors and encoded through an embedding layer. Next, the text feature vectors and metadata vectors are concatenated and input into a gating attention mechanism. The gating layer calculates weights using a fully connected network and a sigmoid function, and then weights the two types of features to obtain a fused feature representation.
[0126] The fused feature vectors are input into a hierarchical encoder. The hierarchical encoder comprises: a bottom-level encoder, consisting of 12 stacked Transformer layers, employing the BERT-based Chinese model, primarily used to capture general language semantic features; and a top-level encoder, consisting of 4 stacked Transformer layers, specifically pre-trained for government affairs corpora. In this stage, a large amount of government hotline work order data is used to perform an MLM task, and a government affairs-specific vocabulary (such as "urban management enforcement," "sanitation operations," and "medical insurance reimbursement") and a policy and regulatory terminology database are introduced to enhance the model's understanding of citizens' colloquial expressions and government professional vocabulary, thereby outputting domain-enhanced semantic representation vectors.
[0127] In the multi-label prediction stage, the co-occurrence frequency of labels in historical work orders is first statistically analyzed, and the co-occurrence probability between each pair of labels is calculated. Using each label as a node, an adjacency matrix of the label graph is constructed based on the co-occurrence probability. The domain-enhanced semantic representation vector is input into the graph neural network decoder, and features are propagated and updated between label nodes through a message passing mechanism, thereby learning the dependencies between labels. Finally, the representation vector of each label node is obtained, and the probability distribution of each label is output through a classifier. Based on a set probability threshold, the final multi-label prediction result is output.
[0128] To address the issue of severely imbalanced work order label distribution, this embodiment employs a focus loss function as the primary training loss function to reduce attention to easily classified samples and enhance the learning ability for minority class samples. Simultaneously, a meta-learning framework is used for training. First, a small number of samples are divided into a support set and a validation set. Fast gradient updates are performed on the support set to obtain task-specific parameters. On the validation set, the loss is further calculated, and global parameters are updated. Through multiple rounds of meta-task training, globally initialized parameters capable of rapidly adapting to new event types are obtained.
[0129] After training, the high-performing BERT teacher model is used as the knowledge source to train a lightweight BiLSTM student model through knowledge distillation. This significantly reduces the number of model parameters and computational complexity while maintaining the predictive power of the teacher model. Finally, the student model is deployed on the NPU hardware platform of the government hotline system. Combined with model quantization and pruning optimization, the real-time inference performance of the model on the hardware is improved to meet the high concurrency requirements of the business.
[0130] In actual operation, if the model predicts labels incorrectly, human reviewers will correct them. The system automatically collects these manual corrections, forming a new incremental dataset. Based on the incremental data, the teacher model is fine-tuned, and then the student model is updated through knowledge distillation before being redeployed to the NPU platform. This process is repeated cyclically, enabling the model to continuously learn new label patterns and colloquial expressions from citizens, gradually improving prediction accuracy and forming a self-evolving online feedback loop.
[0131] Compared to the traditional rule-based government hotline service platform data tag training system, the effectiveness verification data of this system is shown in Table 1. This invention demonstrates significant advantages in work order classification accuracy, cross-departmental work order dispatch error rate, new event response speed, single work order processing latency, and average annual maintenance cost. The corresponding graphs for the improved technical performance are shown below. Figure 3 As shown, the classification accuracy increased from 71.2% to 92.7%, the cross-departmental dispatch error rate decreased from 22.6% to 4.1%, the new event response delay was shortened from 48 hours to 30 minutes, the single work order processing delay was reduced from 2000ms to 50ms, and the average annual operation and maintenance cost decreased from 2 million yuan to 300,000 yuan. This fully verifies the comprehensive superiority of the invention in terms of accuracy, timeliness, and economy, and demonstrates its wide application value in actual government hotline work order processing scenarios.
[0132] Table 1. Performance verification data of this invention compared to rule engines.
[0133] index Rule Engine The device of the present invention Increase Work order classification accuracy (F1) 71.2% 92.7% ↑21.5% Cross-departmental order dispatch error rate 22.6% 4.1% ↓82% New event response delay 48 hours 30 minutes ↓99% Single-work order processing delay 2000ms 50ms ↓97.5% Annual operating and maintenance costs 2 million yuan 300,000 yuan ↓85%
[0134] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A training method for a hotline platform label model based on a deep neural network, characterized in that, include: The input hotline work order text is preprocessed, and the structured metadata of the work order is extracted. A pre-trained BERT model is used to semantically encode the pre-processed work order text to obtain text feature vectors. The structured metadata is converted into metadata vectors, and the text feature vectors and metadata feature vectors are dynamically fused through a gating attention mechanism. The fused feature vectors are input into the hierarchical encoder, and the second-stage pre-training of the MLM task is performed using government hotline work order data to learn government terminology and citizens' spoken language expression patterns, and output domain-enhanced semantic representation vectors. The co-occurrence probability of labels in historical work orders is statistically analyzed, a label association graph adjacency matrix is constructed, the domain-enhanced semantic representation vector is input into the graph neural network decoder, and the dependencies between labels are learned through the message passing mechanism to achieve multi-label collaborative prediction and output the label probability distribution. The problem of imbalanced samples is solved by using label probability distribution and focus loss function, and the model is trained and optimized by combining meta-learning framework, so that the model can adaptively optimize new event type labels based on a small number of samples. The trained teacher model is compressed into a student model through knowledge distillation, and the student model is deployed to the NPU hardware platform to achieve lightweight deployment. An online feedback closed-loop mechanism is established to collect manual correction results for incremental training.
2. The hotline platform label model training method based on deep neural networks according to claim 1, characterized in that, The data preprocessing steps include: Based on a dialect dictionary, the work order text is normalized to convert local colloquial expressions into standard language expressions. Named entity recognition technology is used to identify sensitive information such as ID card numbers, phone numbers, and detailed addresses in work order texts, and these are replaced with special mask tags for entity desensitization. Sequence labeling models are used to detect and correct typos in work order texts, thereby improving text quality.
3. The hotline platform label model training method based on deep neural networks according to claim 1, characterized in that, The dynamic fusion steps of the gated attention mechanism include: The joint feature vector is obtained by concatenating the text feature vector and the metadata vector. The joint feature vector is linearly transformed by a fully connected layer, and the gating weights are calculated using the sigmoid activation function. The text feature vector and metadata vector are weighted and combined based on gating weights to generate a fused feature vector.
4. The hotline platform label model training method based on deep neural networks according to claim 1, characterized in that, The layered encoder includes a bottom-level architecture and a top-level architecture. The bottom-level architecture is a general BERT model composed of 12 Transformers, and the top-level architecture is a government affairs domain enhancement layer composed of 4 Transformers.
5. The hotline platform label model training method based on deep neural networks according to claim 1, characterized in that, The steps for outputting the domain-enhanced semantic representation vector include: The underlying encoder is used to capture general semantic features, and the top-level encoder is used to perform masked language model pre-training tasks based on government hotline work order corpus. During the pre-training process, specific vocabularies and terminology databases from the government sector are introduced to semantically enhance colloquial expressions used by citizens, policy and regulatory terminology, and entities of government agencies.
6. The hotline platform label model training method based on deep neural networks according to claim 1, characterized in that, The output label probability distribution step includes: Statistical analysis was performed on the multi-label annotation results of historical hotline work orders, and the co-occurrence probability between each label pair was calculated. Each label is used as a node in the label association graph neural network, and an adjacency matrix of the label association graph neural network is constructed based on the co-occurrence probability. The domain-enhanced semantic representation vector is input into the label association graph neural network, and the feature propagation and update between label nodes are performed by message passing mechanism to obtain the representation vector of each label node. The updated label node representation vectors are classified, the probability distribution of each label is calculated, and the final multi-label prediction result is output based on the set probability threshold.
7. The hotline platform label model training method based on deep neural networks according to claim 1, characterized in that, The adaptive optimization steps include: In multi-label classification training, the focus loss function is used as the task loss, and the class weights are adaptively set by combining the label frequency in the historical work order or meta-task support set. The training is carried out using a model-independent meta-learning framework. In each meta-task, a small number of samples are divided into a support set and a validation set. The model parameters are first updated with gradients on the support set based on the task loss to obtain task-specific adaptation parameters. Continue to calculate the loss value based on the task loss on the validation set, and backpropagate the gradient to update the global initialization parameters; Through multi-round meta-task training, global initialization parameters with fast convergence capability are obtained.
8. The hotline platform label model training method based on deep neural networks according to claim 1, characterized in that, The lightweight deployment steps include: The trained teacher model is set as a pre-trained BERT model, and a lightweight student model is trained as a BiLSTM network using the knowledge distillation method, so that the student model can maintain the predictive ability of the teacher model while reducing the number of model parameters and computational complexity. The trained student model is loaded onto the NPU hardware platform, and inference speed and computational efficiency are further improved through quantization, model pruning, or other lightweight optimization techniques.
9. The hotline platform label model training method based on deep neural networks according to claim 1, characterized in that, The steps of the online feedback closed-loop mechanism include: Collect manually corrected hotline work order annotations and automatically add them to the training dataset; Incremental fine-tuning of the teacher model enables it to learn new label patterns or correct erroneous predictions; The updated model continues to be deployed in the inference system to achieve continuous adaptation to new event types or minority class labels and improve prediction accuracy; By repeatedly executing the above steps, an online feedback loop for model self-evolution is formed.
Citation Information
Cited By
Entity recognition method for Chinese spoken text and related equipment
CN121787414A
Text and structured data classification method based on cascade model and feature fusion
CN122112735A