Lightweight work order business type identification model training method, lightweight work order business type identification method and lightweight work order business type identification device
By cleaning and text-enhancing government work order sample data, combining word embedding models and time vectorization, training teacher and student models, and using knowledge distillation and intermediate feature alignment loss, the problem of inaccurate classification of government work orders in existing technologies is solved, and efficient and accurate classification is achieved in resource-constrained environments.
Patent Information
- Application Number
- CN202510778509.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-11
AI Technical Summary
In actual application scenarios with limited resources, existing technologies make it difficult to accurately and efficiently classify government work orders, especially in CPU environments and when training samples are insufficient, the advantages of pre-trained models cannot be fully utilized.
By cleaning and text-enhancing the work order sample data to expand the sample quantity, combining the word embedding model to process the text content, and vectorizing the generation date and time, we train a teacher model dominated by a linear fully connected layer and a student model dominated by a long short-term memory network and an attention network, and use knowledge distillation and intermediate feature alignment loss for model training.
It achieves efficient and accurate classification of government work orders in a resource-constrained environment, saves computing resources, reduces the demand for hardware capabilities, and improves the accuracy of model recognition.
Smart Images

Figure CN120670952A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of work order data processing, and in particular to a lightweight work order business type recognition model training method, recognition method and device. Background Art
[0002] In existing technologies, government service platforms are responsible for accepting and distributing a large number of citizen complaints, but manual sorting and classification can no longer meet the needs of efficient processing. Traditional text classification methods are mainly divided into two steps: feature extraction and machine learning model classification. Although they are efficient, they have low accuracy and poor generalization ability. In recent years, deep learning models represented by the Transformer model, such as BERT, have been widely used in text classification. However, they have high requirements for hardware resources and the number of training data samples, and usually require a high-performance computing environment with GPUs and a large amount of labeled data. Government work order data has the characteristics of being unstructured, colloquial, and having a wide range of topics. In addition, the time when the work order is generated will affect the type of work order. For example, government work orders are mostly concentrated during working hours on weekdays. In actual application scenarios with limited resources, existing technologies find it difficult to fully utilize the advantages of pre-trained models and combine the characteristics of government work order data to achieve efficient and accurate government text classification. Therefore, a new work order type identification solution is urgently needed. Summary of the Invention
[0003] In view of this, an embodiment of the present invention provides a lightweight work order business type identification model training method, identification method and device to eliminate or improve one or more defects existing in the existing technology and solve the problem that the existing technology is difficult to accurately and efficiently classify government work orders under resource constraints.
[0004] Therefore, the present invention provides a lightweight work order business type identification model training method, which includes the following steps: Performing data cleaning on a plurality of work order sample data, wherein the work order sample data includes time information and text content; The sample data of the work order is expanded by using text enhancement, and the work order type is added as a label for the classification task to construct a training sample set; the work order type includes government work orders and non-government work orders; Processing the text content of the work order sample data in the training sample set through a preset word embedding model to obtain a first embedding vector, looking up a table to retrieve a second embedding vector indicating whether it belongs to a working day based on the generation date in the time information, and looking up a table to retrieve a third embedding vector indicating the formation time period based on the generation time in the time information; concatenating the first embedding vector, the second embedding vector, and the third embedding vector into an overall embedding vector; Using the training sample set to train a teacher model dominated by a linear fully connected layer, the teacher model takes the overall embedding vector as input and outputs a first prediction value for the classification task, constructs a first classification loss based on the deviation between the first prediction value and the label, and updates the parameters of the teacher model in combination with a time constraint loss; The training sample set is used to train the student model dominated by the long short-term memory network and the attention network. The student model takes the overall embedding vector as input and outputs a second prediction value of the classification task. A second classification loss is constructed based on the deviation between the second prediction value and the label. The parameters of the student model are updated by combining the distillation loss of the difference in prediction distribution between the student model and the updated teacher model and the intermediate feature alignment loss to obtain a work order business type recognition model.
[0005] In some embodiments, the data cleaning includes: removing special characters, invalid content, repeated text, Chinese word segmentation, and stop words; The sample data of the work order is expanded by using text enhancement, including: Based on the work order sample data, the sample is expanded by using synonym replacement and sentence transformation based on HowNet synonym dictionary; By constructing guide words, a preset large language model is used to generate similar sample expansion based on the work order sample data, and the large language model adopts ChatGPT or Deepseek; The word embedding model is the RoBERTa model.
[0006] In some embodiments, the teacher model includes a first fully connected layer, an activation function layer, a Dropout layer, a second fully connected layer and a first softmax layer; the activation function layer adopts a Tanh function; The student model includes a bidirectional long short-term memory network, an attention layer, a third fully connected layer and a second softmax layer.
[0007] In some embodiments, the output of the teacher model is: ; ; in, represents the first predicted value; E represents the overall embedding vector, 、 are the parameters of the first fully connected layer, 、 is the parameter of the second fully connected layer; z represents the output of the second fully connected layer; Then the calculation formula for the first classification loss is: ; in, represents the first classification loss, BCELoss is the binary cross entropy loss function; y represents the label; The calculation formula of the time constraint loss is: ; in, represents the time constraint loss, represents the third embedding vector, is the Frobenius norm, is used to constrain the similarity of the third embedding vector corresponding to adjacent generation times; ; The total loss calculation formula of the first classification loss combined with the time constraint loss in the method is: ; in, is the total loss of the teacher model, is the weight of the time constraint loss.
[0008] In some embodiments, the second classification loss is calculated as: ; in, represents the second classification loss; represents the second predicted value; y represents the label; The calculation formula of the distillation loss is: ; in, represents the distillation loss, is the output probability distribution of the second fully connected layer in the teacher model, represents the output probability distribution of the third fully connected layer in the student model; KL represents the calculation of KL divergence; is a fixed parameter; The calculation formula of the intermediate feature alignment loss is: ; in, represents the intermediate feature alignment loss, represents the output features of the Dropout layer in the teacher model, represents the student model; In the method, the total loss calculation formula of the second classification loss combined with the distillation loss and the intermediate feature alignment loss is: ; Among them, β and γ are weight coefficients.
[0009] In some embodiments, the method further uses an Adma optimizer to update parameters of the teacher model and the student model.
[0010] On the other hand, the present invention also provides a method for identifying a work order business type, the method comprising the following steps: Acquire work order data to be identified, the work order data including time information and text content; Processing the text content of the work order data through a preset word embedding model to obtain a first embedding vector, looking up a table to retrieve a second embedding vector indicating whether it belongs to a working day based on the generation date in the time information, and looking up a table to retrieve a third embedding vector indicating the formation time period based on the generation time in the time information; and concatenating the first embedding vector, the second embedding vector, and the third embedding vector into an overall embedding vector; The overall embedding vector is input into the work order business type recognition model in the above-mentioned lightweight work order business type recognition model training method, and the recognition result of the work order category is output.
[0011] On the other hand, the present invention also provides a work order business type identification device, including a processor, a memory and a computer program / instruction stored in the memory, wherein the processor is used to execute the computer program / instruction, and when the computer program / instruction is executed, the device implements the steps of the above method.
[0012] On the other hand, the present invention further provides a computer-readable storage medium having a computer program / instruction stored thereon, which implements the steps of the above method when executed by a processor.
[0013] On the other hand, the present invention also provides a computer program product, comprising a computer program / instruction, which implements the steps of the above method when executed by a processor.
[0014] The lightweight work order business type identification model training method, identification method, and device described in the present invention expand the sample size by cleaning and text-enhancing the work order sample data. Based on word embedding of the text content of the work order sample data, the generation date and generation time are also vectorized to explore the relationship between the work order classification results and the generation time. After training a large-scale parameter teacher model dominated by a linear fully connected layer, a student model dominated by a long short-term memory network and an attention network is trained based on knowledge distillation. During the parameter update process, the hard loss of the prediction results, the distillation loss of the difference in the prediction distribution between the student model and the teacher model, and the intermediate feature alignment loss are combined to ensure model recognition accuracy on the basis of obtaining a lightweight work order business classification and identification model, save computing resources, and reduce the demand for hardware capabilities.
[0015] Additional advantages, objects, and features of the present invention will be set forth in part in the following description and will become apparent to those skilled in the art upon examination of the following or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained by the structures particularly pointed out in the description and drawings.
[0016] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present invention are not limited to the above specific descriptions, and the above and other purposes that can be achieved by the present invention will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation of the present invention. In the drawings: Figure 1 The figure is a flowchart of a lightweight work order business type identification model training method according to an embodiment of the present invention.
[0018] Figure 2 This is a flow chart of a lightweight BERT-based automatic classification solution for government work orders according to another embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0020] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, the accompanying drawings only show structures and / or processing steps closely related to the solutions according to the present invention, while other details that are not closely related to the present invention are omitted.
[0021] It should be emphasized that the term "include / comprises" when used herein refers to the existence of features, elements, steps or components, but does not exclude the existence or addition of one or more other features, elements, steps or components.
[0022] It should also be noted that, unless otherwise specified, the term "connection" herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.
[0023] With the digital transformation of cities and the widespread deployment of government service platforms, citizens can submit a variety of information, including requests, suggestions, inquiries, and complaints, through various channels such as telephone, apps, WeChat mini-programs, and websites. The platforms quickly and accurately distribute these text messages to the relevant departments, ensuring timely responses and resolution. However, with the continuous increase in citizen participation, the number of requests received daily by government service platforms has increased exponentially, and manual sorting and classification are no longer sufficient to meet the needs of efficient processing. The work order text data generated by government service platforms is unstructured, colloquial, and covers a wide range of topics. Traditional text classification methods mainly involve two steps: first, extracting text features such as BoW, N-gram, and TF-IDF to represent the text data, and then using machine learning models such as logistic regression, support vector machines, and tree models represented by XGBoost for classification. Although these methods are relatively efficient, they have low accuracy and poor generalization. In recent years, end-to-end deep learning models represented by the Transformer model, such as pre-training models such as BERT, have been widely used in text classification. However, pre-training models such as BERT have high requirements for hardware resources and the number of samples of training data, and usually require a high-performance computing environment with a GPU and a large amount of labeled data. In addition, unlike ordinary texts, the time when government work orders are generated will also affect the type of work orders. For example, government work orders are generally concentrated in working hours on weekdays, such as between 8 am and 5 pm. Therefore, in actual application scenarios with limited resources, such as in a CPU environment and when training samples are insufficient, how to give full play to the advantages of pre-training models and combine the characteristics of government work order data to achieve efficient and accurate government text classification is the problem to be solved by the present invention.
[0024] Specifically, the present invention provides a lightweight work order business type recognition model training method, such as Figure 1 As shown, the method includes the following steps S101 to S105: Step S101: performing data cleaning on a plurality of work order sample data, where the work order sample data includes time information and text content.
[0025] Step S102: The sample data of the work order is expanded by using text enhancement, and the work order type is added as a label of the classification task to construct a training sample set; the work order type includes government work orders and non-government work orders.
[0026] Step S103: Process the text content of the work order sample data in the training sample set through a preset word embedding model to obtain a first embedding vector. A second embedding vector is retrieved based on the generation date in the time information to indicate whether it belongs to a working day. A third embedding vector is retrieved based on the generation time in the time information to indicate the formation period. The first, second, and third embedding vectors are concatenated into an overall embedding vector.
[0027] Step S104: Use the training sample set to train the teacher model dominated by the linear fully connected layer. The teacher model takes the overall embedding vector as input and outputs the first prediction value of the classification task. The first classification loss is constructed based on the deviation between the first prediction value and the label, and the parameters of the teacher model are updated in combination with the time constraint loss.
[0028] Step S105: Use the training sample set to train the student model dominated by the long short-term memory network and the attention network. The student model takes the overall embedding vector as input and outputs the second prediction value of the classification task. The second classification loss is constructed based on the deviation between the second prediction value and the label. The student model parameters are updated by combining the distillation loss of the difference in prediction distribution between the student model and the updated teacher model and the intermediate feature alignment loss to obtain a work order business type recognition model.
[0029] In step S101, the purpose of data cleaning is to remove noise and extract clean, valid text data to lay the foundation for subsequent processing. In some embodiments, data cleaning includes: removing special characters, invalid content, repeated text, Chinese word segmentation, and stop words.
[0030] The purpose of removing special characters and invalid content is to clean up noise in the text, making it cleaner and easier to process. Common special characters include line feed (\n), tab (\t), non-breaking space (\u3000), and zero-width space (\u200b). Furthermore, ticket text may contain irrelevant information such as "ticket source," "ticket flow," "message title," and "message body." This information is generally unhelpful for classification and needs to be removed as well.
[0031] For example, the original text is: "Work Order Source: Phone Message\n\nMessage Title: Road Maintenance Issue\u3000\u200b\n\nMessage Body: Hello, I would like to report that there are many potholes on the road near my home, which affects travel safety." The text after cleaning is: "Hello, I would like to report that there are many potholes on the road near my home, which affects travel safety." Due to the possibility of the same user submitting the same problem repeatedly or multiple users complaining about the same matter, there may be texts that are exactly the same or highly similar in the work orders. To reduce redundancy, improve data quality and processing efficiency, it is necessary to remove duplicate texts. For exactly duplicate texts, the duplicate items can be directly deleted. For texts that are highly similar in content, techniques such as MinHash (Minimum Hash Algorithm) can be used for deduplication.
[0032] Chinese word segmentation is to split continuous Chinese text into individual lexical units because Chinese expresses meaning based on words. Accurate word segmentation helps with subsequent text representation and analysis. Commonly used Chinese word segmentation tools include Jieba, and the word segmentation mode can be selected as "accurate mode" to ensure the maximum degree of restoring the user's expression intention.
[0033] Stop words refer to words that frequently appear in the text but have little significance in distinguishing the theme meaning of the text, such as "de", "le", "shi", "zai", "he", etc. Removing stop words can reduce the data dimension, reduce the complexity of model training, and at the same time help improve the performance and interpretability of the model. Usually, a common stop word list and a custom reserved word list are loaded to clean the word segmentation results and delete common stop words. In addition, high-frequency invalid words that appear in a large number of texts can be further removed according to the document frequency, and these high-frequency words can be extended and incorporated into the stop word list.
[0034] In step S102, the text enhancement technology can effectively alleviate the problem of insufficient training sample quantity for government work orders. Specifically, the rule enhancement method is based on historical work order texts and uses methods such as synonym replacement and sentence pattern transformation to expand the work order texts, such as using HowNet's Chinese Thesaurus of Synonyms. Additionally, using the Few-Shot text generation method, combined with Prompt design and a small amount of historical government work order texts, similar complaint texts for government work orders are generated with the help of large language models such as ChatGPT, Deepseek, Kimi, etc.
[0035] Specifically, the sample expansion is carried out on the work order sample data by means of text enhancement, including steps S1021 and S1022: Step S1021: Based on the work order sample data, sample expansion is carried out by means of synonym replacement and sentence pattern transformation based on HowNet's Chinese Thesaurus of Synonyms.
[0036] HowNet is a large-scale Chinese lexical semantic knowledge base, whose synonym dictionary provides rich synonym information. We analyze the vocabulary in the work order sample and search for synonyms in the HowNet synonym dictionary for key expressions. We then replace these synonyms, for example, replacing "the road is bumpy" with "the road surface is bumpy." This generates new samples and enriches the text's expression. Based on the lexical semantic relationships and grammatical structure information provided by HowNet, we adjust the sentence structure of the work order text. For example, we transform "I discovered that the garbage at the community gate was not cleared in a timely manner" into "I discovered that the garbage at the community gate was not cleared in a timely manner." This further expands the sample and enables the model to learn the same semantic features expressed in different sentence structures.
[0037] Step S1022: by constructing guide words, using a preset large language model to generate similar sample expansion based on the work order sample data, the large language model adopts ChatGPT or Deepseek.
[0038] Guide words guide the large language model to generate text that meets the requirements. In the context of government work order text enhancement, guide words can be keywords in the work order subject, common complaint and suggestion sentences, and so on. The guide words and original work order sample data are provided as input to a large language model (such as ChatGPT or Deepseek). Based on the guide word prompts and the semantic information of the original sample, the model generates new samples that are similar to the original work order in terms of semantics and style, thereby expanding the dataset and increasing sample diversity.
[0039] Furthermore, step S102 defines the work order types as government work orders and non-government work orders, thereby dividing the core objectives of executing classified tasks. In other embodiments, government work orders can be further divided into government service, urban management, public service, and economic management categories based on business areas.
[0040] According to the degree of urgency, government work orders can be further divided into: urgent work orders, relatively urgent work orders and regular work orders.
[0041] According to the complaint type, government work orders can be further divided into: service quality complaints, service efficiency complaints and facilities and equipment complaints.
[0042] According to the type of suggestion, government work orders can be further divided into: service optimization suggestion category, policy formulation suggestion category and facility improvement suggestion category.
[0043] According to the scope of impact, government work orders can be further divided into: personal impact work orders, local impact work orders and wide impact work orders.
[0044] According to the nature of the problem, government work orders can be further divided into: public safety, environmental protection and social livelihood.
[0045] In step S103, the text content of the work order sample data in the training sample set is processed through a preset word embedding model to obtain a first embedding vector. Here, the preset word embedding model can adopt the Chinese RoBERTa pre-training model, and the first token of the text (i.e., the vector corresponding to [CLS]) or the average pooled vector of the first and last tokens is taken as the embedding vector of the main content field. According to the generation date in the time information, the second embedding vector is taken out to mark whether it belongs to a working day. If it is a working day, it is marked as 1, otherwise it is 0, and the corresponding embedding vector is obtained by table lookup. According to the generation time in the time information, the third embedding vector of the time period is taken out to mark the formation time. The time of the day is divided into 24 hours, the embedding vector of the hour when the work order is generated is learned, and the embedding vector of the hour is taken out by table lookup. Finally, the first embedding vector, the second embedding vector, and the third embedding vector are connected to form an overall embedding vector. The overall embedding vector contains the characteristics of the text content and time information.
[0046] In step S104, the teacher model takes the overall embedding vector as input and outputs the first prediction value of the classification task. A first classification loss is constructed based on the deviation between the first prediction value and the label. The classification loss of the teacher model can be calculated using a binary cross entropy loss function. In order to better learn the embedding vector of time, a time constraint loss is also added. Considering that the time characteristics of two adjacent time slots are similar, the corresponding time embedding vectors should also be similar, and a specific loss function is used to incorporate the above-mentioned time characteristics. The total loss of the teacher model is calculated, where the weight of the time constraint loss is 0.5 in the embodiment of the present invention. The Adam optimizer is used to backpropagate the teacher model loss and update the parameters of the teacher model. The optimizer learning rate can be set to 0.001 to complete the parameter update of the teacher model.
[0047] In some embodiments, the teacher model includes a first fully connected layer, an activation function layer, a Dropout layer, a second fully connected layer and a first softmax layer; the activation function layer adopts the Tanh function.
[0048] In some embodiments, the output of the teacher model is: ; ; in, represents the first predicted value; E represents the overall embedding vector, 、 are the parameters of the first fully connected layer, 、 is the parameter of the second fully connected layer; z represents the output of the second fully connected layer.
[0049] The calculation formula for the first classification loss is: ; in, represents the first classification loss, BCELoss is the two-class cross entropy loss function; y represents the label; The calculation formula of time constraint loss is: ; in, represents the time constraint loss, represents the third embedding vector, is the Frobenius norm, is used to constrain the similarity of the third embedding vector corresponding to adjacent generation times; ; The total loss calculation formula of the first classification loss combined with the time constraint loss in the method is: ; in, is the total loss of the teacher model, is the weight of the time constraint loss.
[0050] In step S105, the fine-tuned teacher model is used to predict the training set samples, obtaining the teacher model's classification scores, which are then smoothed and used as soft labels for the student model. After text preprocessing, the student model randomly initializes word vectors, first passes them through a bidirectional long short-term memory (BiLSTM) network to obtain hidden representations, and then enters the attention layer to obtain a text representation. The student model's hard loss is calculated, which is the loss constructed by the deviation of the student model's classification score from the label after the probability distribution of the student model's predicted category is obtained through the layers. The distillation loss is calculated based on the difference between the probability distribution predicted by the student model and the teacher model's soft labels, using the KL divergence. The intermediate feature alignment loss is calculated using the cosine similarity function using the fine-tuned teacher model text embedding vector and the student model text representation. The total loss of the student model is obtained by taking the weighted sum of the hard loss, distillation loss, and intermediate feature alignment loss. The hard loss weight can be set to 0.3, and the distillation loss weight can be set to 0.7. The Adam optimizer is used to backpropagate the total loss of the student model and update the parameters of the student model. The optimizer learning rate is set to 0.001, and finally a work order business type recognition model is obtained, which can predict and judge the business type of new government work orders added daily.
[0051] In some embodiments, the student model includes a bidirectional long short-term memory network, an attention layer, a third fully connected layer, and a second softmax layer.
[0052] In some embodiments, the second classification loss is calculated as: ; in, represents the second classification loss; represents the second predicted value; y represents the label.
[0053] The calculation formula for distillation loss is: ; in, represents the distillation loss, is the output probability distribution of the second fully connected layer in the teacher model, represents the output probability distribution of the third fully connected layer in the student model; KL represents the calculation of KL divergence; is a fixed parameter.
[0054] The calculation formula for the intermediate feature alignment loss is: ; in, represents the intermediate feature alignment loss, represents the output features of the Dropout layer in the teacher model, Represents the student model.
[0055] In the method, the total loss calculation formula of the second classification loss combined with the distillation loss and the intermediate feature alignment loss is: ; Among them, β and γ are weight coefficients.
[0056] On the other hand, the present invention further provides a method for identifying a work order business type, the method comprising the following steps S201 to S203: Step S201: Acquire work order data to be identified, where the work order data includes time information and text content.
[0057] Step S202: Process the text content of the work order data through a preset word embedding model to obtain a first embedding vector. A second embedding vector is retrieved based on the date of creation in the time information to indicate whether the work order data belongs to a working day. A third embedding vector is retrieved based on the time of creation in the time information to indicate the time period. The first, second, and third embedding vectors are concatenated to form an overall embedding vector.
[0058] Step S203: input the entire embedding vector into the work order business type recognition model in the above-mentioned lightweight work order business type recognition model training method, and output the recognition result of the work order category.
[0059] On the other hand, the present invention also provides a work order business type identification device, including a processor, a memory and a computer program / instruction stored in the memory, wherein the processor is used to execute the computer program / instruction, and when the computer program / instruction is executed, the device implements the steps of the above method.
[0060] On the other hand, the present invention further provides a computer-readable storage medium having a computer program / instruction stored thereon, which implements the steps of the above method when executed by a processor.
[0061] On the other hand, the present invention also provides a computer program product, comprising a computer program / instruction, which implements the steps of the above method when executed by a processor.
[0062] The present invention will be described below in conjunction with a specific embodiment: This embodiment proposes a lightweight BERT-based automatic classification solution for government work orders, which specifically includes the following modules: Work order text preprocessing module: This module preprocesses the text of government work orders based on the characteristics of government work orders, including removing special characters and invalid content, removing duplicate text, Chinese word segmentation, and removing stop words.
[0063] Work order text enhancement module: In view of the small number of government work orders, rule enhancement and generation technology based on pre-trained language models are used to generate artificially synthesized text.
[0064] Government work order embedding vectorization initialization module: First, the Chinese RoBERTa pre-trained model is used to obtain the embedding vector of the work order text data. Then, based on the time field of the work order, it is determined whether it is a weekday and its hourly features are extracted. This results in the embedding vector of the weekday feature and the hourly embedding vector. These three are added together to obtain the final initial embedding vector of the government work order.
[0065] Teacher model training module: During the teacher model training phase, the initial embedding vector of the government work order is input into the BERT model, and then the labeled training samples are trained through a two-layer feedforward neural network.
[0066] Student model training module: During the student model training phase, the knowledge of the BERT model is transferred to a lightweight student model through the knowledge distillation method. The student model uses a bidirectional long short-term memory model based on the attention mechanism.
[0067] Government Affairs Work Order Classification Prediction Model: For the government affairs work orders that need to be predicted, first perform text preprocessing, and then use the trained student model for prediction.
[0068] As Figure 2 shown, the specific implementation steps are as follows: 1. Work Order Text Preprocessing: Extract the main content fields in the government affairs work order and perform preprocessing operations on the text, specifically including: 1.1 Remove special characters and invalid content: The original text is mixed with various special characters and symbols, such as \n, \t, \u3000, \u200b, etc. In addition, information such as "Work Order Source", "Work Order Transfer", "Message Title", and "Message Body" is common at the beginning and end of many message texts. Use regular expressions and string replacement methods to clear the above noise data. After removing the above noise, work orders with less than a certain number of words are directly excluded to improve the accuracy of subsequent processing. In the example of this invention, the number of words is taken as 10.
[0069] 1.2 Remove duplicate texts: Due to the same user submitting repeatedly or multiple people complaining about the same matter, there are texts with exactly the same or highly similar content in the work orders. Use the minhash technology to remove duplicates for approximately duplicate texts.
[0070] 1.3 Chinese text word segmentation: After completing text cleaning, use the Jieba word segmentation tool to segment the text. Segment the text in "accurate mode" and import the entity retention word list to ensure the maximum restoration of the user's expression intention.
[0071] 1.4 Remove stop words: First, load the common stop word list and the custom retention word list, clean the word segmentation results, and delete common stop words, such as "de", "le", "qingkuang", "yixia", etc. At the same time, to further improve the feature sparsity and discrimination, calculate the document frequency of each word and eliminate high-frequency invalid words that appear in a large number of texts. Then expand the high-frequency vocabulary into the stop word list to form a reusable text preprocessing component. [[ID=2,1]]
[0072] 2. Work Order Text Enhancement: In view of the situation that the number of government affairs work orders is small, use text enhancement technology to generate synthetic texts, specifically including: 2.1 Rule Enhancement Method: Based on the historical work order text, use the HowNet synonym thesaurus in Chinese to expand the work order text by means of synonym replacement and sentence pattern transformation; 2.2 Few-Shot Text Generation Method: Combine Prompt design and a small amount of historical government affairs work order text, and use large language models such as ChatGPT, Deepseek, Kimi, etc. to generate complaint texts for similar government affairs work orders.
[0073] The design of the guide word prompt can refer to the example "Please imitate the complaint message styles of the following three citizens in the "12345 hotline" to generate a new citizen government affairs appeal text. The topic should be close to public services and people's livelihood issues, with a natural tone, clear logic, and content of no less than 150 words. Example 1:..., Example 2:..., Example 3:..."
[0074] 3. Initialize the embedding vector of government work orders: 3.1 First, extract the main content field in the government work order, that is, the complaint text, load the output layer of the Chinese RoBERTa pre-trained model, and take the first token of the text (that is, the vector corresponding to [CLS]) or the average pooled vector of the first and last tokens as the embedding vector of the main content field ; 3.2 Determine whether it is a working day based on the time the work order is generated, learn the embedding vector of whether the work order is a working day, if it is a working day, it is 1, otherwise it is 0, and retrieve the embedding vector of whether it is a working day by looking up the table ; 3.3 Divide a day into 24 hours and learn the hourly embedding vector generated by the work order. The hourly embedding vector is dimensional matrix, each row represents the hour. Take out the hour generated by the government work order, ranging from 0 to 23, and take out the embedding vector of the hour by looking up the table .
[0075] 3.4 Adding the above three embedding vectors gives the final embedding vector of the government work order, which is expressed as: .
[0076] 4. For the embedding vector E of the work order obtained in step (3), it first passes through a fully connected layer, then activates with the Tanh function, then passes through a Dropout layer, and then passes through a fully connected layer to realize the classification function, output the relative scores of various labels, and finally convert the scores into probabilities through the Softmax function. In the embodiment of the present invention, the output dimension of the last fully connected layer is 2, which mainly distinguishes between government work orders and non-government work orders. The output of the final model is for: ; ; in, 、 、 、 are learnable parameters.
[0077] 5. Calculate the classification loss of the teacher model , where y is the true label of the work order and BCELoss is the binary cross entropy loss function.
[0078] 6. Calculate the teacher model time constraint loss In order to better learn the temporal embedding vector This embodiment adds a time constraint loss, that is, considering that two adjacent time slots (such as 9:00-10:00 and 10:00-11:00 in the morning) should have similar time characteristics, so the corresponding time embedding vectors should also be similar. The present invention uses the following loss function To incorporate the above time characteristics: ; in, represents the time constraint loss, represents the third embedding vector, is the Frobenius norm, is used to constrain the similarity of the third embedding vector corresponding to adjacent generation times; ; 7. Calculate the total loss of the teacher model: ; in, is the total loss of the teacher model, is the weight of the time constraint loss. In this embodiment The value of is 1.
[0079] 8. Fine-tune the teacher model. Use the Adam optimizer to backpropagate the teacher model's loss and update the parameters of the Chinese RoBERTa pre-trained model. In this implementation, the optimizer learning rate is set to 0.001, and the number of model training iterations is set to 50.
[0080] 9. For the training set samples, use the fine-tuned teacher model to predict and obtain the classification score of the teacher model , and then smoothed as a soft label for the student model : ; in, is the temperature coefficient, which is mainly responsible for controlling the degree of distillation of the model. The higher the probability prediction distribution, the flatter it is, and the richer the knowledge information the student model learns from the teacher model. .
[0081] 10. Student model training. After text preprocessing, the training samples are randomly initialized with word vectors. They are first passed through a bidirectional long short-term memory network BiLSTM to obtain the hidden representation h i , then enter an attention layer to get the text representation h s , the attention weight is calculated as: ; Then the text representation h s It is obtained from the following formula: ; Among them, W and b are the parameters that the model needs to learn.
[0082] 11. Calculate the hard loss of the student model . The text representation h s Input a fully connected layer to get the classification score z of the student model student Then, after a softmax layer, the probability distribution of the student model prediction category is obtained , then the hard loss of the student model : ; 12. Calculate the distillation loss of the student model First, calculate the soft labels of the probability distribution predicted by the student model : ; Calculating distillation loss based on KL divergence : ; in, Distribution for teachers and students KL divergence of .
[0083] 13. Calculate the intermediate feature alignment loss of the student model and the teacher model . Fine-tuned teacher model text embedding vector ,but , where cosine is the cosine similarity function.
[0084] 14. Calculate the total loss of the student model : ; Wherein, β and γ are weight coefficients. , .
[0085] 15. Student model training. Use Adam optimizer to optimize the total loss of the student model. Back propagation is performed to update the parameters of the student model. In this implementation example, the optimizer learning rate is set to 0.001 and the number of model iteration training is set to 50.
[0086] 16. Government work order classification: For each newly added government work order, it is input into the trained student model for prediction to determine whether its business type is government-related.
[0087] Corresponding to the above method, the present invention also provides an apparatus / system, which includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the apparatus / system implements the steps of the method described above.
[0088] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements the steps of the aforementioned edge computing server deployment method. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the art.
[0089] In summary, the lightweight work order business type recognition model training method, recognition method, and device described in the present invention, by cleaning and text-enhancing the work order sample data to expand the sample quantity, and based on the word embedding of the text content of the work order sample data, also vectorizes the generation date and generation time to explore the relationship between the work order classification result and the generation time. After training the large-scale parameter teacher model dominated by the linear fully connected layer, the student model dominated by the long short-term memory network and the attention network is trained based on the knowledge distillation method. In the parameter update process, the hard loss of the prediction result, the distillation loss of the difference in the prediction distribution between the student model and the teacher model, and the intermediate feature alignment loss are combined to ensure the model recognition accuracy on the basis of obtaining a lightweight work order business classification recognition model, save computing resources and reduce the demand for hardware capabilities.
[0090] It should be understood by those skilled in the art that the various exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether to implement the system in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention. When implemented in hardware, it may be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via a data signal carried in a carrier wave.
[0091] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.
[0092] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.
[0093] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations to the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A lightweight work order business type recognition model training method, characterized by: The method comprises the following steps: Performing data cleaning on a plurality of work order sample data, wherein the work order sample data includes time information and text content; The sample data of the work order is expanded by using text enhancement, and the work order type is added as a label for the classification task to construct a training sample set; the work order type includes government work orders and non-government work orders; Processing the text content of the work order sample data in the training sample set through a preset word embedding model to obtain a first embedding vector, looking up a table to retrieve a second embedding vector indicating whether it belongs to a working day based on the generation date in the time information, and looking up a table to retrieve a third embedding vector indicating the formation time period based on the generation time in the time information; concatenating the first embedding vector, the second embedding vector, and the third embedding vector into an overall embedding vector; Using the training sample set to train a teacher model dominated by a linear fully connected layer, the teacher model takes the overall embedding vector as input and outputs a first prediction value for the classification task, constructs a first classification loss based on the deviation between the first prediction value and the label, and updates the parameters of the teacher model in combination with a time constraint loss; The training sample set is used to train the student model dominated by the long short-term memory network and the attention network. The student model takes the overall embedding vector as input and outputs a second prediction value of the classification task. A second classification loss is constructed based on the deviation between the second prediction value and the label. The parameters of the student model are updated by combining the distillation loss of the difference in prediction distribution between the student model and the updated teacher model and the intermediate feature alignment loss to obtain a work order business type recognition model.
2. The lightweight work order business type recognition model training method according to claim 1 is characterized in that: The data cleaning includes: removing special characters, invalid content, repeated text, Chinese word segmentation and stop words; The sample data of the work order is expanded by using text enhancement, including: Based on the work order sample data, the sample is expanded by using synonym replacement and sentence transformation based on HowNet synonym dictionary; By constructing guide words, a preset large language model is used to generate similar sample expansion based on the work order sample data, and the large language model adopts ChatGPT or Deepseek; The word embedding model is the RoBERTa model.
3. The lightweight work order business type recognition model training method according to claim 1 is characterized in that: The teacher model includes a first fully connected layer, an activation function layer, a Dropout layer, a second fully connected layer and a first softmax layer; the activation function layer adopts the Tanh function; The student model includes a bidirectional long short-term memory network, an attention layer, a third fully connected layer and a second softmax layer.
4. The lightweight work order business type recognition model training method according to claim 3 is characterized in that: The output of the teacher model is: ; ; in, represents the first predicted value; E represents the overall embedding vector, 、 are the parameters of the first fully connected layer, 、 is the parameter of the second fully connected layer; z represents the output of the second fully connected layer; Then the calculation formula for the first classification loss is: ; in, represents the first classification loss, BCELoss is the binary cross entropy loss function; y represents the label; The calculation formula of the time constraint loss is: ; in, represents the time constraint loss, represents the third embedding vector, is the Frobenius norm, is used to constrain the similarity of the third embedding vector corresponding to adjacent generation times; ; The total loss calculation formula of the first classification loss combined with the time constraint loss in the method is: ; in, is the total loss of the teacher model, is the weight of the time constraint loss.
5. The lightweight work order business type recognition model training method according to claim 3 is characterized in that: The calculation formula for the second classification loss is: ; in, represents the second classification loss; represents the second predicted value; y represents the label; The calculation formula of the distillation loss is: ; in, represents the distillation loss, is the output probability distribution of the second fully connected layer in the teacher model, represents the output probability distribution of the third fully connected layer in the student model; KL represents the calculation of KL divergence; is a fixed parameter; The calculation formula of the intermediate feature alignment loss is: ; in, represents the intermediate feature alignment loss, represents the output features of the Dropout layer in the teacher model, represents the student model; In the method, the total loss calculation formula of the second classification loss combined with the distillation loss and the intermediate feature alignment loss is: ; Among them, β and γ are weight coefficients.
6. The lightweight work order business type recognition model training method according to claim 3 is characterized in that: The method also uses an Adma optimizer to update parameters of the teacher model and the student model.
7. A method for identifying the business type of a work order, characterized in that: The method comprises the following steps: Acquire work order data to be identified, the work order data including time information and text content; Processing the text content of the work order data through a preset word embedding model to obtain a first embedding vector, looking up a table to retrieve a second embedding vector indicating whether it belongs to a working day based on the generation date in the time information, and looking up a table to retrieve a third embedding vector indicating the formation time period based on the generation time in the time information; and concatenating the first embedding vector, the second embedding vector, and the third embedding vector into an overall embedding vector; The overall embedding vector is input into the work order business type recognition model in the lightweight work order business type recognition model training method described in any one of claims 1 to 6, and the recognition result of the work order category is output.
8. A device for identifying a work order business type, comprising a processor, a memory, and a computer program / instruction stored in the memory, characterized in that: The processor is configured to execute the computer program / instructions. When the computer program / instructions are executed, the device implements the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Intelligent sensing method for non-cooperative target component, electronic equipment and storage medium
CN114359651A
Domain knowledge utilization system, domain knowledge utilization method, and domain knowledge utilization program
US20250165814A1