Work order auditing method and device, electronic equipment and readable storage medium

By combining the characteristics of the general pre-trained language model and the domain knowledge pre-trained language model, the problem of limited performance in the telecommunications operation and maintenance field is solved, and the accuracy of the prediction of suspended audit scenarios and the reliability of the automated audit process is achieved.

CN120012772APending Publication Date: 2025-05-16CHINA TELECOM CORP LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411898116.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the field of telecommunications operations and maintenance, text classification algorithms based on fine-tuning of pre-trained language models are difficult to effectively deal with various suspending reasons, resulting in limited effectiveness in automatic work order review.

Method used

By obtaining the knowledge base in the communication field and specifying work tickets, the feature fusion of the general pre-trained language model and the domain knowledge pre-trained language model is used to determine the target review scenario category to which the work ticket belongs.

Benefits of technology

The model's prediction accuracy of suspended audit scenarios is improved, and the efficiency and reliability of the automated suspended audit process is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012772A_ABST
    Figure CN120012772A_ABST
Patent Text Reader

Abstract

The invention provides a work order auditing method and device, electronic equipment and a readable storage medium, and the method comprises the steps: obtaining a knowledge base and a specified work order related to the communication field, inputting the specified work order into a preset universal pre-training language model, obtaining a first hidden layer feature of the specified work order in a hidden layer of the pre-training language model, retrieving the key entity words in a knowledge base, obtaining target item knowledge base entries related to the key entity words, inputting the target item knowledge base entries into a domain knowledge pre-training language model pre-trained through the knowledge base, and obtaining a domain knowledge pre-training language model; according to the method, the first hidden layer feature of the target item knowledge base item in the hidden layer of the domain knowledge pre-training language model is acquired, the second hidden layer feature of the target item knowledge base item in the hidden layer of the domain knowledge pre-training language model is acquired, the first hidden layer feature and the second hidden layer feature are fused to acquire the target hidden layer feature, and the target auditing scene category to which the specified work order belongs is determined through the target hidden layer feature, so that the auditing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing, and in particular relates to a work order review method, device, electronic equipment and readable storage medium. Background Art

[0002] With the rise of a new round of scientific and technological revolution, network communication, as the core technology of "new infrastructure", is driving the continuous expansion of operators' network scale. In particular, the rapid development and widespread application of 5G communication technology have significantly increased the complexity of network architecture. Against this background, the number of operator network operation and maintenance work orders continues to grow, which puts higher requirements on monitoring and maintenance work. In order to ensure the stable operation of the network, the network monitoring and maintenance department needs to complete fault handling and apply for settlement within the prescribed time period according to the different fault levels. However, some work orders involving special processing requirements often need to go through a temporary suspension process to avoid failure to complete the settlement within the prescribed time limit, and conduct compliance checks on the reasons for the suspension. This process aims to avoid the risk of decision-making errors and ensure the quality and stability of user service experience.

[0003] The traditional work order pending review process relies on manual execution, and work order review is usually conducted by designated teams or personnel. However, this manual review model has problems such as low efficiency and inconsistent review standards. In addition, personal subjective judgment and experience differences lead to a high misjudgment rate, difficulty in quantifying review results, and uneven review quality. In addition, with the transformation and upgrading of business, the complexity of network operation and maintenance work continues to increase. Traditional work order pending review methods can no longer support the current operation and maintenance work needs, so the transformation of operation and maintenance models to intelligent ones has become an inevitable trend.

[0004] In recent years, the rise of artificial intelligence deep learning technology has provided a new idea for intelligent operation and maintenance models. The intelligent suspension review process can liberate a large amount of human resources and significantly improve the efficiency of network operation and maintenance. However, building a high-precision suspension review model still faces challenges. Although the text classification algorithm fine-tuned based on the pre-trained language model can realize the compliance judgment of suspension actions, in the field of telecommunications operation and maintenance, due to the lack of professional domain knowledge, its effectiveness is limited when dealing with various suspension reasons (such as disaster risks, city power outages, equipment updates, etc.). Summary of the invention

[0005] The present invention provides a work order review method, device, electronic device and readable storage medium to solve the problem that when a telecommunications work order is currently automatically suspended for review, a text classification algorithm fine-tuned based on a pre-trained language model can realize compliance judgment of the suspension action, but in the field of telecommunications operation and maintenance, due to the lack of professional field knowledge, its performance is limited when dealing with various suspension reasons (such as disaster risks, city power outages, equipment updates, etc.).

[0006] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:

[0007] In a first aspect, the present invention provides a work order review method, the method comprising:

[0008] Acquire a knowledge base and a designated work order in the field of communications, wherein the knowledge base includes a plurality of knowledge base entries, and the designated work order includes a plurality of key entity words;

[0009] Inputting the specified work order into a preset universal pre-trained language model, and obtaining a first hidden layer feature of the specified work order in a hidden layer of the pre-trained language model;

[0010] Searching the key entity word in the knowledge base to obtain a target item knowledge base entry related to the key entity word;

[0011] Inputting the target item knowledge base entry into a domain knowledge pre-trained language model pre-trained by the knowledge base, and obtaining a second hidden layer feature of the target item knowledge base entry in a hidden layer of the domain knowledge pre-trained language model;

[0012] Obtaining a target hidden layer feature by fusing the first hidden layer feature and the second hidden layer feature;

[0013] The target audit scenario category to which the designated work order belongs is determined through the target hidden layer features.

[0014] Optionally, before acquiring the knowledge base and designated work order in the communication field, the method further includes:

[0015] Get some original work orders in the field of communications;

[0016] For any of the original work orders, entity words, entity descriptions of the entity words, and entity types are extracted from the original work order through a preset large language model and a preset prompt engineering framework;

[0017] The entity word, the entity description about the entity word and the entity type are stored in a preset format to obtain a knowledge base entry about the entity word.

[0018] Optionally, after acquiring the knowledge base and designated work order in the communication field, the method further includes:

[0019] For any knowledge base entry, obtaining a target entity word from the knowledge base entry;

[0020] Replacing pre-selected tokens in the target entity word with mask tokens;

[0021] Inputting the knowledge base entry after the replacement operation into a domain knowledge pre-trained language model to obtain a first sentence-level vector;

[0022] The first sentence-level vector is projected into a preset vocabulary space to obtain a prediction result about the pre-selected word.

[0023] Optionally, after projecting the first sentence-level vector to obtain a prediction result about the pre-selected word, the method further includes:

[0024] Determining a prediction result of the pre-selected word-unit and a first loss value of the pre-selected word-unit by using a preset loss function;

[0025] If the first loss value is greater than a preset value, updating the model parameters in the domain knowledge pre-trained language model;

[0026] Obtaining a prediction result of the pre-selected word unit and a second loss value of the pre-selected word unit by using the domain knowledge pre-trained language model after model parameter update;

[0027] If the second loss value is less than a preset value, the updated model parameters are determined as target model parameters in the domain knowledge pre-trained language model.

[0028] Optionally, the step of fusing the first hidden layer feature with the second hidden layer feature to obtain a target hidden layer feature includes:

[0029] Using the first hidden layer feature as a query condition in an attention mechanism, and using the second hidden layer feature as a key and a value in the attention mechanism;

[0030] For any hidden layer, the query condition and the key are compared to obtain a comparison result;

[0031] Determining the attention weight of the hidden layer according to the comparison result, wherein different comparison results correspond to different attention weights;

[0032] extracting a target feature from the value by using the attention weight;

[0033] The target feature is fused with the first hidden layer feature to obtain a target hidden layer feature.

[0034] Optionally, before determining the target audit scenario category to which the designated work order belongs through the target hidden layer features, the method further includes:

[0035] Get several different pending review scenario categories;

[0036] Obtain the number and feature vector of the pending review scenario categories;

[0037] A multidimensional mapping space is generated through the number of suspended audit scenario categories and feature vectors, wherein the dimension of the multidimensional mapping space is determined by the number of suspended audit scenario categories.

[0038] Optionally, determining the target review scenario category to which the designated work order belongs by using the target hidden layer features includes:

[0039] Obtaining a second sentence-level vector of the target hidden layer feature through a preset operator;

[0040] Projecting the second sentence-level vector to the multidimensional mapping space, obtaining a probability value that the specified work order belongs to the pending review scenario category;

[0041] If any of the probability values ​​is greater than a preset threshold, the pending review scenario category corresponding to the probability value is determined as the target review scenario category to which the designated work order belongs;

[0042] If the probability values ​​are all smaller than the preset thresholds, it is determined that the designated work order belongs to the non-pending review scenario category.

[0043] In a second aspect, the present invention provides a work order review device, the device comprising:

[0044] A first acquisition module is used to acquire a knowledge base and a designated work order in the field of communications, wherein the knowledge base includes a plurality of knowledge base entries, and the designated work order includes a plurality of key entity words;

[0045] A second acquisition module is used to input the specified work order into a preset universal pre-trained language model to obtain a first hidden layer feature of the specified work order in a hidden layer of the pre-trained language model;

[0046] A third acquisition module is used to search the key entity word in the knowledge base to obtain a target item knowledge base entry related to the key entity word;

[0047] A fourth acquisition module is used to input the target item knowledge base entry into a domain knowledge pre-trained language model pre-trained by the knowledge base, and obtain a second hidden layer feature of the target item knowledge base entry in a hidden layer of the domain knowledge pre-trained language model;

[0048] A fifth acquisition module, configured to acquire a target hidden layer feature by fusing the first hidden layer feature and the second hidden layer feature;

[0049] The first determination module is used to determine the target review scenario category to which the specified work order belongs based on the target hidden layer features.

[0050] In a third aspect, the present invention provides an electronic device, comprising: a transceiver, a memory, a processor, and a program stored in the memory and executable on the processor;

[0051] The processor is used to read the program in the memory to implement any of the work order review methods described above.

[0052] In a fourth aspect, the present invention provides a computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium, and when the computer-readable storage medium is run on a computer, the computer executes any of the above-mentioned work order review methods.

[0053] In the present application, a knowledge base and a designated work order regarding the field of communications are obtained, wherein the knowledge base includes a plurality of knowledge base entries, and the designated work order includes a plurality of key entity words. The designated work order is input into a preset general pre-trained language model to obtain the first hidden layer feature of the designated work order in the hidden layer of the pre-trained language model; the key entity words are retrieved in the knowledge base to obtain the target item knowledge base entry regarding the key entity words; the target item knowledge base entry is input into a domain knowledge pre-trained language model pre-trained by the knowledge base to obtain the second hidden layer feature of the target item knowledge base entry in the hidden layer of the domain knowledge pre-trained language model; the target hidden layer feature is obtained by fusing the first hidden layer feature and the second hidden layer feature, and the feature fusion of the two models effectively promotes the knowledge transfer and feature fusion between the two models, realizes cross-model knowledge injection, and further enhances the model's ability to understand and apply knowledge in specific domains. The target hidden layer feature is used to determine the target review scenario category to which the designated work order belongs, effectively improving the model's prediction accuracy for the suspended review scenario, thereby ensuring the efficiency and reliability of the automated suspended review process. Based on the use of a general pre-trained language model, this application combines the use of a domain knowledge pre-trained language model that is pre-trained based on a knowledge base in the communications field to perform suspension review on designated work orders, thereby avoiding the lack of professional domain knowledge in the field of telecommunications operation and maintenance, which leads to limited performance when dealing with various suspension reasons (such as disaster risks, city power outages, equipment updates, etc.), and effectively improves the prediction accuracy of suspension review scenarios, thereby ensuring the efficiency and reliability of the automated suspension review process. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the present application or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0055] Figure 1 It is a step flow chart of a work order review method provided by this application;

[0056] Figure 2 It is a flow chart of another work order review method provided by this application;

[0057] Figure 3 yes Figure 1 The flowchart of step 105 in a work order review method provided by the present application is shown;

[0058] Figure 4 yes Figure 1 The flowchart of step 106 in a work order review method provided by the present application is shown;

[0059] Figure 5 It is a structural diagram of a work order review device provided by this application;

[0060] Figure 6 It is a structural diagram of an electronic device provided by this application. DETAILED DESCRIPTION

[0061] The following will be combined with the drawings in this application to clearly and completely describe the technical solutions in this application. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0062] With the transformation and upgrading of business, the complexity of network operation and maintenance work continues to increase. The traditional work order suspension review mechanism can no longer meet the needs of modern operation and maintenance efficiency and accuracy. The transformation of operation and maintenance mode to intelligence has become an inevitable trend. Among them, the application of artificial intelligence deep learning technology has opened up new possibilities. However, although the fine-tuning text classification algorithm based on PLM can realize the compliance judgment of suspension actions, in the field of telecommunications operation and maintenance, due to the lack of professional domain knowledge, its effectiveness is limited when dealing with various reasons for suspension (such as disaster risks, city power outages, equipment updates, etc.). In order to help the model acquire domain knowledge, domain knowledge enhancement based on knowledge graphs is a commonly used technical means. However, due to the complexity of entity relationships in actual applications, the construction of knowledge graphs is not easy, and it is difficult to become an efficient external knowledge base construction method to improve model performance. Based on this, the present application proposes a work order review method.

[0063] Before understanding this application, the technical terms related to this application are introduced as follows:

[0064] PLM (Pre-trained Language Model) is a deep learning model that uses a large amount of corpus for unsupervised pre-training, and can then be used for various downstream tasks such as text classification, named entity recognition, sentiment analysis, etc. Currently, popular pre-trained language models include BERT (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), RoBERTa (Robustly Optimized BERTPretraining Approach), etc.

[0065] DK-PLM (Domain Knowledge Pre-trained Language Model) is a language model that is pre-trained for specific domain knowledge to enhance the model's ability to understand and apply domain knowledge. This model can better capture and represent the semantic structure of a specific domain, thereby showing higher accuracy and professionalism when processing domain-related tasks.

[0066] LLM (Large Language Model) is an artificial intelligence model designed to understand and generate human language. Compared with PLM, it can reach billions or trillions of parameter scales. It learns a lot of language knowledge during the training process, so it has high flexibility and scalability, and its reasoning ability is excellent. Currently, popular open-source and closed-source large language models include chat-GPT, chat-GLM, Vicuna, qWen2, llama3.1 series, etc.

[0067] Based on the understanding of the above technical terms, refer to Figure 1 , Figure 1 This is a flowchart of the steps of a work order review method provided by this application, such as Figure 1 As shown, the method may include:

[0068] Step 101, obtaining a knowledge base and a designated work order in the field of communications, wherein the knowledge base includes a plurality of knowledge base entries, and the designated work order includes a plurality of key entity words.

[0069] In the modern network operation and maintenance system, the work order system, as the hub of information flow, carries the key fault diagnosis and repair tasks. The work order system includes various types of uploaded work orders. Among them, the core elements in the work order content - entity information such as alarm name, fault cause and solution, can be used as an important part of the domain knowledge base. However, the way to build such a materialized knowledge base usually requires the in-depth participation of a large number of experts in the field, and to parse, record and strictly verify the entities one by one. This process is extremely arduous and consumes huge manpower costs. Therefore, in this application, through a large language model, combined with a prompt engineering framework, entity words, entity descriptions and entity types are extracted from the original work order, and then a simple review by domain technical experts can be conducted, so that a knowledge base about the communication field can be efficiently constructed. Among them, the large language model can be LLM, and the prompt engineering framework is used to guide the large language model on how to extract entity words, entity descriptions and entity types from the original work order. The prompt engineering framework can include specific prompt templates, instructions and rules. For example, "You are a network operation and maintenance expert. Please provide me with a detailed description of the following professional terms: 1.XX, 2XX..." The specific steps include:

[0070] Get some original work orders in the field of communications;

[0071] For any original work order, the preset large language model and preset prompt engineering framework are used to extract entity words, entity descriptions and entity types from the original work order.

[0072] The entity word, the entity description about the entity word and the entity type are stored in a preset format to obtain a knowledge base entry about the entity word.

[0073] The preset format may be a json format. A knowledge base entry in the knowledge base is a structured entity word, entity description and entity type triple. Then the knowledge base entry in the knowledge base is stored in the following general form:

[0074] 1.{entity1:string,description1:string,type1:string}

[0075] 2.{entity2:string,description2:string,type2:string}

[0076] …

[0077] n.{entityN:string,descriptionN:string,typeN:string}

[0078] Among them, entity is the entity word, description is the entity description of the entity word generated by the large language model, and type is the entity type of the entity word.

[0079] For example, if the combination of the prompt engineering framework and the output text is "You are a network operation and maintenance expert. Please provide me with a detailed description of the following professional terms: 1. Transmission relay failure, 2. City power outage, 3. Replacement of optical module." By organizing it through the large language model, the following form of knowledge base data can be constructed:

[0080] {

[0081] "data":[

[0082] {"entity":"Transmission relay failure",

[0083] "description":"Transmission relay failure refers to the situation in which data transmission fails or quality deteriorates due to problems with relay equipment during network transmission.",

[0084] "type":"Alarm Name"},

[0085] {"entity":"Power outage",

[0086] "description":"A city power outage refers to a situation where the city power grid stops supplying power due to various reasons (planned or unplanned).",

[0087] "type":"Failure reason"},

[0088] {"entity":"Optical cable interruption",

[0089] "description":"Cable interruption refers to the interruption of data transmission caused by physical damage to the optical cable line or other reasons.",

[0090] "type":"Failure reason"},

[0091] {"entity":"Replace optical module",

[0092] "description":"Replacing an optical module refers to the operation of replacing the optical transceiver module used to send and receive optical signals in network devices (such as routers, switches, etc.).",

[0093] "type":"Solution"}, ...... ]

[0096] }

[0097] By utilizing advanced large language models combined with a prompt engineering framework, entity definitions and descriptions can be constructed quickly and efficiently, greatly reducing the need for manual participation and significantly reducing labor costs. At the same time, the constructed knowledge base can help the model better understand and process work order information in the communications field during subsequent model training, thereby improving the accuracy of model recognition.

[0098] Step 102: input the specified work order into a preset universal pre-trained language model to obtain the first hidden layer feature of the specified work order in the hidden layer of the pre-trained language model.

[0099] The universal pre-trained language model in the embodiment of the present invention can be a PLM, and the designated work order is a work order during network operation and maintenance. There can be multiple ones. For example, there are 3 designated work orders, and the content of designated work order 1 is "due to heavy rain affecting the city power outage causing the alarm, apply for a 4-hour suspension order for network inspection", and the content of designated work order 2 is "after inspection, the cause of the over-temperature alarm is the damage to the air-conditioning compressor. The manufacturer has been contacted for on-site replacement, and a temporary suspension order of 48 hours is applied", and the content of designated work order 3 is "check that the device port is normal, no operation has been performed, the alarm information needs to be verified, and an application for suspension order is applied for observation until 8 pm".

[0100] It should be noted that each designated work order is input into the preset universal pre-trained language model, and the first hidden layer features corresponding to the work order will be obtained. In order to facilitate identification, a work order ID can be set for each designated work order. When extracting features, this work order ID will be added to the acquired features. The universal pre-trained language model performs semantic encoding on the input designated work order, which processes information from the perspective of language understanding.

[0101] Step 103: search the knowledge base for the key entity word to obtain a target knowledge base entry related to the key entity word.

[0102] The embodiment of the present invention uses a general pre-trained language model to extract features from a specified work order. However, in the field of telecommunications operation and maintenance, the general pre-trained language model lacks professional domain knowledge, which results in limited effectiveness when dealing with various reasons for suspension (such as disaster risks, city power outages, equipment updates, etc.). Therefore, a domain knowledge pre-trained language model is also introduced. The domain knowledge pre-trained language model is trained based on the content of a knowledge base constructed in the field of communications. Therefore, it is necessary to use an entity extraction algorithm to extract key entity words from the specified work order, and then extract entity descriptions and entity types related to the specified work order from the knowledge base based on the key entity words, and input them into the domain knowledge pre-trained language model according to the data organization form of the content in the knowledge base for understanding and encoding domain-specific knowledge.

[0103] Step 104 , input the target item knowledge base entry into the domain knowledge pre-trained language model pre-trained by the knowledge base, and obtain the second hidden layer feature of the target item knowledge base entry in the hidden layer of the domain knowledge pre-trained language model.

[0104] In the embodiment of the present invention, after finding the corresponding target item knowledge base entry from the knowledge base according to the key entity words of the specified work order, the target item knowledge base entry is input into the domain knowledge pre-trained language model pre-trained by the knowledge base to process the information from the perspective of professional domain knowledge. Among them, the domain knowledge pre-trained language model needs to be pre-trained. When training the domain knowledge pre-trained language model, multiple knowledge base entries are obtained from the knowledge base about the communication field for training. However, during training, one knowledge base entry is trained at a time, and after the training is completed, the next knowledge base entry is selected for training. During the pre-training process, any knowledge base entry is denoted as s i , which is designed to include entity words, entity descriptions, and entity type triples. Its specific form is as follows:

[0105] s i ={entity i ,description i ,type i}

[0106] Among them, entity i Refers to entity words, description i Is the entity description, type i It is the entity type. First, the masked language modeling technology (MLM) is used to build the contextual relationship of the domain knowledge pre-trained language model, that is, pre-selecting words from the entity words, and then replacing the pre-selected words with [MASK] mask tags. For example, after replacement, entity i ={w 1 ,[MASK],...,w n}, the knowledge base entries after the substitution operation are input into the domain knowledge pre-trained language model, and then the first sentence-level vector is generated through the AvgPool operation. i represents the first sentence-level vector, x i =AvgPool(DK-PLM(s i )), where DK-PLM is a domain knowledge pre-trained language model, DK-PLM(s i ) is the domain knowledge pre-trained language model pair s i The output result of AvgPool is a pooling operation. Then x iProjecting into K-dimensional space, where K corresponds to the size of the vocabulary, obtains prediction results for pre-selected tokens. The specific steps include:

[0107] For any knowledge base entry, obtain the target entity word from the knowledge base entry;

[0108] Replace the pre-selected tokens in the target entity word with mask tokens;

[0109] Input the knowledge base entry after the replacement operation into the domain knowledge pre-trained language model to obtain the first sentence level vector;

[0110] The first sentence-level vector is projected into a preset vocabulary space to obtain prediction results about pre-selected word units.

[0111] Among them, when the first sentence-level vector is projected to the preset vocabulary space, a linear transformation (such as a fully connected layer) is usually used to achieve feature fusion and semantic refinement. The projected vector representation is now in the vocabulary space, and each dimension corresponds to a word in the vocabulary. This vector can be interpreted as the probability distribution of each word in the vocabulary. For example, any knowledge base entry selected is: entity word: Marie Curie, entity description: Polish-born physicistand chemist, entity type: Person; the target entity word is "Marie Curie", assuming that our pre-selected word is "Curie", we replace it with a mask tag, entity word: Marie[MASK], the first sentence-level vector obtained according to the replaced knowledge base entry is [0.234, 0.567,..., 0.890], the first sentence-level vector is projected to the preset vocabulary space, and the prediction results for the pre-selected word "[MASK]" are obtained, 1. Curie (probability: 0.97), 2. (Probability: 0.02), 3. Einstein (Probability: 0.01)...

[0112] Through the above steps, it is possible to use the domain knowledge pre-trained language model to predict the masked words in the knowledge base entries. This not only takes into account the contextual information around the selected words, but also deeply understands the intrinsic relationship between the masked words and the overall text, significantly improves the ability to capture entity semantic details based on entities, and also enhances the automatic completion and reasoning capabilities of the knowledge base.

[0113] Furthermore, in order to quantify the gap between the prediction result of the model and the actual selected word-unit label, a loss function is also introduced in the embodiment of the present invention to explicitly guide the word-unit prediction task, and the parameters in the model are adjusted through the loss function to better fit the training data, thereby improving the performance of the model. When adjusting, it is necessary to first use a preset loss function (such as cross entropy loss) to calculate the difference between the prediction result and the pre-selected word-unit to obtain a first loss value. If the first loss value is greater than a preset threshold, it means that the prediction result of the model is not accurate enough and the parameters need to be updated. At this time, gradient descent or other optimization algorithms can be used to update the model parameters according to the gradient of the loss function, and then the updated model parameters are used to re-predict, and the preset loss function is used again to calculate the difference between the prediction result and the actual label to obtain a second loss value. If the second loss value is less than the preset threshold, it means that the prediction result of the model is accurate enough. The updated model parameters are determined as the target model parameters in the domain knowledge pre-trained language model. The specific steps include:

[0114] Determining a prediction result of a pre-selected word and a first loss value of the pre-selected word by using a preset loss function;

[0115] If the first loss value is greater than a preset value, the model parameters in the domain knowledge pre-trained language model are updated;

[0116] Pre-training the language model with domain knowledge after updating model parameters to obtain a prediction result of a pre-selected word and a second loss value of the pre-selected word;

[0117] If the second loss value is less than the preset value, the updated model parameters are determined as the target model parameters in the domain knowledge pre-trained language model.

[0118] For example, assuming that the word we pre-selected is "Einstein", we use the domain knowledge pre-trained language model to predict "Einstein" and get the prediction result. We use the cross entropy loss function to calculate the difference between the prediction result and the true label "Einstein", and get the first loss value of 0.5. Assuming the preset value is 0.3, because 0.5>0.3, we use the gradient descent algorithm to update the parameters of the model, use the updated model to re-predict "Einstein", and use the cross entropy loss function again to calculate the difference between the prediction result and the true label "Einstein", and get the second loss value of 0.2. Because 0.2<0.3, the updated model parameters are determined as the target model parameters.

[0119] Among them, the loss function can be expressed as:

[0120]

[0121] in, is a linear transformer, x i is the first sentence level vector, y i is the vector of the pre-selected word, K is the size of the preset vocabulary space, and N is the number of prediction results.

[0122] It should be noted that if the second loss value is greater than the preset value, the model parameters in the domain knowledge pre-trained language model continue to be updated until the loss value calculated by the loss function is less than the preset value, and the updating of the model parameters is terminated. The preset value is set according to actual needs, and the present invention does not make specific limitations here.

[0123] Step 105, obtaining target hidden layer features by fusing the first hidden layer features with the second hidden layer features.

[0124] In the embodiment of the present invention, the target item knowledge base entry about the key entity word is obtained from the knowledge base according to the specified work order. After the domain knowledge pre-trained language model is pre-trained, the characteristics of the target item knowledge base entry can be obtained by using the domain knowledge pre-trained language model. In the embodiment of the present invention, the universal pre-trained language model and the domain knowledge pre-trained language model present a dual-branch parallel structure, for example, Figure 2 As shown in the figure, information is processed from the perspectives of general language understanding and professional domain knowledge respectively. Information is processed from the perspectives of general language understanding and professional domain knowledge respectively. The general pre-trained language model directly processes the specified work order. Through the key entity words, the knowledge base entry of the target item is obtained from the knowledge base and input into the domain knowledge pre-trained language model. In order to promote information exchange and fusion between branches, a layer-by-layer Cross-Attention mechanism is introduced. The hidden layer features extracted by the general pre-trained language model are used as queries, while the hidden layer features of the domain knowledge pre-trained language model act as keys and values ​​to achieve cross-model knowledge transfer and injection. In addition, knowledge injection and feature fusion are required in each layer. After the features are fused, they are output through the general pre-trained language model. The output features are called target hidden layer features. The target audit scenario category of the specified work order is determined through the target hidden layer features. Attention Mechanism is an important technology in deep learning, especially in the field of natural language processing. The attention mechanism allows the model to pay more attention to certain specific parts when processing input data, thereby improving the performance of the model.

[0125] It should be noted that the key is used to compare with the query condition, and the value is used to extract the target feature. Therefore, the query condition (Query) is compared with the key (Key), usually through the dot product or the scaling form of the dot product. The comparison result represents the similarity between the query condition and the key. The comparison result is a matrix that represents the similarity between each query condition and each key. The attention weight is calculated based on the comparison result. The attention weight represents the degree of attention paid to different parts when the model processes the input data. The comparison result is usually normalized by the softmax function to obtain the attention weight. The attention weight is used to perform weighted summation on the value (Value) to obtain the target feature. The target feature is the part that the model focuses on when processing the input data. The target feature is fused with the query condition (the first hidden layer feature) to obtain the final hidden layer feature. The fusion can be performed by weighted summation or splicing.

[0126] Further, step 105, such as Figure 3 As shown:

[0127] Step 1051, using the first hidden layer features as query conditions in the attention mechanism, and using the second hidden layer features as keys and values ​​in the attention mechanism.

[0128] Step 1052, for any hidden layer, compare the query condition and the key to obtain a comparison result.

[0129] Step 1053, determining the attention weight of the hidden layer through the comparison results, wherein different comparison results correspond to different attention weights.

[0130] Step 1054, extract the target features from the values ​​through the attention weights.

[0131] Step 1055, fuse the target feature with the first hidden layer feature to obtain the target hidden layer feature.

[0132] For example, assume that the output feature of the first hidden layer is Query = [0.1, 0.2, 0.3, ..., 0.9], the output feature of the second hidden layer is: Key = [0.2, 0.3, 0.4, ..., 1.0], Value = [0.3, 0.4, 0.5, ..., 1.1], perform dot product operation on Query and Key to obtain the comparison result, Dot Product = Query * Key = [0.02, 0.06, 0.12, ..., 0.90], normalize the comparison result through the softmax function, and obtain the attention weight: AttentionWeights = softmax(Dot Product) = [0.1, 0.2, 0.3, ..., 0.9], use the attention weight to perform weighted summation on Value to obtain the target feature: Target Feature = Attention Weights*Value=[0.03,0.08,0.15,...,0.99], perform weighted summation of the target feature and the first hidden layer feature to obtain the target hidden layer feature: Target Hidden Layer Feature=Query+Target Feature=[0.13,0.28,0.45,...,1.89].

[0133] In the above-mentioned progressive fusion strategy, the weights of features at each level are dynamically adjusted, which enables the knowledge injection process to gradually transition from basic features to advanced features, meeting the needs of multi-level feature learning.

[0134] Step 106, determining the target review scenario category to which the specified work order belongs through the target hidden layer features.

[0135] The audit scenario categories in the embodiment of the present invention include suspended audit scenario categories and non-suspended audit scenario categories. In order to better determine the target audit scenario category to which a specified work order belongs, it is necessary to first construct a multidimensional mapping space for different categories of suspended audit scenarios. The construction steps include:

[0136] Get several different pending review scenario categories;

[0137] Get the number and feature vector of pending review scenario categories;

[0138] A multidimensional mapping space is generated by the number of pending review scenario categories and feature vectors, wherein the dimension of the multidimensional mapping space is determined by the number of pending review scenario categories.

[0139] Through the above steps, we can generate a multidimensional mapping space to represent different pending review scenario categories and their feature vectors. This multidimensional mapping space can help us better understand and analyze pending review scenarios, thereby optimizing the review process and improving review efficiency.

[0140] It should be noted that the fused target hidden layer features are output by the general pre-trained language model. After obtaining the target hidden layer features and the multi-dimensional mapping space, the second sentence-level vector is first generated by a preset operator (for example, the AvgPool operator). For example, t' i Indicates a specified work order, using PLM(t' i ) represents the target hidden layer feature, and x i ' represents the second sentence level vector, then x i =AvgPool(PLM(t' i )). The second sentence-level vector is projected into the multidimensional mapping space to obtain the probability value of the specified work order belonging to each pending review scenario category, and each probability value is checked to see whether it is greater than a preset threshold value. The pending review scenario category (text review) corresponding to the probability value of 0.5 is determined as the target review scenario category to which the specified work order belongs. If all probability values ​​are less than the preset threshold value, it is determined that the specified work order belongs to the non-pending review scenario category.

[0141] Further, step 106, such as Figure 4 As shown:

[0142] Step 1061, obtaining the second sentence-level vector of the target hidden layer feature through a preset operator.

[0143] Step 1062, projecting the second sentence-level vector to the multidimensional mapping space, and obtaining the probability value that the specified work order belongs to the pending review scenario category.

[0144] Step 1063: If any probability value is greater than a preset threshold, the suspended review scenario category corresponding to the probability value is determined as the target review scenario category to which the specified work order belongs.

[0145] Step 1064: If the probability values ​​are all less than the preset threshold, it is determined that the specified work order belongs to the non-pending review scenario category.

[0146] It should be noted that if there are multiple probability values ​​greater than the preset threshold, the suspended review scenario category corresponding to the largest probability value is selected to be determined as the target review scenario category to which the specified work order belongs.

[0147] For example, three designated work orders are reviewed. The content of designated work order 1 is "the alarm is caused by the power outage caused by heavy rain, and the order is pending for 4 hours for network inspection", the content of designated work order 2 is "after inspection, the cause of the over-temperature alarm is the damage of the air-conditioning compressor, the manufacturer has been contacted for on-site replacement, and the order is temporarily pending for 48 hours", and the content of designated work order 3 is "check that the device port is normal, no operation has been performed, the alarm information needs to be verified, and the order is pending for observation until 8 pm". Combined with the model prediction results, it can be determined that designated work order 1 and designated work order 2 meet the pending review conditions, and their semantic features are in the pending review scene feature space of "power outage" and "equipment replacement", while designated work order 3 does not extract relevant entity information, and the model output result does not meet the suspension conditions. The sample is effectively identified and recalled, and this work order is determined to be a non-pending review scene category.

[0148] Furthermore, in order to improve the prediction accuracy of the model for the pending review scenario, a loss function is introduced to collaboratively update the learnable parameters of the PLM model and the DK-PLM model through iterations of a large amount of data.

[0149] In the present application, a knowledge base and a designated work order regarding the field of communications are obtained, wherein the knowledge base includes a plurality of knowledge base entries, and the designated work order includes a plurality of key entity words. The designated work order is input into a preset general pre-trained language model to obtain the first hidden layer feature of the designated work order in the hidden layer of the pre-trained language model; the key entity words are retrieved in the knowledge base to obtain the target item knowledge base entry regarding the key entity words; the target item knowledge base entry is input into a domain knowledge pre-trained language model pre-trained by the knowledge base to obtain the second hidden layer feature of the target item knowledge base entry in the hidden layer of the domain knowledge pre-trained language model; the target hidden layer feature is obtained by fusing the first hidden layer feature and the second hidden layer feature, and the feature fusion of the two models effectively promotes the knowledge transfer and feature fusion between the two models, realizes cross-model knowledge injection, and further enhances the model's ability to understand and apply knowledge in specific domains. The target hidden layer feature is used to determine the target review scenario category to which the designated work order belongs, effectively improving the model's prediction accuracy for the suspended review scenario, thereby ensuring the efficiency and reliability of the automated suspended review process. Based on the use of a general pre-trained language model, this application combines the use of a domain knowledge pre-trained language model that is pre-trained based on a knowledge base in the communications field to perform suspension review on designated work orders, thereby avoiding the lack of professional domain knowledge in the field of telecommunications operation and maintenance, which leads to limited performance when dealing with various suspension reasons (such as disaster risks, city power outages, equipment updates, etc.), and effectively improves the prediction accuracy of suspension review scenarios, thereby ensuring the efficiency and reliability of the automated suspension review process.

[0150] Figure 5This is a structural diagram of a work order review device provided by the present application, which may include:

[0151] The first acquisition module 201 is used to acquire a knowledge base and a designated work order in the communication field, wherein the knowledge base includes a plurality of knowledge base entries, and the designated work order includes a plurality of key entity words.

[0152] The second acquisition module 202 is used to input the specified work order into a preset universal pre-trained language model to obtain the first hidden layer feature of the specified work order in the hidden layer of the pre-trained language model.

[0153] The third acquisition module 203 is used to search the key entity word in the knowledge base to obtain the target item knowledge base entry related to the key entity word.

[0154] The fourth acquisition module 204 is used to input the target item knowledge base entry into the domain knowledge pre-trained language model pre-trained by the knowledge base, and obtain the second hidden layer feature of the target item knowledge base entry in the hidden layer of the domain knowledge pre-trained language model.

[0155] The fifth acquisition module 205 is used to acquire the target hidden layer feature by fusing the first hidden layer feature and the second hidden layer feature.

[0156] The first determination module 206 is used to determine the target review scenario category to which the specified work order belongs through target hidden layer features.

[0157] Optionally, the work order review device further includes:

[0158] The sixth acquisition module is used to acquire several original work orders in the communication field.

[0159] The extraction module is used to extract entity words, entity descriptions of entity words and entity types from any original work order through a preset large language model and a preset prompt engineering framework.

[0160] The storage module is used to store the entity word, the entity description about the entity word and the entity type according to a preset format to obtain a knowledge base entry about the entity word.

[0161] The seventh acquisition module is used to acquire a target entity word from any knowledge base entry.

[0162] The replacement module is used to replace the pre-selected tokens in the target entity word with mask tokens.

[0163] The eighth acquisition module is used to input the knowledge base entry after the replacement operation into the domain knowledge pre-trained language model to obtain the first sentence level vector.

[0164] The ninth acquisition module is used to project the first sentence-level vector to a preset vocabulary space to obtain a prediction result about a pre-selected word.

[0165] The second determination module is used to determine the prediction result of the pre-selected word and the first loss value of the pre-selected word by using a preset loss function.

[0166] The parameter updating module is used to update the model parameters in the domain knowledge pre-trained language model if the first loss value is greater than a preset value.

[0167] The tenth acquisition module is used to obtain the prediction result of the pre-selected word unit and the second loss value of the pre-selected word unit through the domain knowledge pre-training language model after the model parameters are updated.

[0168] The third determination module is used to determine the updated model parameters as the target model parameters in the domain knowledge pre-trained language model if the second loss value is less than a preset value.

[0169] Optionally, the fifth acquisition module 205 specifically includes:

[0170] Set up submodules to use the first hidden layer features as query conditions in the attention mechanism, and the second hidden layer features as keys and values ​​in the attention mechanism.

[0171] The first acquisition submodule is used to compare the query condition and the key for any hidden layer to obtain the comparison result.

[0172] The first determination submodule is used to determine the attention weight of the hidden layer through the comparison results, wherein different comparison results correspond to different attention weights.

[0173] The extraction submodule is used to extract target features from the values ​​through attention weights.

[0174] The second acquisition submodule is used to fuse the target feature with the first hidden layer feature to obtain the target hidden layer feature.

[0175] Optionally, the work order review device further includes:

[0176] The eleventh acquisition module is used to obtain several different categories of suspended review scenarios.

[0177] The twelfth acquisition module is used to obtain the number and feature vector of pending review scenario categories.

[0178] A generation module is used to generate a multidimensional mapping space according to the number of pending review scenario categories and feature vectors, wherein the dimension of the multidimensional mapping space is determined by the number of pending review scenario categories.

[0179] Optionally, the first determining module 206 specifically includes:

[0180] The third acquisition submodule is used to obtain the second sentence-level vector of the target hidden layer feature through a preset operator.

[0181] The fourth acquisition submodule is used to project the second sentence-level vector to the multidimensional mapping space to obtain the probability value that the specified work order belongs to the pending review scenario category.

[0182] The second determination submodule is used to determine the suspended review scenario category corresponding to the probability value as the target review scenario category to which the specified work order belongs if any probability value is greater than a preset threshold.

[0183] The third determination submodule is used to determine that the specified work order belongs to the non-pending review scenario category if the probability values ​​are all less than a preset threshold.

[0184] In the present application, a knowledge base and a designated work order regarding the field of communications are obtained, wherein the knowledge base includes a plurality of knowledge base entries, and the designated work order includes a plurality of key entity words. The designated work order is input into a preset general pre-trained language model to obtain the first hidden layer feature of the designated work order in the hidden layer of the pre-trained language model; the key entity words are retrieved in the knowledge base to obtain the target item knowledge base entry regarding the key entity words; the target item knowledge base entry is input into a domain knowledge pre-trained language model pre-trained by the knowledge base to obtain the second hidden layer feature of the target item knowledge base entry in the hidden layer of the domain knowledge pre-trained language model; the target hidden layer feature is obtained by fusing the first hidden layer feature and the second hidden layer feature, and the feature fusion of the two models effectively promotes the knowledge transfer and feature fusion between the two models, realizes cross-model knowledge injection, and further enhances the model's ability to understand and apply knowledge in specific domains. The target hidden layer feature is used to determine the target review scenario category to which the designated work order belongs, effectively improving the model's prediction accuracy for the suspended review scenario, thereby ensuring the efficiency and reliability of the automated suspended review process. Based on the use of a general pre-trained language model, this application combines the use of a domain knowledge pre-trained language model that is pre-trained based on a knowledge base in the communications field to perform suspension review on designated work orders, thereby avoiding the lack of professional domain knowledge in the field of telecommunications operation and maintenance, which leads to limited performance when dealing with various suspension reasons (such as disaster risks, city power outages, equipment updates, etc.), and effectively improves the prediction accuracy of suspension review scenarios, thereby ensuring the efficiency and reliability of the automated suspension review process.

[0185] The present invention also provides an electronic device, such as Figure 6As shown, it includes a processor 301, a communication interface 302, a memory 303 and a communication bus 304, wherein the processor 301, the communication interface 302, and the memory 303 communicate with each other through the communication bus 304.

[0186] Memory 303, used for storing computer programs;

[0187] The processor 301 is used to execute the program stored in the memory 303 to implement the following steps:

[0188] Acquire a knowledge base and a designated work order in the field of communications, wherein the knowledge base includes a plurality of knowledge base entries, and the designated work order includes a plurality of key entity words;

[0189] Inputting the specified work order into a preset universal pre-trained language model, and obtaining a first hidden layer feature of the specified work order in a hidden layer of the pre-trained language model;

[0190] Searching the key entity word in the knowledge base to obtain a target item knowledge base entry related to the key entity word;

[0191] Inputting the target item knowledge base entry into a domain knowledge pre-trained language model pre-trained by the knowledge base, and obtaining a second hidden layer feature of the target item knowledge base entry in a hidden layer of the domain knowledge pre-trained language model;

[0192] Obtaining a target hidden layer feature by fusing the first hidden layer feature and the second hidden layer feature;

[0193] The target audit scenario category to which the designated work order belongs is determined through the target hidden layer features.

[0194] The communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into a physical address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0195] The communication interface is used for communication between the above terminal and other devices.

[0196] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0197] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0198] The present invention also provides a readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the work order review method of the aforementioned embodiment.

[0199] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0200] The algorithm and display provided herein are not inherently related to any particular computer, virtual system or other equipment. According to the above description, it is obvious that the structure required for constructing such a system is. In addition, the present invention is not directed to any specific programming language either. It should be understood that various programming languages ​​can be utilized to realize the content of the present invention described herein, and the description of the above specific language is for disclosing the best mode of the present invention.

[0201] In the description provided herein, a large number of specific details are described. However, it is understood that embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.

[0202] Similarly, it should be understood that in order to streamline the present invention and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of the present invention, various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, this disclosed method should not be interpreted as reflecting the following intention: that the claimed invention requires more features than those expressly recited in each claim. Rather, as reflected in the claims below, inventive aspects lie in less than all of the features of the individual embodiments disclosed above. Therefore, the claims that follow the detailed description are hereby expressly incorporated into the detailed description, with each claim itself serving as a separate embodiment of the present invention.

[0203] Those skilled in the art will appreciate that the modules in the devices in the embodiments may be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments may be combined into one module or unit or component, and in addition they may be divided into a plurality of submodules or subunits or subcomponents. Except that at least some of such features and / or processes or units are mutually exclusive, all features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed in this manner may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature providing the same, equivalent or similar purpose.

[0204] The various component embodiments of the present invention may be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. It should be understood by those skilled in the art that a microprocessor or digital signal processor (DSP) may be used in practice to implement some or all of the functions of some or all of the components in the sorting device according to the present invention. The present invention may also be implemented as a device or apparatus program for executing part or all of the methods described herein. Such a program for implementing the present invention may be stored on a computer-readable medium, or may be in the form of one or more signals. Such a signal may be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0205] It should be noted that the above embodiments illustrate the present invention rather than limit it, and that those skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbol between brackets shall not be construed as a limitation on the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "one" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention may be implemented by means of hardware comprising a number of different elements and by means of a suitably programmed computer. In a unit claim enumerating a number of devices, several of these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc., does not indicate any order. These words may be interpreted as names.

[0206] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0207] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

[0208] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

[0209] It should be noted that the various data-related processes in the embodiments of the present application are all carried out in compliance with the corresponding data protection laws and policies of the country where the device is located, and with the authorization given by the owner of the corresponding device.

Claims

1. A work order review method, the method comprising: Acquire a knowledge base and a designated work order in the field of communications, wherein the knowledge base includes a plurality of knowledge base entries, and the designated work order includes a plurality of key entity words; Inputting the specified work order into a preset universal pre-trained language model, and obtaining a first hidden layer feature of the specified work order in a hidden layer of the pre-trained language model; Searching the key entity word in the knowledge base to obtain a target item knowledge base entry related to the key entity word; Inputting the target item knowledge base entry into a domain knowledge pre-trained language model pre-trained by the knowledge base, and obtaining a second hidden layer feature of the target item knowledge base entry in a hidden layer of the domain knowledge pre-trained language model; Obtaining a target hidden layer feature by fusing the first hidden layer feature and the second hidden layer feature; The target audit scenario category to which the designated work order belongs is determined through the target hidden layer features.

2. The method according to claim 1, characterized in that: Before obtaining the knowledge base and designated work order in the communication field, it also includes: Get some original work orders in the field of communications; For any of the original work orders, entity words, entity descriptions of the entity words, and entity types are extracted from the original work order through a preset large language model and a preset prompt engineering framework; The entity word, the entity description about the entity word and the entity type are stored in a preset format to obtain a knowledge base entry about the entity word.

3. The method according to claim 2, characterized in that: After obtaining the knowledge base and designated work order in the communication field, the method further includes: For any knowledge base entry, obtaining a target entity word from the knowledge base entry; Replacing pre-selected tokens in the target entity word with mask tokens; Inputting the knowledge base entry after the replacement operation into a domain knowledge pre-trained language model to obtain a first sentence-level vector; The first sentence-level vector is projected into a preset vocabulary space to obtain a prediction result about the pre-selected word.

4. The method according to claim 3, characterized in that: After projecting the first sentence-level vector to obtain a prediction result about the pre-selected word, the method further includes: Determining a prediction result of the pre-selected word-unit and a first loss value of the pre-selected word-unit by using a preset loss function; If the first loss value is greater than a preset value, updating the model parameters in the domain knowledge pre-trained language model; Obtaining a prediction result of the pre-selected word unit and a second loss value of the pre-selected word unit by using the domain knowledge pre-trained language model after model parameter update; If the second loss value is less than a preset value, the updated model parameters are determined as target model parameters in the domain knowledge pre-trained language model.

5. The method according to claim 1, characterized in that: The step of fusing the first hidden layer feature and the second hidden layer feature to obtain a target hidden layer feature includes: Using the first hidden layer feature as a query condition in an attention mechanism, and using the second hidden layer feature as a key and a value in the attention mechanism; For any hidden layer, the query condition and the key are compared to obtain a comparison result; Determining the attention weight of the hidden layer according to the comparison result, wherein different comparison results correspond to different attention weights; extracting a target feature from the value by using the attention weight; The target feature is fused with the first hidden layer feature to obtain a target hidden layer feature.

6. The method according to claim 1, characterized in that: Before determining the target audit scenario category to which the designated work order belongs through the target hidden layer features, the method further includes: Get several different pending review scenario categories; Obtain the number and feature vector of the pending review scenario categories; A multidimensional mapping space is generated through the number of suspended audit scenario categories and feature vectors, wherein the dimension of the multidimensional mapping space is determined by the number of suspended audit scenario categories.

7. The method according to claim 6, characterized in that: Determining the target audit scenario category to which the designated work order belongs through the target hidden layer features includes: Obtaining a second sentence-level vector of the target hidden layer feature through a preset operator; Projecting the second sentence-level vector to the multidimensional mapping space, obtaining a probability value that the specified work order belongs to the pending review scenario category; If any of the probability values ​​is greater than a preset threshold, the pending review scenario category corresponding to the probability value is determined as the target review scenario category to which the designated work order belongs; If the probability values ​​are all smaller than the preset thresholds, it is determined that the designated work order belongs to the non-pending review scenario category.

8. A work order review device, characterized in that: The device comprises: A first acquisition module is used to acquire a knowledge base and a designated work order in the field of communications, wherein the knowledge base includes a plurality of knowledge base entries, and the designated work order includes a plurality of key entity words; A second acquisition module is used to input the specified work order into a preset universal pre-trained language model to obtain a first hidden layer feature of the specified work order in a hidden layer of the pre-trained language model; A third acquisition module is used to search the key entity word in the knowledge base to obtain a target item knowledge base entry related to the key entity word; A fourth acquisition module is used to input the target item knowledge base entry into a domain knowledge pre-trained language model pre-trained by the knowledge base, and obtain a second hidden layer feature of the target item knowledge base entry in a hidden layer of the domain knowledge pre-trained language model; A fifth acquisition module, configured to acquire a target hidden layer feature by fusing the first hidden layer feature and the second hidden layer feature; The first determination module is used to determine the target review scenario category to which the specified work order belongs through the target hidden layer features.

9. An electronic device, characterized in that: include: A transceiver, a memory, a processor, and a program stored on the memory and executable on the processor; The processor is used to read the program in the memory to implement the steps in the work order review method as described in any one of claims 1-7.

10. A readable storage medium for storing a program, characterized in that: When the stored program is executed by the processor, the steps in the work order review method as described in any one of claims 1-7 are implemented.

Citation Information

Cited By

  • Vertical field-oriented model training method, electronic equipment, storage medium and program product

    CN121094031A

  • Work order auditing method and device, electronic equipment and computer readable storage medium

    CN121504396A