A framework for extracting configuration command knowledge for operation and maintenance
Through the configuration command knowledge extraction framework for the operation and maintenance field, deep learning and pre-trained language models are used, combined with string editing distance and syntactic attention mechanism, the problem of configuration command knowledge extraction in the operation and maintenance field is solved, and efficient extraction of configuration command entities and relationships is achieved.
Patent Information
- Application Number
- CN202210181458.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-02-25
AI Technical Summary
It is difficult to extract configuration command knowledge in the field of operation and maintenance, and it is difficult for the existing technology to effectively solve the extraction of relationships between configuration commands, especially when faced with long texts, uncommon words, many similar entities and implicit expressions of relationships.
A framework for the extraction of configuration command knowledge for operation and maintenance is proposed, including knowledge template construction module, entity extraction module, relationship classification module and data enhancement module based on bootstrap. The framework uses deep learning models and pre-trained language models, combining string editing distance and syntactic attention mechanisms to extract and classify configuration command entities and relationships.
It effectively solves the cold start problem and insufficient data problem of configuration command knowledge extraction in the operation and maintenance field, improves the accuracy of configuration command entity recognition and relationship extraction, and meets business usage requirements.
Smart Images

Figure CN114547250B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge base construction, and in particular to a framework for extracting configuration command knowledge oriented to the operation and maintenance field. Background Art
[0002] With the rapid development of the Internet, the utilization rate of network equipment has increased, and the frequency of network configuration optimization has increased dramatically. A large amount of digital text information such as manuals, documents, and data has accumulated in the field of operation and maintenance. The massive amount of operation and maintenance data contains rich and important configuration command knowledge, which is an important means for operation and maintenance personnel to configure network equipment, diagnose network faults, and optimize network performance.
[0003] Extracting the relationship between configuration commands in the operation and maintenance field is a key step in building a configuration command knowledge base, which aims to identify the relationship between configuration command entity pairs from the text. However, the operation and maintenance field data uses a unique language form to express rich and complex knowledge. For example, the command entity text in the operation and maintenance data is long and contains many rare and unconventional English words; the text mentions a large number of similar configuration command entities; the relationship is implicitly expressed; and there are many overlapping relationships. Current technology is difficult to solve these pain points.
[0004] In the existing technology, there are two main types of relationship extraction methods: one is the relationship extraction model, and the other is the relationship classification model. The input of the relationship extraction model is only text, and the output is a triple consisting of entities and relationships. According to the different training methods of the model, it can be divided into a joint extraction model and a pipeline model.
[0005] The joint extraction model extracts entities and relations from triples at the same time. This type of method not only requires a large amount of training data, but also requires that the sentences expressing the relations should not be too complex. The pipeline model is divided into multiple models to extract different components from triples. Each model is trained or predicted in a pipeline mode. The model depends on the results of its previous model. This type of method is prone to error accumulation.
[0006] The input of the relation classification model is the text and the given head and tail entity pairs, and the output is the relationship between entities in the text. Depending on the length of the input text and whether the entity information spans sentences, the relation classification model can be divided into sentence-level relation classification model and document-level relation classification model.
[0007] The model design of the sentence-level relationship classification task mainly considers how to obtain the vector representation of the head and tail entities. After obtaining the vector representation of the entity, the final relationship prediction result is output based on the classification module. The document-level relationship classification method is generally a model based on graphs and templates, which uses NLP tools to construct static graph structures or predefine templates. From the model side, the quality of the static graph structure often determines the difficulty of model learning; from the perspective of the graph structure construction method, the quality of the graph structure obtained by NLP tools for long document-level texts is often low. These classification methods often have high requirements for sentence representation and the interference items of the entities cannot be complex.
[0008] On the standard relation extraction dataset NYT, the joint relation extraction model achieved excellent performance, such as the triple extraction F1 of the CasRel model was about 89.6%. Therefore, 200 operation and maintenance domain data were randomly selected and the joint relation extraction model was used for experiments. These methods could not meet the business requirements. For example, the triple extraction F1 of the CasRel model was about 72%. Through case analysis, it is believed that the main reasons for the performance degradation are: the domain data presents the characteristics of low resources and small samples; the configuration command entity text is long and contains many rare and unconventional English words; there are many similar entities in the operation and maintenance data; the relationship between configuration commands is implicitly expressed and there are many overlapping relationships. Summary of the invention
[0009] The present invention is made to solve the above-mentioned problem, and aims to provide a framework for extracting configuration command knowledge oriented to the operation and maintenance field.
[0010] The present invention provides a framework for configuration command knowledge extraction in the field of operation and maintenance, which has the following characteristics: a knowledge template construction module, which defines a configuration command relationship set according to the business requirements of configuration commands in the field of operation and maintenance, and constructs a text containing predefined relationships in a user manual, generalizing it into a knowledge description template; an entity extraction module, which performs fuzzy matching on configuration command entities in combination with the edit distance of a string to extract command entities in the text; a relationship classification module, which uses a deep learning model to model the semantics of the text, and generalizes the rules by learning the configuration command relationships in the text, thereby mining more texts expressing similar relationships; a data enhancement module based on bootstrap, which uses slots to replace the configuration command entities mentioned in the text, regards the generalized text as a high-quality knowledge description template, and adds the high-quality knowledge description template to a template library, and when the newly generated high-quality knowledge description template is less than a threshold, the Bootstrap data expansion and enhancement iterative convergence. Among them, the relationship classification module includes an encoder part based on a pre-trained language model, a syntactic attention part, and a classification part based on PU Learning.
[0011] In the framework of configuration command knowledge extraction for the operation and maintenance field provided by the present invention, there may also be such a feature: wherein the configuration command relationship set at least includes a pre-order relationship and a mutually exclusive relationship. The construction method is to construct by keyword retrieval or sentence matching.
[0012] In the framework of configuration command knowledge extraction for the operation and maintenance field provided by the present invention, there may also be such a feature: wherein the encoder part based on the pre-trained language model is used to combine the text with the configuration command entity to model the semantic features of the text, that is:
[0013] h context =BERT([CLS, c 1 , ..., c n , SEP, cmd 1 , cmd 2 ])
[0014] In the formula, c i (i≤n) represents the i-th token, n represents the length of the text, cmd 1 , cmd 2 They represent two command entities respectively, and CLS and SEP represent BERT special markers respectively.
[0015] In the framework of configuration command knowledge extraction for the operation and maintenance field provided by the present invention, the following features may also be provided: wherein the syntactic attention unit is used to enable the relational classification model to capture the features at the syntactic level, firstly, the knowledge description template is subjected to dependency syntactic analysis to construct a syntactic attention weight matrix, and a syntactic enhanced text vector is obtained by combining the multi-head attention mechanism, specifically including the following steps: Step 1, initializing the attention weight matrix of all zeros, and calculating each token c by using dependency syntactic analysis i The set of all ancestor nodes p i , and update the weight matrix, which is calculated as follows:
[0016]
[0017] Where token c j is token c i Ancestor node; Step 2, use the multi-head attention mechanism to enhance the contextual semantic vector of the encoder:
[0018]
[0019] h sg =W o Concat([A 1 ·V 1 ;...;A i ·V i;...;A H ·V H ])
[0020] Where i (1≤i≤H) represents the i-th head attention module, which is calculated by the multi-head attention matrix, Q, V respectively represent the query Query and key Key, which are used to calculate the attention weights A, A i.j Represents token c i and token c i The original weights between the directly dependent tokens in the syntactic tree are retained, and the semantic representation of the text h c+sg Combining the encoded contextual representation with the syntactically enhanced contextual representation using a weighted fusion strategy:
[0021] h c+sg =α·h context +(1-a)·h sg .
[0022] In the framework of configuration command knowledge extraction for the operation and maintenance field provided by the present invention, there may also be such a feature: wherein the classification part based on PU Learning uses the nnPU loss function in PU learning to optimize the model parameters:
[0023] J mnpu =π·E p(x|y=1) [l(g(x))]+max{0,E p(x) [l(-g(x))]-π·E p(x|y=1) [l(-g(x))]}
[0024]
[0025] In the formula, π is the category prior probability, g is the decision function used to output the probability of the final prediction being a positive sample (y=1), and l is the loss function.
[0026] Functions and Effects of the Invention
[0027] According to the framework of configuration command knowledge extraction for the operation and maintenance field involved in the present invention, it includes: a knowledge template construction module, which defines a configuration command relationship set according to the business needs of configuration commands in the operation and maintenance field, and constructs a text containing predefined relationships in the user manual, generalizing it into a knowledge description template; an entity extraction module, which combines the edit distance of the string to perform fuzzy matching on the configuration command entity to extract the command entity in the text; a relationship classification module, which uses a deep learning model to model the semantics of the text, generalizes the rules by learning the configuration command relationship in the text, and thus mines more texts expressing similar relationships; a data enhancement module based on bootstrap, which uses slots to replace the configuration command entities mentioned in the text, regards the generalized text as a high-quality knowledge description template, and adds the high-quality knowledge description template to the template library. When the newly generated high-quality knowledge description template is less than the threshold, the Bootstrap data expansion and enhancement iterative convergence. Among them, the relationship classification module includes an encoder part based on a pre-trained language model, a syntactic attention part, and a classification part based on PULearning.
[0028] Therefore, the framework knowledge template construction module of the configuration command knowledge extraction for the operation and maintenance field of the present invention mainly solves the data cold start problem in the operation and maintenance field, and the entity extraction module mainly solves the problems of configuration command entities, such as entities that are too long and there are many entities of the same type. The relationship classification module mainly solves the problem of the relationship between configuration commands, and the data enhancement template based on bootstrap mainly solves the problems of insufficient training data and high annotation cost of configuration commands in the operation and maintenance field. This framework solves the cold start problem of configuration command knowledge extraction in the operation and maintenance field, and provides an effective construction method for the construction of a configuration command knowledge base.
[0029] In addition, since the joint relationship model cannot extract configuration command knowledge in the operation and maintenance field, the two-stage configuration command relationship extraction model proposed in the present invention first constructs a configuration command prefix tree and combines the string edit distance to identify configuration command entities from unstructured text, and then uses a deep model to model the semantic relationships expressed in the text for relationship classification, thereby realizing the extraction of configuration command knowledge in the operation and maintenance field.
[0030] In addition, to address the problem of simple samples and single form, the bootstrap-based data enhancement method proposed in the present invention mainly discovers texts that may express knowledge from unlabeled texts by combining a deep relationship classification model with PUlearning, and then expands the knowledge description model to achieve iterative data enhancement.
[0031] In addition, the present invention effectively combines data enhancement and deep classification models to solve the problem of low resources and few samples of configuration command data in the field of operation and maintenance. The data of configuration commands in the field of operation and maintenance present the characteristics of low resources and few samples, which greatly reduces the performance of traditional knowledge extraction models, seriously restricts the extraction ability of the model, and is far from meeting the requirements of business use. The present invention proposes a bootstrap-based data enhancement method and a deep classification model to utilize a small amount of expert knowledge to expand the amount of configuration command data in the field of operation and maintenance and improve the quality of knowledge extraction. This is a first in the knowledge extraction of configuration commands in the field of operation and maintenance.
[0032] In addition, the present invention effectively uses the method of pre-tree and edit distance to solve the problems of long entities and many similar entities. Traditional relationship extraction can only solve entities with short spans in text, and it is difficult to identify similar entities, which is far from meeting the needs of command configuration entity recognition. The present invention proposes to use the configuration command pre-tree, combined with the string edit distance, to not only solve the problems of long entities and many similar entities, but also solves the problem that the deep model cannot link the entity to the command entity node of the configuration command graph object. This is a first in the field of operation and maintenance command configuration entity recognition.
[0033] Finally, the present invention effectively combines expert background knowledge to solve the model cold start problem. The configuration command data in the operation and maintenance field is fragmented, with little knowledge precipitation and reuse, and it is impossible to directly extract configuration command knowledge through deep models. The present invention constructs a knowledge description template based on expert background knowledge, and uses deep models to improve the generalization of the knowledge description template, expands the knowledge description template, and solves the model cold start problem caused by lack of data. This is a first in the construction of operation and maintenance configuration command data. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a framework diagram of configuration command knowledge extraction for the operation and maintenance field in an embodiment of the present invention;
[0035] Figure 2 is a schematic diagram of a BootRE extraction model in an embodiment of the present invention;
[0036] Figure 3 Schematic diagram of the BootRE relational classification model in an embodiment of the present invention. DETAILED DESCRIPTION
[0037] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments and accompanying drawings specifically illustrate the framework of configuration command knowledge extraction for the operation and maintenance field of the present invention.
[0038] This embodiment aims to extract configuration command knowledge from the operation and maintenance manual. Since this knowledge is relatively scattered in the manual and appears in different description forms, it is difficult to obtain through the existing extraction model.
[0039] This embodiment provides a framework for extracting configuration command knowledge in the field of operation and maintenance.
[0040] Figure 1 It is a framework diagram of configuration command knowledge extraction for the operation and maintenance field in this embodiment.
[0041] like Figure 1 As shown, the framework of configuration command knowledge extraction for the operation and maintenance field of this embodiment includes a knowledge template construction module, an entity extraction module, a relationship classification module and a bootstrap-based data enhancement module.
[0042] The knowledge template construction module defines a set of configuration command relationships according to the business requirements of configuration commands in the operation and maintenance field, and constructs text containing predefined relationships in the user manual, generalizing it into a knowledge description template.
[0043] Due to the lack of annotated configuration command knowledge in the operation and maintenance field, the remote supervision method cannot be directly applied. This embodiment defines a configuration command relationship set, such as a pre-order relationship, a mutually exclusive relationship, etc., according to the business needs of configuration commands in the operation and maintenance field. Some texts containing predefined relationships are constructed in the user manual through keyword retrieval or sentence matching, and generalized into a knowledge description template.
[0044] For example, "lacpforce-forward and maxactive-linknumber are mutually exclusive. This command cannot be configured after the upper and lower limit thresholds of the Eth-Trunk interface member links are configured." This can be generalized into a knowledge description template "{cmd_1} and {cmd_2} are mutually exclusive." In this way, different numbers of knowledge description templates are constructed for each predefined relationship to solve the cold start problem and prepare for the subsequent mining of accurate seed samples.
[0045] The entity extraction module combines the edit distance of the string to perform fuzzy matching on the configuration command entities to extract the command entities in the text.
[0046] The text matched by the knowledge description module may express a certain semantic relationship, but lacks entity information. In addition, the same configuration command entity with different parameter values will have different functions. Therefore, in order to facilitate operation and maintenance personnel to accurately locate the regional information of the configuration command entity, this embodiment proposes a prefix tree method based on the configuration command.
[0047] Figure 2 Schematic diagram of the BootRE extraction model in this embodiment.
[0048] like Figure 2As shown in the figure, this method combines the edit distance of the string to perform fuzzy matching on the configuration command entity to extract the command entity in the text. This method uses a rule-based and data mining algorithm to solve the problem of insufficient training data for deep models, and uses the prefix tree constructed by the complete configuration command library to achieve high-quality entity extraction.
[0049] The relationship classification module uses a deep learning model to model the semantics of the text and generalizes the rules by learning the configuration command relationships in the text, thereby mining more texts that express similar relationships.
[0050] The data obtained through the knowledge description template is essentially a strong rule matching method and does not have generalization. To address this problem, this embodiment proposes a deep relationship classification model.
[0051] Figure 3 Schematic diagram of the BootRE relational classification model in this embodiment.
[0052] like Figure 3 As shown in the figure, the model uses a deep learning model to model the semantics of the text and generalizes the rules by learning the configuration command relationship in the text in order to mine more texts expressing similar relationships.
[0053] The relation classification module includes an encoder part based on a pre-trained language model, a syntactic attention part, and a classification part based on PULearning.
[0054] The encoder part based on the pre-trained language model is used to combine the text with the configuration command entity to model the semantic features of the text, namely:
[0055] h context =BERT([CLS, c 1 , ..., c n ,SEP,cmd 1 , cmd 2 ])
[0056] In the formula, c i (i≤n) represents the i-th token, n represents the length of the text, cmd 1 , cmd 2 They represent two command entities respectively, and CLS and SEP represent BERT special markers respectively.
[0057] The syntactic attention part is used to enable the relation classification model to capture syntactic features. First, the knowledge description template is subjected to dependency syntactic analysis to construct a syntactic attention weight matrix, and then the syntactic enhanced text vector is obtained by combining the multi-head attention mechanism. Specifically, the following steps are included:
[0058] Step 1: Initialize the attention weight matrix to zero and use dependency parsing to calculate each token c i The set of all ancestor nodes p i , and update the weight matrix, which is calculated as follows:
[0059]
[0060] Where token c j is token c i ancestor node.
[0061] Step 2: Use the multi-head attention mechanism to enhance the encoder’s contextual semantic vector:
[0062]
[0063] h sg =W o Concat([A 1 ·V 1 ;...;A i ·V i ;...;A H ·V H ])
[0064] Where i (1≤i≤H) represents the i-th head attention module, which is calculated by the multi-head attention matrix, Q, V respectively represent the query Query and key Key, which are used to calculate the attention weights A, A i.j Represents token c i and token c i The attention score between them guides the model to focus on the information between tokens with high scores.
[0065] In order to enable the model to better enhance the modeling ability of text with the help of external syntactic information, this embodiment retains the original weights between the directly dependent associated tokens in the syntactic tree, and the semantic representation of the text h c+sg Combining the encoded contextual representation with the syntactically enhanced contextual representation using a weighted fusion strategy:
[0066] h c+sg =α·h context +(1-a)·h sg .
[0067] The classification part based on PU Learning is mainly to solve the problem of false negative data in negative samples and enhance the generalization of classification boundaries. Therefore, the nnPU loss function in PU learning is used to optimize the model parameters:
[0068] Jmnpu =π·E p(x|y=1) [l(g(x))]+max{0,E p(x) [l(-g(x))]-π·E p(x|y=1) [l(-g(x))]}
[0069]
[0070] In the formula, π is the category prior probability, g is the decision function used to output the probability of the final prediction being a positive sample (y=1), and l is the loss function.
[0071] In view of the lack of sample diversity in the seed training dataset, the bootstrap-based data enhancement module uses the generalization ability of the deep model to perform data enhancement. The specific process is as follows:
[0072] The configuration command entities mentioned in the text are replaced by slots, and the generalized text is regarded as a high-quality knowledge description template. Since the number of new knowledge templates generated in each Bootstrap iteration is small, manual work is introduced to control the quality of the relationship classification generalization template.
[0073] Add high-quality knowledge description templates to the template library to build larger-scale data, improve the generalization performance of the relational classification model, and discover more diverse texts containing knowledge. When the number of newly generated high-quality knowledge description templates is less than the threshold, Bootstrap data augmentation and enhancement iterative convergence.
[0074] Functions and Effects of the Embodiments
[0075] According to the framework of configuration command knowledge extraction for the operation and maintenance field involved in this embodiment, it includes: a knowledge template construction module, which defines a configuration command relationship set according to the business needs of configuration commands in the operation and maintenance field, and constructs a text containing predefined relationships in the user manual, generalizing it into a knowledge description template; an entity extraction module, which combines the edit distance of the string to perform fuzzy matching on the configuration command entity to extract the command entity in the text; a relationship classification module, which uses a deep learning model to model the semantics of the text, and generalizes the rules by learning the configuration command relationship in the text, so as to mine more texts expressing similar relationships; a data enhancement module based on bootstrap, which uses slots to replace the configuration command entities mentioned in the text, regards the generalized text as a high-quality knowledge description template, and adds the high-quality knowledge description template to the template library. When the newly generated high-quality knowledge description template is less than the threshold, the Bootstrap data expansion and enhancement iterative convergence. Among them, the relationship classification module includes an encoder part based on a pre-trained language model, a syntactic attention part, and a classification part based on PULearning.
[0076] Therefore, the framework knowledge template construction module for configuration command knowledge extraction in the operation and maintenance field of this embodiment mainly solves the problem of cold start of data in the operation and maintenance field, and the entity extraction module mainly solves problems such as configuration command entities, such as entities that are too long and there are many entities of the same type. The relationship classification module mainly solves the problem of the relationship between configuration commands, and the data enhancement template based on bootstrap mainly solves the problem of insufficient training data and high annotation cost of configuration commands in the operation and maintenance field. This framework solves the cold start problem of configuration command knowledge extraction in the operation and maintenance field, and provides an effective construction method for the construction of a configuration command knowledge base.
[0077] In addition, since the joint relationship model cannot extract configuration command knowledge in the operation and maintenance field, the two-stage configuration command relationship extraction model proposed in this embodiment first constructs a configuration command prefix tree and combines the string edit distance to identify configuration command entities from unstructured text, and then uses a deep model to model the semantic relationships expressed in the text for relationship classification, thereby realizing the extraction of configuration command knowledge in the operation and maintenance field.
[0078] In addition, in order to address the problem of simple samples and single form, the bootstrap-based data enhancement method proposed in this embodiment mainly discovers texts that may express knowledge from unlabeled texts by combining a deep relationship classification model with PUlearning, and then expands the knowledge description model to achieve iterative data enhancement.
[0079] In addition, this embodiment effectively combines data enhancement and deep classification models to solve the problem of low resources and few samples of configuration command data in the operation and maintenance field. The data of configuration commands in the operation and maintenance field presents the characteristics of low resources and few samples, which greatly reduces the performance of traditional knowledge extraction models, seriously restricts the extraction ability of the model, and is far from meeting the business use requirements. This embodiment proposes a bootstrap-based data enhancement method and a deep classification model to use a small amount of expert knowledge to expand the amount of configuration command data in the operation and maintenance field and improve the quality of knowledge extraction. This is a first in the knowledge extraction of configuration commands in the operation and maintenance field.
[0080] In addition, this embodiment effectively uses the method of pre-tree and edit distance to solve the problems of long entities and many similar entities. Traditional relationship extraction can only solve entities with short spans in text, and it is difficult to identify similar entities, which is far from meeting the needs of command configuration entity recognition. This embodiment proposes to use the configuration command pre-tree, combined with the string edit distance, not only to solve the problems of long entities and many similar entities, but also to solve the problem that the deep model cannot link the entity to the command entity node of the configuration command graph object. This is a first in the field of operation and maintenance command configuration entity recognition.
[0081] Finally, this embodiment effectively combines expert background knowledge to solve the model cold start problem. The configuration command data in the operation and maintenance field is fragmented, and there is little knowledge precipitation and reuse, so it is impossible to directly extract configuration command knowledge through the deep model. This embodiment constructs a knowledge description template based on expert background knowledge, and uses a deep model to improve the generalization of the knowledge description template, expand the knowledge description template, and solve the model cold start problem caused by lack of data. This is a first in the construction of operation and maintenance configuration command data.
[0082] The above-mentioned embodiments are preferred examples of the present invention and are not intended to limit the protection scope of the present invention.
Claims
1. A framework for extracting configuration command knowledge in the field of operation and maintenance, characterized in that: include: The knowledge template construction module defines the configuration command relationship set according to the business requirements of configuration commands in the operation and maintenance field, and constructs text containing predefined relationships in the user manual to generalize it into a knowledge description template; An entity extraction module performs fuzzy matching on configuration command entities in combination with the edit distance of the character string to extract command entities in the text; A relationship classification module uses a deep learning model to model the semantics of the text, generalizes rules by learning the configuration command relationship in the text, and thus mines more texts expressing similar relationships; Based on the bootstrap data enhancement module, the configuration command entity mentioned in the text is replaced by a slot, the generalized text is regarded as a high-quality knowledge description template, and the high-quality knowledge description template is added to the template library. When the newly generated high-quality knowledge description template is less than a threshold, the Bootstrap data expansion and enhancement iterative convergence is completed. The relation classification module includes an encoder part based on a pre-trained language model, a syntactic attention part, and a classification part based on PU Learning. The syntactic attention unit is used to enable the relational classification model to capture syntactic features. First, the knowledge description template is subjected to dependency syntactic analysis to construct a syntactic attention weight matrix, and a syntactic enhanced text vector is obtained by combining a multi-head attention mechanism. Specifically, the following steps are included: Step 1: Initialize the attention weight matrix to zero and use dependency parsing to calculate each token c i The set of all ancestor nodes p i , and update the weight matrix, which is calculated as follows: Where token c j is token c i Ancestor node of Step 2: Use the multi-head attention mechanism to enhance the encoder’s contextual semantic vector: h sg =W o ·Concat([A1·V1;…;A i ·V i ;…;A H ·V H ]) In the formula, i represents the i-th head attention module, which is calculated by the multi-head attention matrix, Q, V represents the query Query and key Key, respectively, which are used to calculate the attention weights A, A i.j Represents token c i and token c i The attention score between them guides the model to focus on the information between tokens with high scores. The original weights between the directly dependent tokens in the syntactic tree are retained, and the semantic representation of the text h c+sg Combining the encoded contextual representation with the syntactically enhanced contextual representation using a weighted fusion strategy: h c+sg =α·h context +(1-a)·h sg The PU Learning-based classification unit uses the nnPU loss function in PU learning to optimize model parameters: J mnpu =π·E p(x|y=1) [l(g(x))]+max{0,E p(x) [l(-g(x))]-π·E p(x|y=1) [l(-g(x))]} In the formula, π is the category prior probability, g is the decision function used to output the probability that the final prediction is a positive sample y=1, and l is the loss function.
2. The configuration command knowledge extraction framework for the operation and maintenance field according to claim 1 is characterized by: in, The configuration command relationship set at least includes a pre-order relationship and a mutually exclusive relationship. The construction method is to construct through keyword retrieval or sentence matching.
3. The configuration command knowledge extraction framework for the operation and maintenance field according to claim 1 is characterized in that: in, The encoder part based on the pre-trained language model is used to combine the text with the configuration command entity to model the semantic features of the text, namely: h context =BERT([CLS,c1,…,c n ,SEP,cmd1,cmd2]) In the formula, c i , i≤n, represents the i-th token, n represents the length of the text, cmd1 and cmd2 represent two command entities respectively, and CLS and SEP represent BERT special markers respectively.
Citation Information
Patent Citations
Knowledge graph relational data extraction method based on semantic syntax interaction network
CN111241295A
Information system-oriented knowledge graph construction method, device and electronic equipment
CN111782817A