An information extraction method and device for unstructured detection text
By combining BERT, SVM, and LSTM-CRF models with the PageRank algorithm, the problems of low efficiency and poor accuracy in unstructured text processing are solved, achieving efficient and accurate information extraction and event chain construction, supporting intelligent decision-making in power material management.
Patent Information
- Application Number
- CN202511167960.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-20
Smart Images

Figure CN120670589B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of language processing, and in particular to an information extraction method and device for unstructured detection text. BACKGROUND
[0002] In the field of power material management, the processing of unstructured detection reports has always been a key and complex link. With the rapid development of smart grids and the increasing diversification of power equipment, the detection of power materials has generated a large amount of unstructured detection text data. These data cover various forms such as equipment detection reports, fault records, and operation logs. These detection texts contain rich key information such as detection items, detection values, units, and detection results, which are crucial for the quality evaluation, fault prediction, and operation and maintenance decision of power materials.
[0003] Patent CN119645826A proposes an AI-based unstructured test case conversion method and device. The method includes preprocessing the test case to be converted and generating a text vector based on the preprocessing result. The target category is determined based on the preset classification template and the text vector. The preset classification template includes at least two target categories. The test case to be converted is converted based on the target category and the preset output document to obtain a structured test case. In this method, the classification template is used for target classification, which has certain limitations. First, it relies on manually developed rules or templates. Although this method is simple and direct, it lacks flexibility and adaptability. Second, this method has obvious shortcomings in processing text semantic understanding, and cannot accurately capture deep semantic information and context relationships in the text, which limits the accuracy and completeness of information extraction. In addition, this method cannot construct information association relationships, lacks systematicness and logicality, and is difficult to form a complete and coherent event information chain, thereby affecting the value and efficiency of information utilization. In summary, the existing technology has the disadvantages of low efficiency, poor accuracy, insufficient flexibility, and weak information association when processing unstructured text detection files, which cannot meet the needs of modern power material management for efficient, accurate, and intelligent information processing. SUMMARY
[0004] The present application provides an information extraction method and device for unstructured detection text, which can improve the accuracy and flexibility of unstructured detection text information extraction.
[0005] An information extraction method for unstructured detection text, comprising:
[0006] Obtain the unstructured detection text to be extracted and preprocess it to obtain a preprocessed text. Encode the preprocessed text to obtain a detection text sequence.
[0007] inputting the detection text sequence into a BERT model for deep semantic coding, and outputting a semantic feature vector;
[0008] inputting the semantic feature vector into a pre-trained SVM model for classification and recognition, extracting a key detection text segment from the unstructured detection text according to a classification and recognition result;
[0009] inputting the key detection text segment into a pre-trained LSTM-CRF model for attribute recognition, and outputting detection text attribute information;
[0010] extracting relevant standard detection information in an industry standard, establishing a logical mapping relationship rule set according to the standard detection information, determining an association relationship between each information element in the detection text attribute information according to the logical mapping relationship rule set, and forming an information association rule set;
[0011] constructing a detection event information chain according to the information association rule set.
[0012] Further, the unstructured detection text to be extracted is preprocessed, including:
[0013] performing character recognition on the unstructured detection text, removing redundant words, and obtaining the preprocessed text;
[0014] encoding the preprocessed text to obtain a detection text sequence, including:
[0015] encoding the words in the preprocessed text to obtain the detection text sequence.
[0016] Further, the BERT model includes a multi-layer Transformer encoder, which is used for word embedding processing and coding of each word in the detection text sequence to generate the semantic feature vector.
[0017] Further, the SVM model is used to determine whether a corresponding word is a key word according to the semantic feature vector; if the corresponding word is a key word, a corresponding text segment is extracted according to the position of the key word in the preprocessed text as the key detection text segment.
[0018] Further, according to the position of the key word in the preprocessed text, a corresponding text segment is extracted as the key detection text segment, including:
[0019] positioning the position of the key word in the preprocessed text, searching from the position to both ends, stopping the search when a terminal symbol is searched, and extracting the text between the two terminal symbols as an initial key detection text segment;
[0020] The obtained initial key detection text segment is de-duplicated to obtain the key detection text segment.
[0021] Further, the LSTM-CRF model comprises an input layer, an embedding layer, a forward LSTM, a backward LSTM and a CRF layer.
[0022] The input layer is configured to receive the key detection text segment and generate a word sequence according to the key detection text segment.
[0023] The embedding layer is configured to map each word in the word sequence to a word vector to obtain a word vector sequence.
[0024] The forward LSTM is configured to process the word vector sequence from left to right according to position to obtain a forward hidden state, and the backward LSTM is configured to process the word vector sequence from right to left according to position to obtain a backward hidden state.
[0025] The forward hidden state and the backward hidden state are spliced to obtain a hidden state vector.
[0026] The CRF layer is configured to perform attribute prediction on the hidden state vector to obtain a label sequence about the attribute prediction result, match the label sequence with the words in the key detection text segment, and generate detection text attribute information.
[0027] Further, a logical mapping relationship rule set is established according to the standard detection information, comprising:
[0028] Standard detection content is extracted from the standard detection information, and the attributes of each standard detection content are determined, each standard detection content having an associated relationship is associated mapped according to the attributes of the standard detection content, and the logical mapping relationship rule set is formed.
[0029] Further, the information elements in the detection text attribute information are detection content words, and each detection content word is provided with a corresponding attribute label.
[0030] According to the logical mapping relationship rule set, the associated relationship between each information element in the detection text attribute information is determined to form an information association rule set, comprising:
[0031] Select two detection content words with different attributes to form a to-be-confirmed word group, compare the attribute of the detection content word in the to-be-confirmed word group with the attribute of the standard detection content having the association relationship in the logical mapping relationship rule set, if the attribute of the detection content word is same as the attribute of the standard detection content having the association relationship in the logical mapping relationship rule set, confirm that the detection content word in the to-be-confirmed word group has the association relationship;
[0032] Obtain all detection content words having the association relationship to form the information association rule set.
[0033] Further, the association relationship includes a subordinate relationship and a cause-effect relationship;
[0034] According to the information association rule set, a detection event information chain is constructed, including:
[0035] The detection content word in the information association rule set is taken as a node, the association relationship is taken as an edge, the direction of the edge is determined according to the subordinate relationship and the cause-effect relationship between the nodes, and an association relationship directed graph is obtained;
[0036] Based on the association relationship directed graph, a PageRank algorithm is used to calculate the weight of the node, and the weight of each node is obtained;
[0037] Each node is traversed, for each current node, the node with the highest weight value in the nodes having the association relationship with the current node is determined as a final association node, and the current node and the final association node form an initial information chain;
[0038] According to the direction between the nodes in the initial information chain and the front and back position relationship of the nodes in the preprocessed text, the nodes are integrated and sorted, and the detection event information chain is obtained.
[0039] Further, the weight of the node is calculated by the following formula:
[0040] ;
[0041] Wherein, PR(u i ) represents the weight of the node u i , d represents a damping factor, IN(u i ) represents a set of all nodes pointing to the node u i , u j is a node in the set IN(u i ), PR(u j ) represents the weight of the node u j , represents the number of nodes pointed to by the node u j , represents the number of nodes pointed to by the node u jPointing to node u i The weighting coefficients.
[0042] An information extraction device for unstructured detection text, comprising:
[0043] The preprocessing module is used to acquire the unstructured detection text to be extracted and preprocess it to obtain preprocessed text, and encode the preprocessed text to obtain a detection text sequence.
[0044] The encoding module is used to input the detected text sequence into the BERT model for deep semantic encoding and output semantic feature vectors;
[0045] The classification module is used to input the semantic feature vector into a pre-trained SVM model for classification and recognition, and extract key detection text fragments from the unstructured detection text based on the classification and recognition results.
[0046] The detection module is used to input the key detection text fragments into a pre-trained LSTM-CRF model for attribute recognition and output the detection text attribute information.
[0047] The association module is used to extract relevant standard testing information from industry standards, establish a logical mapping relationship rule set based on the standard testing information, determine the association relationship between each information element in the test text attribute information based on the logical mapping relationship rule set, and form an information association rule set.
[0048] The construction module is used to construct a detection event information chain based on the information association rule set.
[0049] Furthermore, the preprocessing module preprocesses the extracted unstructured detection text, including:
[0050] The unstructured detection text is subjected to text recognition to remove redundant words, thereby obtaining the preprocessed text;
[0051] The preprocessing module encodes the preprocessed text to obtain a detection text sequence, including:
[0052] The words in the preprocessed text are encoded to obtain the detection text sequence.
[0053] Furthermore, the BERT model includes a multi-layer Transformer encoder, which is used to perform word embedding processing and encoding on each word in the detected text sequence to generate the semantic feature vector.
[0054] Furthermore, the SVM model is used to determine whether the corresponding word is a keyword based on the semantic feature vector; if the corresponding word is a keyword, the corresponding text segment is extracted as the key detection text segment based on the position of the keyword in the preprocessed text.
[0055] Based on the location of the keyword in the preprocessed text, the corresponding text fragment is extracted as the key detection text fragment, including:
[0056] Locate the position of the keyword in the preprocessed text, search from the position to both ends, stop the search when the end symbol is found, and extract the text between the two end symbols as the initial key detection text segment;
[0057] The obtained initial key detection text fragments are deduplicated to obtain the key detection text fragments.
[0058] Furthermore, the LSTM-CRF model includes an input layer, an embedding layer, a forward LSTM, a backward LSTM, and a CRF layer;
[0059] The input layer is used to receive the key detection text fragments and generate a word sequence based on the key detection text fragments;
[0060] The embedding layer is used to map each word in the word sequence to a word vector to obtain a word vector sequence;
[0061] The forward LSTM is used to process the word vector sequence from left to right according to its position to obtain the forward hidden state, and the backward LSTM is used to process the word vector sequence from right to left according to its position to obtain the backward hidden state.
[0062] The forward hidden state and the backward hidden state are concatenated to obtain the hidden state vector;
[0063] The CRF layer is used to predict the attributes of the hidden state vector, obtain a label sequence of the attribute prediction results, and match the label sequence with words in the key detection text segment to generate detection text attribute information.
[0064] Furthermore, the association module establishes a set of logical mapping relationship rules based on the standard detection information, including:
[0065] The standard testing content is extracted from the standard testing information, and the attributes of each standard testing content are determined. Based on the attributes of the standard testing content, the related standard testing content is mapped to form the logical mapping relationship rule set.
[0066] Furthermore, the information elements in the detected text attribute information are detected content words, and each detected content word is set with a corresponding attribute label;
[0067] The association module determines the association relationships between various information elements in the detected text attribute information based on the logical mapping relationship rule set, forming an information association rule set, including:
[0068] Select detection content words with different attributes to form a word group to be confirmed. Compare the attributes of the detection content words in the word group to be confirmed with the attributes of the standard detection content that has a relationship in the logical mapping relationship rule set. If the attributes of the detection content words are the same as the attributes of the standard detection content that has a relationship in the logical mapping relationship rule set, then it is confirmed that the detection content words in the word group to be confirmed have a relationship.
[0069] The information association rule set is formed by acquiring all related words in the detected content.
[0070] Furthermore, the aforementioned relationships include subordinate relationships and causal relationships;
[0071] The graph construction module constructs a detection event association graph based on the information association rule set, including:
[0072] Using the detected content words in the information association rule set as nodes and the association relationships as edges, the direction of the edges is determined based on the subordinate and causal relationships between nodes, thus obtaining a directed graph of association relationships.
[0073] Based on the directed graph of the aforementioned relationships, the PageRank algorithm is used to calculate the node weights to obtain the weight of each node;
[0074] Traverse each node. For each current node, determine the node with the highest weight value among the nodes that are related to the current node as the final related node, and form an initial information chain between the current node and the final related node.
[0075] The nodes are integrated and sorted according to the pointers between nodes in the initial information chain and the sequential positional relationship of the nodes in the preprocessed text to obtain the detection event information chain.
[0076] Furthermore, the weight of the node is calculated using the following formula:
[0077] ;
[0078] Among them, PR(u i ) represents node u i The weights, d represents the damping factor, IN(u i ) represents all pointers to node ui The set of nodes, u j For set IN(u i A node in ) PR(u j ) represents node u j The weight, Represents node u j The number of nodes it points to. Represents node u j Pointing to node u i The weighting coefficients.
[0079] The present invention provides a method and apparatus for information extraction from unstructured detection text, which has at least the following beneficial effects:
[0080] (1) By combining the BERT model, SVM model and LSTM-CRF model, the key detection text fragments in unstructured detection text were extracted. The BERT model can deeply understand the deep semantics of the text and generate semantic feature vectors. The SVM model locates keywords based on these feature vectors. The LSTM-CRF model further extracts attribute information from the key detection text fragments, such as the name of the detection project, the value, the unit and the detection result. This not only improves the speed of information processing, but also reduces errors caused by human factors and improves the accuracy of information extraction, providing reliable data support for power material management.
[0081] (2) By defining a logical mapping rule set and constructing an information association rule set, the PageRank algorithm is used to calculate the node weights and capture the complex relationships between each information element. This forms a detection event information chain that includes key information from the start of detection to the result determination, enabling power material management personnel to deeply understand the detection events and providing information basis for material quality assessment, fault prediction and power grid safe operation.
[0082] (3) The final obtained detection event information chain only focuses on key detection events in unstructured detection text, which effectively saves the time of power material management personnel and improves production efficiency. Attached Figure Description
[0083] Figure 1 This is a flowchart of one embodiment of the information extraction method for unstructured detection text provided by the present invention.
[0084] Figure 2 This is a flowchart of one embodiment of the information association rule set in the information extraction method for unstructured detection text provided by the present invention.
[0085] Figure 3This is a flowchart of one embodiment of the information extraction method for unstructured detection text provided by the present invention for constructing a detection event association graph.
[0086] Figure 4 This is a flowchart of one embodiment of the information extraction device for unstructured text detection provided by the present invention. Detailed Implementation
[0087] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0088] refer to Figure 1 In some embodiments, a method for extracting information from unstructured detection text is provided, including:
[0089] S1. Obtain the unstructured detection text to be extracted and preprocess it to obtain preprocessed text. Encode the preprocessed text to obtain the detection text sequence.
[0090] S2. Input the detected text sequence into the BERT model for deep semantic encoding and output semantic feature vectors;
[0091] S3. Input the semantic feature vector into the pre-trained SVM model for classification and recognition, and extract key detection text fragments from the unstructured detection text based on the classification and recognition results;
[0092] S4. Input the key detected text fragments into the pre-trained LSTM-CRF model for attribute recognition and output the detected text attribute information.
[0093] S5. Extract relevant standard testing information from industry standards, establish a logical mapping relationship rule set based on the standard testing information, determine the association relationship between each information element in the test text attribute information based on the logical mapping relationship rule set, and form an information association rule set.
[0094] S6. Construct a detection event information chain based on the information association rule set.
[0095] Specifically, in step S1, the unstructured detection text to be extracted is preprocessed, including:
[0096] The unstructured detection text is subjected to text recognition to remove redundant words, thereby obtaining the preprocessed text;
[0097] The preprocessed text is encoded to obtain a detection text sequence, including:
[0098] The words in the preprocessed text are encoded to obtain the detection text sequence.
[0099] Specifically, in step S1, there may be words such as interjections and exclamations that are irrelevant to the detected content. Therefore, during the preprocessing process, redundant words such as these are identified to obtain the preprocessed text.
[0100] The words in the preprocessed text are encoded to obtain the detection text sequence.
[0101] The detected text sequence T can be represented as T = [t1, t2, ..., tt] n ], where t i This represents the i-th word.
[0102] Further, in step S2, the detected text sequence T is input into the BERT model for deep semantic encoding. The BERT model includes a multi-layer Transformer encoder, which is used to perform word embedding processing and encoding on each word in the detected text sequence to generate the semantic feature vector.
[0103] Specifically, the detected text sequence T is input into the BERT model and is first converted into a word embedding matrix:
[0104] ;
[0105] Where E(T) represents the word embedding matrix, e(t) i ) represents the embedding vector of the i-th word.
[0106] The word embedding matrix is bidirectionally context-encoded through a layer-by-layer Transformer encoder, and the final output semantic feature vector is shown below:
[0107] ;
[0108] Where H represents the output semantic feature vector matrix, H=[h1, h2, ..., h n ], hi represents the i-th semantic feature vector, L represents the number of Transformer encoders, Transformer L This represents the L-th layer Transformer encoder.
[0109] Further, in step S3, the semantic feature vector is input into the pre-trained SVM model for classification and recognition to determine whether the corresponding word is a keyword.
[0110] The output function of the SVM model is shown below:
[0111] ;
[0112] Where f(hi) represents the output of the SVM model. When the output is +1, it indicates that the corresponding word is a keyword; when the output is -1, it indicates that it is a non-keyword. A single semantic feature vector is extracted from the semantic feature matrix H output by the BERT model. For symbolic functions, Let α be the semantic feature vector corresponding to the training sample. j For Lagrange multipliers, y j For the labels of the training samples, K(h) i h j ) is the kernel function, b is the bias term, and m is the number of training samples.
[0113] If the corresponding word is a keyword, then the corresponding text segment is extracted as the key detection text segment based on the position of the keyword in the preprocessed text, specifically including:
[0114] Locate the position of the keyword in the preprocessed text, search from the position to both ends, stop the search when the end symbol is found, and extract the text between the two end symbols as the initial key detection text segment;
[0115] The obtained initial key detection text fragments are deduplicated to obtain the key detection text fragments.
[0116] Among them, the terminator can be a punctuation mark in the text that indicates the end of a sentence, such as a period or semicolon.
[0117] Further, in step S4, the key detection text fragments are input into the LSTM-CRF model, which includes an input layer, an embedding layer, a forward LSTM, a backward LSTM, and a CRF layer;
[0118] The input layer is used to receive the key detection text fragments and generate a word sequence based on the key detection text fragments;
[0119] The embedding layer is used to map each word in the word sequence to a word vector to obtain a word vector sequence;
[0120] The forward LSTM is used to process the word vector sequence from left to right according to its position to obtain the forward hidden state, and the backward LSTM is used to process the word vector sequence from right to left according to its position to obtain the backward hidden state.
[0121] The forward hidden state and the backward hidden state are concatenated to obtain the hidden state vector;
[0122] The CRF layer is used to predict the attributes of the hidden state vector, obtain a label sequence of the attribute prediction results, and match the label sequence with words in the key detection text segment to generate detection text attribute information.
[0123] Specifically, the input layer converts the received key detection text fragments into a word sequence W, represented as: W=[w1, w2, ... w2]. n ];w i Let be the i-th word in the word sequence W.
[0124] The embedding layer will embed each word w in the word sequence W. i Mapped to 300-dimensional word vector p i The word vector sequence P is obtained by the following formula:
[0125] p i =Embedding(w i );
[0126] Here, Embedding is a word embedding function.
[0127] Furthermore, the forward LSTM is used to process the word vector sequence from left to right according to its position, capturing the context from left to right and obtaining the forward hidden state. The specific calculation formula is as follows:
[0128] ;
[0129] in, The computation function for the forward LSTM. Let be the word vector at time t. For forward LSTM in The forward hidden state at any given moment Let t be the forward hidden state.
[0130] Furthermore, the backward LSTM is used to process the word vector sequence from right to left according to its position, capturing the right-to-left context and obtaining the backward hidden state. The specific calculation formula is as follows:
[0131] ;
[0132] in, The computation function for the backward LSTM. For backward LSTM in The hidden state at all times Let t be the backward hidden state.
[0133] The forward hidden state and the backward hidden state are concatenated, and the calculation formula is as follows: ,in, This represents the vector concatenation operation, s t The dimension is 256, and it captures text context features, outputting the hidden state vector s at each position. t .
[0134] The CRF layer performs attribute prediction on the hidden state vector to obtain a label sequence based on the attribute prediction results. This label sequence is then matched with words in the key detection text fragment to generate detection text attribute information. Attributes may include the detection item name, numerical value, unit, detection result, etc.
[0135] Further, in step S5, a set of logical mapping relationship rules is established based on the standard detection information, including:
[0136] The standard testing content is extracted from the standard testing information, and the attributes of each standard testing content are determined. Based on the attributes of the standard testing content, the related standard testing content is mapped to form the logical mapping relationship rule set.
[0137] Specifically, standard testing content can be extracted from the standard testing information through keyword recognition, and the attributes of each standard testing content can be determined, including the testing item name, value, unit, judgment result, judgment rule, etc. Based on the attributes of the standard testing content, related standard testing content is mapped to form the logical mapping relationship rule set.
[0138] Furthermore, the information elements in the detection text attribute information output by the CRF layer are detection content words, and each detection content word has a corresponding attribute label. For example, the attribute label for "insulation resistance" is "detection item name", the attribute label for "1000" is "value", the attribute label for "Ω" is "unit", and the attribute label for "qualified" is "detection result".
[0139] refer to Figure 2 In step S5, the association relationships between various information elements in the detected text attribute information are determined according to the logical mapping relationship rule set, forming an information association rule set, including:
[0140] S51. Select detection content words with different attributes to form a group of words to be confirmed. Compare the attributes of the detection content words in the group of words to be confirmed with the attributes of the standard detection content that has a relationship in the logical mapping relationship rule set. If the attributes of the detection content words are the same as the attributes of the standard detection content that has a relationship in the logical mapping relationship rule set, then confirm that the detection content words in the group of words to be confirmed have a relationship.
[0141] S52. Obtain all related words in the detected content to form the information association rule set.
[0142] Specifically, two detection content words with different attributes can be selected to form a word group to be confirmed. It is necessary to confirm whether there is a relationship between these two detection content words. The relationship is then compared with the attributes of the standard detection content with a relationship in the logical mapping relationship rule set. The calculation formula is as follows:
[0143] ;
[0144] Among them, Map(x) i x j () represents two different attribute detection content words x i and x j The result of determining whether there is a correlation is 1, which indicates that the detected content words x have different attributes. i and x j A correlation of 0 indicates that the detected content word x has two different attributes. i and x j There is no correlation, r i and r j Indicates with x i and x j The attributes of two standard detection contents that have the same properties.
[0145] In some embodiments, the relationship includes a subordinate relationship and a causal relationship. For example, the test item "insulation resistance" is subordinate to the value "1000", the value "1000" is subordinate to the unit "Ω", and the test item "insulation resistance" is causally related to the test result "qualified".
[0146] Further, refer to Figure 3 In step S6, a detection event information chain is constructed based on the information association rule set, including:
[0147] S61. Take the detected content words in the information association rule set as nodes and the association relationship as edges. Determine the direction of the edges according to the subordinate and causal relationships between nodes to obtain a directed graph of association relationships.
[0148] S62. Based on the directed graph of the relationship, the PageRank algorithm is used to calculate the node weights to obtain the weight of each node;
[0149] S63. Traverse each node. For each current node, determine the node with the highest weight value among the nodes that are related to the current node as the final related node, and form an initial information chain between the current node and the final related node.
[0150] S64. Based on the pointers between nodes in the initial information chain and the sequential positional relationship of nodes in the preprocessed text, the nodes are integrated and sorted to obtain the detection event information chain.
[0151] Specifically, in step S61, the direction of the edge is determined based on the subordinate and causal relationships between nodes. For subordinate relationships, such as inspection item and value, it can be seen that the inspection item contains the value, so the directed edge between them is the inspection item pointing to the value. If there is a causal relationship between the inspection item and the qualified, then the directed edge between them is the inspection item pointing to the qualified.
[0152] Further, in step S62, the weight of the node is calculated using the following formula:
[0153] ;
[0154] Among them, PR(u i ) represents node u i The weights, d represents the damping factor, IN(u i ) represents all pointers to node u i The set of nodes, u j For set IN(u i A node in ) PR(u j ) represents node u j The weight, Represents node u j The number of nodes it points to. Represents node u j Pointing to node u i The weighting coefficients.
[0155] Further, in step S63, the weight of each node is calculated, and each node is traversed. For each current node, the node with the highest weight among the nodes that are related to the current node is determined as the final associated node, and the current node and the final associated node form an initial information chain. For example, in a directed graph with six nodes A, B, C, D, E, and F, where node A points to nodes B, C, and D, and node B has the highest weight, then node A and node B form an initial information chain. And so on.
[0156] Further, in step S64, the nodes are integrated and sorted according to the pointers between nodes in the initial information chain and the positional relationship of the nodes in the preprocessed text to obtain the detection event information chain.
[0157] First, based on the initial information chain, the nodes are integrated according to their pointing relationships. For example, if the initial information chains are A pointing to B, B pointing to D, C pointing to E, D pointing to C, E pointing to F, and F pointing to C, then integrating them will yield the integrated information chain: A→B→D→C→E→F→A.
[0158] Secondly, the nodes are sorted according to their positions in the preprocessed text: for the integrated information chain, the node at the beginning of the preprocessed text is determined, and the detection event information chain is obtained from the integrated information chain, starting from the node at the beginning of the preprocessed text, based on their pointing relationship.
[0159] If the integrated information chain is a completely continuous loop, then after determining the first node in the preprocessed text as the starting point, the detection event information chain is obtained sequentially based on the pointing relationships of subsequent nodes. For example, in the integrated information chain: A→B→D→C→E→F→A, the first node in the preprocessed text is D. Therefore, taking D as the starting point, the obtained detection event information chain is: D→C→E→F→A→B.
[0160] If the integrated information chain contains both interconnected circular chains and straight chains, the node at the beginning of the circular chain in the preprocessed text is first determined as the starting point to form one event detection information chain. The remaining straight chains then form another event detection information chain. For example, in the integrated information chain: A→B→D→C→E→F→D, where D→C→E→F→D forms a circular chain, and the node at the beginning of the circular chain in the preprocessed text is C, then C→E→F→D is considered one event detection information chain, and A→B is another.
[0161] For example, a chain of test event information can be formed: "Insulation resistance - 1000 - Ω - Pass".
[0162] refer to Figure 4 In some embodiments, an information extraction apparatus for unstructured detection text is provided, comprising:
[0163] Preprocessing module 201 is used to acquire the unstructured detection text to be extracted and preprocess it to obtain preprocessed text, and encode the preprocessed text to obtain a detection text sequence;
[0164] Encoding module 202 is used to input the detected text sequence into the BERT model for deep semantic encoding and output semantic feature vector;
[0165] The classification module 203 is used to input the semantic feature vector into a pre-trained SVM model for classification and recognition, and extract key detection text fragments from the unstructured detection text based on the classification and recognition results.
[0166] Detection module 204 is used to input the key detection text fragment into a pre-trained LSTM-CRF model for attribute recognition and output the detection text attribute information;
[0167] The association module 205 is used to extract relevant standard testing information from industry standards, establish a logical mapping relationship rule set based on the standard testing information, determine the association relationship between each information element in the test text attribute information based on the logical mapping relationship rule set, and form an information association rule set.
[0168] The graph construction module 206 is used to construct a detection event information chain based on the information association rule set.
[0169] Furthermore, the preprocessing module 201 preprocesses the extracted unstructured detection text, including:
[0170] The unstructured detection text is subjected to text recognition to remove redundant words, thereby obtaining the preprocessed text;
[0171] The preprocessing module 201 encodes the preprocessed text to obtain a detection text sequence, including:
[0172] The words in the preprocessed text are encoded to obtain the detection text sequence.
[0173] Furthermore, the BERT model includes a multi-layer Transformer encoder, which is used to perform word embedding processing and encoding on each word in the detected text sequence to generate the semantic feature vector.
[0174] Furthermore, the SVM model is used to determine whether the corresponding word is a keyword based on the semantic feature vector; if the corresponding word is a keyword, the corresponding text segment is extracted as the key detection text segment based on the position of the keyword in the preprocessed text.
[0175] Further, based on the position of the keyword in the preprocessed text, the corresponding text segment is extracted as the key detection text segment, including:
[0176] Locate the position of the keyword in the preprocessed text, search from the position to both ends, stop the search when the end symbol is found, and extract the text between the two end symbols as the initial key detection text segment;
[0177] The obtained initial key detection text fragments are deduplicated to obtain the key detection text fragments.
[0178] Furthermore, the LSTM-CRF model includes an input layer, an embedding layer, a forward LSTM, a backward LSTM, and a CRF layer;
[0179] The input layer is used to receive the key detection text fragments and generate a word sequence based on the key detection text fragments;
[0180] The embedding layer is used to map each word in the word sequence to a word vector to obtain a word vector sequence;
[0181] The forward LSTM is used to process the word vector sequence from left to right according to its position to obtain the forward hidden state, and the backward LSTM is used to process the word vector sequence from right to left according to its position to obtain the backward hidden state.
[0182] The forward hidden state and the backward hidden state are concatenated to obtain the hidden state vector;
[0183] The CRF layer is used to predict the attributes of the hidden state vector, obtain a label sequence of the attribute prediction results, and match the label sequence with words in the key detection text segment to generate detection text attribute information.
[0184] Furthermore, the association module 205 establishes a logical mapping relationship rule set based on the standard detection information, including:
[0185] The standard testing content is extracted from the standard testing information, and the attributes of each standard testing content are determined. Based on the attributes of the standard testing content, the related standard testing content is mapped to form the logical mapping relationship rule set.
[0186] Furthermore, the information elements in the detected text attribute information are detected content words, and each detected content word is set with a corresponding attribute label;
[0187] The association module 205 determines the association relationships between various information elements in the detected text attribute information according to the logical mapping relationship rule set, forming an information association rule set, including:
[0188] Select detection content words with different attributes to form a word group to be confirmed. Compare the attributes of the detection content words in the word group to be confirmed with the attributes of the standard detection content that has a relationship in the logical mapping relationship rule set. If the attributes of the detection content words are the same as the attributes of the standard detection content that has a relationship in the logical mapping relationship rule set, then it is confirmed that the detection content words in the word group to be confirmed have a relationship.
[0189] The information association rule set is formed by acquiring all related words in the detected content.
[0190] Furthermore, the association relationships include subordinate relationships and causal relationships; the construction module 206 constructs a detection event information chain based on the information association rule set, including:
[0191] Using the detected content words in the information association rule set as nodes and the association relationships as edges, the direction of the edges is determined based on the subordinate and causal relationships between nodes, thus obtaining a directed graph of association relationships.
[0192] Based on the directed graph of the aforementioned relationships, the PageRank algorithm is used to calculate the node weights to obtain the weight of each node;
[0193] Traverse each node. For each current node, determine the node with the highest weight value among the nodes that are related to the current node as the final related node, and form an initial information chain between the current node and the final related node.
[0194] The nodes are integrated and sorted according to the pointers between nodes in the initial information chain and the sequential positional relationship of the nodes in the preprocessed text to obtain the detection event information chain.
[0195] The weight of the node is calculated using the following formula:
[0196] ;
[0197] Among them, PR(u i ) represents node u i The weights, d represents the damping factor, IN(u i ) represents all pointers to node u i The set of nodes, u j For set IN(u i A node in ) PR(u j ) represents node u j The weight, Represents node u j The number of nodes it points to. Represents node u j Pointing to node u i The weighting coefficients.
[0198] The information extraction method and apparatus for unstructured detection text provided in the above embodiments have at least the following beneficial effects:
[0199] (1) By combining the BERT model, SVM model and LSTM-CRF model, the key detection text fragments in unstructured detection text were extracted. The BERT model can deeply understand the deep semantics of the text and generate semantic feature vectors. The SVM model locates keywords based on these feature vectors. The LSTM-CRF model further extracts attribute information from the key detection text fragments, such as the name of the detection project, the value, the unit and the detection result. This not only improves the speed of information processing, but also reduces errors caused by human factors and improves the accuracy of information extraction, providing reliable data support for power material management.
[0200] (2) By defining a logical mapping rule set and constructing an information association rule set, the PageRank algorithm is used to calculate the node weights and capture the complex relationships between each information element. This forms a detection event information chain that includes key information from the start of detection to the result determination, enabling power material management personnel to deeply understand the detection events and providing information basis for material quality assessment, fault prediction and power grid safe operation.
[0201] (3) The final obtained detection event information chain only focuses on key detection events in unstructured detection text, which effectively saves the time of power material management personnel and improves production efficiency.
[0202] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if these modifications and modifications of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include these modifications and modifications.
Claims
1. A method for information extraction from unstructured detection text, characterized in that, include: The unstructured detection text to be extracted is obtained and preprocessed to obtain preprocessed text. The preprocessed text is then encoded to obtain a detection text sequence. The detected text sequence is input into the BERT model for deep semantic encoding, and semantic feature vectors are output. The semantic feature vector is input into a pre-trained SVM model for classification and recognition. Based on the classification and recognition results, key detection text fragments are extracted from the unstructured detection text. The key detected text fragments are input into a pre-trained LSTM-CRF model for attribute recognition, and the detected text attribute information is output. Extract relevant standard testing information from industry standards, establish a logical mapping relationship rule set based on the standard testing information, determine the association relationship between each information element in the test text attribute information based on the logical mapping relationship rule set, and form an information association rule set. Based on the information association rule set, a detection event information chain is constructed: the detection content words in the information association rule set are used as nodes, and the association relationships are used as edges. Based on the subordinate and causal relationships between nodes, the direction of the edges is determined to obtain a directed graph of association relationships. Based on the directed graph of the aforementioned relationships, the PageRank algorithm is used to calculate the node weights to obtain the weight of each node; Traverse each node. For each current node, determine the node with the highest weight value among the nodes that are related to the current node as the final related node, and form an initial information chain between the current node and the final related node. The nodes are integrated and sorted according to the pointers between nodes in the initial information chain and the sequential positional relationship of the nodes in the preprocessed text to obtain the detection event information chain; The weight of the node is calculated using the following formula: ; Among them, PR(u i ) represents node u i The weights, d represents the damping factor, IN(u i ) represents all pointers to node u i The set of nodes, u j For set IN(u i A node in ) PR(u j ) represents node u j The weight, Represents node u j The number of nodes it points to. Represents node u j Pointing to node u i The weighting coefficients.
2. The method according to claim 1, characterized in that, The extracted unstructured detection text is preprocessed, including: The unstructured detection text is subjected to text recognition to remove redundant words, thereby obtaining the preprocessed text; The preprocessed text is encoded to obtain a detection text sequence, including: The words in the preprocessed text are encoded to obtain the detection text sequence.
3. The method according to claim 1, characterized in that, The BERT model includes a multi-layer Transformer encoder, which performs word embedding and encoding on each word in the detected text sequence to generate the semantic feature vector.
4. The method according to claim 1, characterized in that, The SVM model is used to determine whether the corresponding word is a keyword based on the semantic feature vector; if the corresponding word is a keyword, the corresponding text segment is extracted based on the position of the keyword in the preprocessed text as the key detection text segment.
5. The method according to claim 4, characterized in that, Based on the location of the keyword in the preprocessed text, the corresponding text fragment is extracted as the key detection text fragment, including: Locate the position of the keyword in the preprocessed text, search from the position to both ends, stop the search when a terminating symbol is found, and extract the text between the two terminating symbols as the initial key detection text fragment; The obtained initial key detection text fragments are deduplicated to obtain the key detection text fragments.
6. The method according to claim 1, characterized in that, A set of logical mapping rules is established based on the standard detection information, including: The standard testing content is extracted from the standard testing information, and the attributes of each standard testing content are determined. Based on the attributes of the standard testing content, the related standard testing content is mapped to form the logical mapping relationship rule set.
7. The method according to claim 6, characterized in that, The information elements in the detected text attribute information are detected content words, and each detected content word is set with a corresponding attribute label; The association relationships between various information elements in the detected text attribute information are determined based on the logical mapping relationship rule set, forming an information association rule set, including: Select detection content words with different attributes to form a word group to be confirmed. Compare the attributes of the detection content words in the word group to be confirmed with the attributes of the standard detection content that has a relationship in the logical mapping relationship rule set. If the attributes of the detection content words are the same as the attributes of the standard detection content that has a relationship in the logical mapping relationship rule set, then it is confirmed that the detection content words in the word group to be confirmed have a relationship. The information association rule set is formed by acquiring all related words in the detected content.
8. An information extraction apparatus for unstructured detection text applied to the method described in any one of claims 1-7, characterized in that, include: The preprocessing module is used to acquire the unstructured detection text to be extracted and preprocess it to obtain preprocessed text, and encode the preprocessed text to obtain a detection text sequence. The encoding module is used to input the detected text sequence into the BERT model for deep semantic encoding and output semantic feature vectors; The classification module is used to input the semantic feature vector into a pre-trained SVM model for classification and recognition, and extract key detection text fragments from the unstructured detection text based on the classification and recognition results. The detection module is used to input the key detection text fragments into a pre-trained LSTM-CRF model for attribute recognition and output the detection text attribute information. The association module is used to extract relevant standard testing information from industry standards, establish a logical mapping relationship rule set based on the standard testing information, determine the association relationship between each information element in the test text attribute information based on the logical mapping relationship rule set, and form an information association rule set. The construction module is used to construct a detection event information chain based on the information association rule set.
Citation Information
Patent Citations
AI-based unstructured test case conversion method and device
CN119645826A
Text information extraction method, device and equipment and readable storage medium
CN112860905A
Performance evaluation method for scientific and technical literature large language model
CN118349818A