Automatic Construction Method of Bridge Maintenance Knowledge Graph Based on Deep Learning

Through deep learning methods, an entity automatic extraction and relationship automatic identification model in the field of bridge management and maintenance was constructed, which solved the problems of low bridge detection information entry efficiency and low degree of automation in knowledge graph construction, and achieved rapid, accurate and reliable automatic construction of bridge management and maintenance knowledge graphs.

CN116450852BActive Publication Date: 2025-06-24SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310461241.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-26
Publication Date
2025-06-24
Estimated Expiration
2043-04-26

AI Technical Summary

Technical Problem

The existing bridge detection information is time-consuming and labor-intensive, storage redundant, complex retrieval, low visualization and utilization efficiency, and the construction method of bridge management knowledge graph is single, relying on professional knowledge and low automation.

Method used

Using a deep learning-based method, the bridge management field is constructed through the BERT+BiLSTM+CRF model, an automatic identification model for the bridge management field relationship is established, and a fully automatic construction of the bridge management knowledge graph is realized.

Benefits of technology

It realizes the rapid, accurate and reliable automatic construction of the bridge management knowledge graph, improves the efficiency of the bridge management knowledge graph construction, reduces organizational difficulty and construction costs, and improves the degree of automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116450852B_ABST
    Figure CN116450852B_ABST
Patent Text Reader

Abstract

The present invention discloses an automatic construction method for a bridge maintenance knowledge graph based on deep learning, including: S1, preprocessing the texts in the field of bridge maintenance and establishing a text annotation database for the field of bridge maintenance; S2, constructing an automatic entity extraction model for the bridge field by using the BERT+BiLSTM+CRF model; S3, establishing an automatic relationship recognition model for the bridge maintenance field: establishing a catalog of entity relationships in the bridge field, using a two-stage relationship recognition algorithm to extract entity pairs from the entities in the bridge maintenance field obtained in step S2, and identifying the matching relationships between the entities in the entity pairs; S4, constructing and visualizing the bridge maintenance knowledge graph. The present invention realizes the automatic construction of the bridge maintenance knowledge graph, solves the problems of large organization difficulty, high construction cost and low automation degree of the traditional knowledge graph construction method in the bridge field, improves the utilization efficiency of bridge detection information, and provides technical support for intelligent bridge maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bridge management and maintenance, and specifically relates to an automatic construction method of a bridge management and maintenance knowledge graph based on deep learning. Background Art

[0002] A large number of conventional and structural inspection reports of bridges contain important information related to structural management and maintenance. However, the traditional input of bridge inspection information is time-consuming and laborious, with redundant storage, complicated retrieval, low visualization level and utilization efficiency. A knowledge graph is a structured semantic knowledge base used to describe concepts in the real world and their interrelationships. It can contain a vast amount of knowledge in a professional field and can simultaneously perform multiple data tasks such as storage, query, processing, and visualization. This makes it an effective means to construct a bridge management and maintenance knowledge graph based on text information such as inspection reports and industry specifications to solve problems such as low efficiency of traditional bridge operation and maintenance management work and poor information utilization rate. At the same time, the bridge management and maintenance knowledge graph can also make full use of and mine the rich information contained in the inspection reports and provide more comprehensive and reliable structural state information. Currently, the construction method of the bridge management and maintenance knowledge graph is single, and limited by the professionalism of bridges. The establishment process of the knowledge graph requires the participation of industry senior experts. In addition, the bridge text information is complicated and the extraction, ablation, and alignment of entities and relationships need to be carried out manually. These problems all lead to great organizational difficulty, high construction cost, and low automation degree in creating the bridge management and maintenance knowledge graph. How to use deep learning technology to automatically, quickly, accurately, and reliably establish a bridge management and maintenance knowledge graph is a technical problem to be solved urgently. Summary of the Invention

[0003] The purpose of the present invention is to solve the related problems of the automatic construction of the bridge management knowledge graph, including the establishment of an annotation database in the bridge management and maintenance field, the establishment of an entity automatic extraction model in the bridge management and maintenance field, the establishment of a relationship automatic recognition model in the bridge management and maintenance field, and the visualization of the bridge management and maintenance knowledge graph and other technical problems. It realizes the full-automatic construction of the bridge management and maintenance knowledge graph based on the bridge management and maintenance text data.

[0004] To achieve the above purpose, the present invention provides the following technical solution: An automatic construction method of a bridge management and maintenance knowledge graph based on deep learning, including the following steps:

[0005] S1. Preprocess the text in the bridge management and maintenance field and establish a text annotation database in the bridge management and maintenance field;

[0006] S2. Use the BERT+BiLSTM+CRF model to construct an entity automatic extraction model in the bridge field: Take the text sequence after cleaning and standardization as the input, and use the Chinese transfer learning pre-trained model BERT to obtain a sequence vector containing semantic features;

[0007] Use the BiLSTM model to capture the bidirectional semantic dependencies of the context information in the text of bridge maintenance and management; use the conditional random field CRF to obtain an optimal prediction sequence through the relationship of adjacent entity tags to identify entities in the field of bridge maintenance and management.

[0008] S3. Establish an automatic relationship recognition model for the field of bridge maintenance and management: establish a catalog of entity relationships in the bridge field, extract the entity pairs of the entities in the field of bridge maintenance and management obtained in step S2, and use a two-stage relationship recognition algorithm to identify the matching relationship between the entities in the entity pair.

[0009] The first stage: Identify the relationship between the entities e i and e j in the field of bridge maintenance and management where there is no text information in the middle. Based on the extracted entities and the preset relationship catalog R, match all the defined entities E and the preset relationship R to establish an entity-relationship-entity tuple T:

[0010] T i =(E i →r i →E i+1 )

[0011] where T i is the E i -r i -E i+1 tuple, E i is the i-th entity, r i is the i-th relationship, and E i+1 is the (i + 1)-th entity; match the entities e i and e j in the entity set after step S2 with all tuples T. The entities within the tuple have a directed relationship, and use a two-way matching algorithm to obtain the relationship r i between the entities, as shown in the following formula:

[0012] r i = match[T i ,(e i ,e j )]

[0013] The second stage: Identify the relationship between a group of adjacent entities in the field of bridge maintenance and management where there is text information in the middle and . First, perform a max pooling operation on the non-entity text between the entities and to obtain a text feature vector If there is no text between the two identified entities, set to 0 to obtain the vector representation of the relationship. Each entity pair obtains two relationship representations, as shown in the following formula:

[0014]

[0015]

[0016] Among them, and are the recognized entity feature vectors;

[0017] Input and these two relationships into a fully connected network, and then activate them using the sigmoid function. This process is expressed as the following formula:

[0018]

[0019] Among them, corresponds to the probabilities of different relationships in the relationship catalog R in the field of bridge maintenance. The position with the largest probability value represents the entity relationship of this group of entities matched relationship r i .

[0020] S4. Establish entity-relationship-entity triples based on the matching relationships between entities in the entity pairs obtained in step S3, and then perform alignment ablation on multiple entities in all entity-relationship-entity triples to obtain the bridge maintenance knowledge graph.

[0021] Furthermore, the aforementioned step S1 includes the following sub-steps:

[0022] S101. Clean and standardize the text in the field of bridge maintenance;

[0023] S102. According to the entity catalog in the field of bridge maintenance, and in accordance with the three-digit sequence annotation method, annotate the beginning part, middle part, and non-entity part of the entities in the text of the field of bridge maintenance.

[0024] Furthermore, the aforementioned step S101 includes the following sub-steps:

[0025] S101-1. Use the Jieba toolkit and a custom dictionary to segment the text in the field of bridge maintenance;

[0026] S101-2. Convert the English expressions in the text information into Chinese expressions, and at the same time remove punctuation marks.

[0027] Furthermore, the aforementioned step S102 includes the following sub-steps:

[0028] S102-1. Establish an entity catalog in the field of bridge maintenance: Define the entities in the knowledge graph of the field of bridge maintenance, and extract the entities in the text information based on the defined entities to create an entity catalog E in the field of bridge maintenance:

[0029]

[0030] Among them, E i is a defined entity, i = {1, 2, …, n}, and n is the number of entities;

[0031] S102-2. According to the entity catalog E in the field of bridge maintenance, using an entity annotation method that combines automatic annotation with the same fields and expert correction, annotate the start, end, and non-entity parts of the text entities related to bridge maintenance to obtain an annotation database for the field of bridge maintenance.

[0032] Furthermore, the aforementioned step S2 includes the following sub-steps:

[0033] S201. Use the Chinese transfer learning pre-trained model BERT to mask some words in the text sequence that has been tokenized in S101-1 and cleaned in S101-2, and then add a special token [CLS] at the beginning of the sequence, and separate sentences with the token [SEP]; input the sequence vector into a bidirectional Transformer encoder for feature extraction; obtain a sequence vector containing semantic features, and the structure of the Transformer encoder is as follows:

[0034]

[0035] Among them, Q, K, V are word vector matrices, and d k is the embedding dimension;

[0036] S202. Use the multi-head attention mechanism to project Q, K, V through multiple different linear transformations, and finally concatenate different Attention results, as shown in the following formulas (3) and (4):

[0037] MultiHead(Q, K, V) = Concat(head1, …, head n )W O

[0038] head i = Attention(QW i Q , KW i K , VW i V )

[0039] Among them, W is the weight matrix, and the model can obtain position information in different spaces;

[0040] S203. The Transformer encoder adds positional encoding before data preprocessing and sums it with the input vector data to obtain the relative position of each word in the sentence. The fully connected feed-forward network of the Transformer encoder consists of two fully connected networks: the activation function of the first layer is ReLU, and the second layer is a linear activation function. The fully connected feed-forward network FFN is expressed as the following formula:

[0041] FFN(Z) = max(0, ZW1 + b1)W2 + b2

[0042] Among them, the output Z of the multi-head attention mechanism, W1 and b1 are the weights and bias vectors of the first fully connected network respectively, and W2 and b2 are the weights and bias vectors of the second fully connected network respectively;

[0043] S204. BiLSTM is used to capture the bidirectional semantic dependencies of the text context information in the field of bridge maintenance; LSTM includes a forget gate, an input gate, an output gate, and a memory Cell structure; the input gate and the forget gate filter out the useless information for entity recognition and pass the useful information to the next moment; the output of the entire structure is obtained by multiplying the output of the memory Cell and the output of the output gate; the sequence is input into the LSTM model, and the output is:

[0044] i t = σ(W xi x t + W hin h t-1 + W ci c t-1 + b i )

[0045] z t = tanh(W xc x t + W hc h t-1 + b c )

[0046] f t = σ(W xf x t + W hf h t-1 + W cf c t-1 + b f )

[0047] c t = f t c t-1 + i t z t

[0048] o t= tanh(W xo x t + W ho h t-1 + W co c t + b0)

[0049] h t = o t tanh(c t )

[0050] where σ is the activation function, W is the weight matrix, b is the bias vector, z t is the content to be added, c t is the updated state at time t, i t , f t , o t are the output results of the input gate, forget gate, and output gate respectively, h t is the output of the entire LSTM cell at time t; BiLSTM takes the forward and backward LSTM for each word sequence respectively, and then combines the outputs at the same moment; for each moment, corresponding to the forward and backward information, the actual output is as follows:

[0051]

[0052] where is the forward output of LSTM, is the backward output of LSTM;

[0053] S205. Calculate the output score of BiLSTM according to the following formula:

[0054] P = Σp t

[0055] where P is the output score matrix of BiLSTM, the size of P is n × k, where k is the number of words and n is the number of entities, and P ij represents the score of the jth entity of the ith word;

[0056] S206. The conditional random field CRF obtains an optimal prediction sequence through the relationship of adjacent entity labels to make up for the shortcoming that BiLSTM is only good at processing long-distance text information but cannot handle the dependencies between adjacent entity labels; for the bridge maintenance text X = (x1, x2,..., x n ) and the predicted bridge maintenance text Y = (y1, y2,..., y n ), its scoring function is as follows:

[0057]

[0058] Among them, A represents the transfer score matrix, and A ij represents the score for entity i to transfer to entity j;

[0059] The probability of generating the predicted bridge maintenance text Y is obtained according to the following formula:

[0060]

[0061] Taking the logarithm of the formula gives the likelihood function of the predicted text sequence:

[0062]

[0063] In the formula, Y represents the predicted labeled text sequence, and Yx represents all possible labeled text sequences. The output sequence with the maximum score is obtained after decoding:

[0064]

[0065] Among them, Y* is the entity in the bridge maintenance field recognized by the model.

[0066] Furthermore, the aforementioned step S3 includes the following sub-steps:

[0067] S301. According to the entity relationship in the bridge maintenance field, establish the entity relationship directory R of the bridge maintenance field:

[0068]

[0069] Among them, r i is the entity relationship in the maintenance field, i = {1, 2,..., m}, and the entity relationship r i is used as the defined relationship group and integrated into the knowledge graph automatic construction model.

[0070] S302. Use the two-stage relationship recognition algorithm to extract the entity pair (e i , e j ) from the bridge maintenance field entities obtained in step S2, and identify the relationship between entity e i and e j .

[0071] Furthermore, the aforementioned step S4 includes the following sub-steps:

[0072] S401. Establish an entity-relationship-entity triple: Establish a directed relationship between the automatically recognized entities (e i , e j ) and the automatically matched relationship r i to form the basic unit of the bridge maintenance knowledge graph - the entity-relationship-entity triple (e i , r i , e j );

[0073] S402. Complete the knowledge graph for bridge management and maintenance: Use cosine similarity to align and eliminate all entities, as shown in the following formula:

[0074]

[0075] where, e i and e k are entities in two different triples of (e i , r i , e i+1 ) and (e k , r k , e k+1 ) respectively;

[0076] S403. According to the cosine value cos(θ) between entities e i and e k , and in combination with a preset value, align and eliminate the entities of e i , e k into one entity e i , so that the two triples (e i , r i , e i+1 ) and (e k , r k , e k+1 ) are fused to form a new graph structure (e i+1 , r i , e i , r k , e k+1 ), and perform iterative loops to align and fuse all triples to construct a unified and complete knowledge graph for bridge management and maintenance;

[0077] S404. Visualize the knowledge graph for bridge management and maintenance: Use all entity-relationship-entity triples and utilize the Neo4j graph database to visualize the knowledge graph for bridge management and maintenance.

[0078] Furthermore, in the aforementioned step S301, the r i is the entity relationship in the field of management and maintenance, i = {1, 2,..., m}, m = 6; the specific entity relationships are defined as: r1 - Location of component, r2 - Diseases generated by component, r3 - Location of disease, r4 - Disease trait category, r5 - Disease trait value, r6 - Suggested measures for disease.

[0079] Furthermore, in the aforementioned step S102: The E iFor defining entities, i = {1, 2, …, n}, n = 8, and the specific entity catalog is defined as: E1 - Bridge components, E2 - Bridge component parts, E3 - Disease categories, E4 - Disease locations, E5 - Disease quantities, E6 - Disease trait categories, E7 - Disease trait values, E8 - Maintenance measures.

[0080] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention utilizes deep learning technology to provide bridge management and maintenance personnel with an automatic construction technology for an automatic, fast, accurate, and reliable bridge management and maintenance knowledge graph, effectively improving the construction efficiency of the bridge management and maintenance knowledge graph and solving the problems of large organization difficulty, high construction cost, and low automation degree of the bridge management and maintenance knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Figure 1 is a flowchart of the automatic construction method for a bridge management and maintenance knowledge graph based on deep learning of the present invention.

[0082] Figure 2 is a diagram of an entity automatic extraction model in the field of bridge management and maintenance.

[0083] Figure 3 is a diagram of a relationship automatic recognition model in the field of bridge management and maintenance.

[0084] Figure 4 is a diagram of an automatic construction model for a bridge management and maintenance knowledge graph.

[0085] Figure 5 is a visualization example diagram of a bridge management and maintenance knowledge graph. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0086] To better understand the technical content of the present invention, specific embodiments are hereby given and described in conjunction with the accompanying drawings as follows.

[0087] In the present invention, various aspects of the present invention are described with reference to the accompanying drawings, and many illustrative embodiments are shown in the drawings. The embodiments of the present invention are not limited to those described in the drawings. It should be understood that the present invention can be implemented by any one of the various concepts and embodiments introduced above and the concepts and implementation manners described in detail below, because the concepts and embodiments disclosed in the present invention are not limited to any implementation manner. In addition, some aspects disclosed in the present invention can be used alone or in any suitable combination with other aspects disclosed in the present invention.

[0088] As Figure 1 shown, the automatic construction method for a bridge management and maintenance knowledge graph based on deep learning of the present invention includes the following steps:

[0089] S1. Preprocess the text in the field of bridge management and maintenance and establish a text annotation database for the field of bridge management and maintenance;

[0090] S2. Use the BERT+BiLSTM+CRF model to build an automatic entity extraction model in the bridge field: take the cleaned and normalized text sequence as input, use the Chinese transfer learning pre-training model BERT to obtain a sequence vector containing semantic features; use the BiLSTM model to capture the bidirectional semantic dependency of the text context information in the bridge maintenance field; use the conditional random field CRF to obtain an optimal prediction sequence through the relationship between neighboring entity labels to identify entities in the bridge maintenance field;

[0091] S3. Establishing an automatic recognition model for bridge maintenance domain relationships: Establishing a bridge domain entity relationship directory, extracting entity pairs of bridge maintenance domain entities obtained in step S2, and using a two-stage relationship recognition algorithm to identify matching relationships between entities in the entity pairs;

[0092] S4. Establish entity-relationship-entity triples according to the matching relationship between entities in the entity pairs obtained in step S3, and then align and ablate multiple entities in all entity-relationship-entity triples to obtain a bridge maintenance knowledge graph.

[0093] Further, as a preferred embodiment of the method for automatically constructing a bridge maintenance knowledge graph based on deep learning of the present invention, step S1 includes the following sub-steps:

[0094] S101. Clean and standardize a large number of bridge maintenance documents, including test reports, industry specifications, expert reports, etc.

[0095] S102. Based on the entity catalog of the bridge management and maintenance field, the entity start part, middle part and non-entity part in the bridge management and maintenance field text are annotated according to the three-digit sequence annotation method to form a bridge management and maintenance field annotation database.

[0096] Further, as a preferred embodiment of the method for automatically constructing a bridge maintenance knowledge graph based on deep learning of the present invention, step S101 includes the following sub-steps:

[0097] S101-1. Use Jieba tool library and custom dictionary to segment text in the field of bridge maintenance. For example, segment "1-20# box girder left side 20 white crystals" into "1-20# box girder left side 20 white crystals"; S101-2. Normalize text information and separate "mm, m, mm" in the text into 2 、m 2 ” and other units are converted into Chinese expressions such as “millimeter, meter, square millimeter, square meter”, and measurement words such as “L, W, S” in the text are converted into Chinese expressions such as “length, width, area”. At the same time, special characters such as period, question mark, exclamation mark, etc. are removed from the text.

[0098] Further, as a preferred embodiment of the automatic construction method of the bridge maintenance knowledge graph based on deep learning of the present invention, step S102 includes the following sub-steps:

[0099] S102-1. Establish an entity directory for the bridge maintenance field: Define the entities of the bridge maintenance knowledge graph, extract the entities in the text information according to the entities in the field, and create an entity directory E for the bridge maintenance field:

[0100]

[0101] Among them, E i is the defined entity, i = {1, 2, …, n}, n = 8, and the specific entity directory is defined as: E1 - Bridge components, E2 - Bridge component parts, E3 - Disease categories, E4 - Disease locations, E5 - Disease quantities, E6 - Disease trait categories, E7 - Disease trait values, E8 - Maintenance measures;

[0102] S102-2. According to the entity directory E of the bridge maintenance field, use an entity annotation method that combines automatic annotation with the same fields and expert correction to annotate the start, end, and non-entity parts of the text entities related to bridge maintenance, and form a bridge maintenance field annotation database. For example, annotate "There are diseases on the bridge pier" as "Bridge (B) pier (I) exists (O) there (O) are (O) diseases (B)".

[0103] Further, as a preferred embodiment of the automatic construction method of the bridge maintenance knowledge graph based on deep learning of the present invention, as Figure 2 shown, step S2 includes the following sub-steps:

[0104] S201. Use the Chinese transfer learning pre-trained model BERT to mask some words in the text sequence that has been segmented in S101-1 and cleaned in S101-2, and then add a special marker [CLS] at the beginning of the sequence, and separate sentences with the marker [SEP]; input the sequence vector into a bidirectional Transformer encoder for feature extraction; obtain a sequence vector containing semantic features, and the structure of the Transformer encoder is as follows:

[0105]

[0106] Among them, Q, K, V are word vector matrices, and d k is the embedding dimension;

[0107] S202. Use the multi-head attention mechanism to project Q, K, V through multiple different linear transformations, and finally splice the different Attention results together, as shown in the following formulas (3) and (4):

[0108] MultiHead(Q, K, V) = Concat(head1, …, head n )W O

[0109] head i = Attention(QW i Q , KW i K , VW i V )

[0110] where W is the weight matrix, and the model can obtain the position information in different spaces;

[0111] S203. The Transformer encoder adds position encoding before data preprocessing and sums it with the input vector data to obtain the relative position of each word in the sentence. The fully connected feed-forward network of the Transformer encoder includes two layers of fully connected networks: the activation function of the first layer is ReLU, and the second layer is a linear activation function; the fully connected feed-forward network FFN is expressed as the following formula:

[0112] FFN(Z) = max(0, ZW1 + b1)W2 + b2 where Z is the output of the multi-head attention mechanism, W1 and b1 are the weights and bias vectors of the first layer of the fully connected network respectively, and W2 and b2 are the weights and bias vectors of the second layer of the fully connected network respectively;

[0113] S204. BiLSTM is used to capture the bidirectional semantic dependencies of the text context information in the field of bridge maintenance; LSTM includes a forget gate, an input gate, an output gate, and a memory Cell structure; both the input gate and the forget gate filter out the useless information for entity recognition and pass the useful information to the next moment; the output of the entire structure is obtained by multiplying the output of the memory Cell by the output of the output gate; the sequence is input into the LSTM model, and the output is:

[0114] i t = σ(W xi x t + W hin h t-1 + W ci c t-1 + b i )

[0115] z t = tanh(W xc x t + W hc h t-1 + b c )

[0116] ft = σ(W xf x t + W hf h t-1 + W cf c t-1 + b f )

[0117] c t = f t c t-1 + i t z t

[0118] o t = tanh(W xo x t + W ho h t-1 + W co c t + b0)

[0119] h t = o t tanh(c t )

[0120] where σ is the activation function, W is the weight matrix, b is the bias vector, z t is the content to be added, c t is the updated state at time t, i t , f t , o t are the output results of the input gate, forget gate, and output gate respectively, h t is the output of the entire LSTM cell at time t; BiLSTM takes the forward and backward LSTM for each word sequence respectively, and then combines the outputs at the same time; for each time, corresponding to the forward and backward information, the actual output is as follows:

[0121]

[0122] where is the forward output of LSTM, is the backward output of LSTM;

[0123] S205. Calculate the output score of BiLSTM according to the following formula:

[0124] P = Σp t

[0125] where P is the output score matrix of BiLSTM, the size of P is n×k, where k is the number of words and n is the number of entities, P ij represents the score of the j-th entity of the i-th word;

[0126] S206. The conditional random field CRF obtains an optimal prediction sequence through the relationship of adjacent entity labels to make up for the disadvantage that BiLSTM is only good at processing long-distance text information but cannot handle the dependencies between adjacent entity labels. For the bridge maintenance text X = (x1, x2, …, x n ) and the predicted bridge maintenance text Y = (y1, y2, …, y n ), its scoring function is as follows:

[0127]

[0128] Among them, A represents the transition score matrix, and A ij represents the score for entity i to transition to entity j;

[0129] The probability of generating the predicted bridge maintenance text Y is obtained according to the following formula:

[0130]

[0131] Taking the logarithm of the formula to obtain the likelihood function of the predicted text sequence:

[0132]

[0133] In the formula, Y represents the predicted labeled text sequence, and Yx represents all possible labeled text sequences. The output sequence with the maximum score is obtained after decoding:

[0134]

[0135] Among them, Y* is the entity in the bridge maintenance field after being recognized by the model.

[0136] Furthermore, as a preferred embodiment of the method for automatically constructing a bridge maintenance knowledge graph based on deep learning in the present invention, step S3 includes the following sub-steps:

[0137] S301. Establish a directory of entity relationships in the bridge maintenance field: According to the professional knowledge in the bridge field, determine the entity relationships in the bridge maintenance field and create a domain relationship directory R:

[0138]

[0139] Among them, r i is the entity relationship in the maintenance field, i = {1, 2, …, m}, and m = 6. The specific entity relationships are defined as: r1 - location where the component is located, r2 - diseases generated by the component, r3 - location where the disease is located, r4 - type of disease trait, r5 - numerical value of disease trait, r6 - recommended measures for the disease. The entity relationship r iAs a defined relationship group, it is integrated into the knowledge graph automatic construction model;

[0140] S302, such as Figure 3 shown, establish an automatic relationship recognition model for the bridge maintenance field. Through the predefined maintenance field relationship catalog R, using a two-stage relationship recognition algorithm, extract entity pairs (e i , e j ) from the bridge maintenance field entities obtained in step S2, and identify the relationship between entity e i and e j as follows:

[0141] The first stage: Identify the relationship between bridge maintenance field entities e i and e j without text information in the middle. Based on the extracted entities and the predefined relationship catalog R, match all the defined entities E and the predefined relationship R to establish an entity-relationship-entity tuple T:

[0142] T i = (E i → r i → E i+1 )

[0143] Among them, T i is the E i -r i -E i+1 tuple, E i is the i-th ontology, r i is the i-th relationship, E i+1 is the (i + 1)-th ontology; Match the entities e i and e j in the entity set after step S2 with all tuples T. The entities within the tuple have a directed relationship, and use a two-way matching algorithm to obtain the relationship r i between the entities, as shown in the following formula:

[0144] r i = match[T i , (e i , e j )]

[0145] For example: The text information "There are multiple longitudinal cracks in the 2# main girder" after step S2 is:

[0146]

[0147] Select the entity pair (e2, e4), and perform two-way matching of the entity pair with T:

[0148]

[0149] Obtain the relationship of this pair of entities (e2 main girder, e4 longitudinal crack) as r2 (diseases generated by components);

[0150] The second stage: Identify the relationship between a group of adjacent entities in the field of bridge maintenance and management with text information in between and First, perform max pooling operation on the non-entity text between entities and e j * to obtain the text feature vector If there is no text between the two identified entities, set to 0 to obtain the vector representation of the relationship. Since the relationship is asymmetric, each entity pair gets two relationship representations, as shown in the following formula:

[0151]

[0152]

[0153] Among them, and are the identified entity feature vectors;

[0154] Input and these two relationships into a fully connected network and then activate them using the sigmoid function. This process is represented as follows:

[0155]

[0156] Among them, corresponds to the probabilities of different relationships in the relationship directory R of the bridge maintenance and management field. The position with the largest probability value represents the entity relationship that is the relationship r matched by this group of entities i .

[0157] For example, the text information "The main disease of the 10# main girder is longitudinal crack" after step S2 is as follows:

[0158]

[0159] Perform max pooling on the text information c "The main disease is" between entities (e2, e3) and automatically identify it through the above algorithm to obtain the relationship of this pair of entities (e2 main girder, e4 longitudinal crack) as r2 (diseases generated by components).

[0160] Furthermore, as a preferred embodiment of the method for automatically constructing a bridge maintenance and management knowledge graph based on deep learning of the present invention, step S4 includes the following sub-steps:

[0161] S401. As shown in Figure 4 , construct entity-relationship-entity triples. Establish a directed relationship between the automatically recognized entities (e i , e j ) and the automatically matched relationship r i to form the basic unit of the bridge maintenance knowledge graph - the "entity-relationship-entity" triple (e i , r i , e j );

[0162] S402. Bridge maintenance graph completion: Use cosine similarity to align and fuse all entities:

[0163]

[0164] Among them, e i and e k are entities in two different triples of (e i , r i , e i+1 ) and (e k , r k , e k+1 ). If the cosine value cos(θ) between entities e i and e k is closer to 1, it indicates that the two entities are similar. Align and fuse the entities e i , e k into one entity e i , so that the two triples (e i , r i , e i+1 ) and (e k , r k , e k+1 ) are fused to form a new graph structure (e i+1 , r i , e i , r k , e k+1 ). Perform iterative loops to align and fuse all triples to construct a unified and complete bridge maintenance knowledge graph;

[0165] S403. As shown in Figure 5 , visualize the bridge maintenance knowledge graph. Use all "entity-relationship-entity" triples and use the Neo4j graph database to visualize the bridge maintenance knowledge graph.

[0166] Although the present invention has been described above with reference to preferred embodiments, it is not intended to limit the present invention. Those of ordinary skill in the art to which the present invention pertains may make various modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the claims.

Claims

1. An automatic construction method for a knowledge graph of bridge management and maintenance based on deep learning, characterized in that It includes the following steps: S1. Preprocess the text in the field of bridge maintenance and management, and establish a text annotation database for the field of bridge maintenance and management; S2. Build an automatic entity extraction model for the bridge field using the BERT+BiLSTM+CRF model: Take the text sequence after cleaning and normalization as the input, and use the Chinese transfer learning pre-trained model BERT to obtain a sequence vector containing semantic features; Use the BiLSTM model to capture the bidirectional semantic dependencies of the context information in the text of the bridge maintenance and management field; Use the conditional random field CRF to obtain an optimal prediction sequence through the relationship of adjacent entity labels, and identify the entities in the bridge maintenance and management field; S3. Build an automatic relationship recognition model for the bridge maintenance and management field: Establish a catalog of entity relationships in the bridge field, extract the entity pairs of the entities in the bridge maintenance and management field obtained in step S2, and use a two-stage relationship recognition algorithm to identify the matching relationship between the entities in the entity pair, specifically as follows: The first stage: Identify the relationship between the entities e in the field of bridge maintenance where there is no text information in the middle. According to the extracted entities and the preset relationship directory R, match all the defined entities E and the preset relationships R, and establish the entity-relationship-entity tuple T: i and e j Identify the relationship between them. Based on the extracted entities and the preset relationship directory R, match all the defined entities E and the preset relationships R, and establish the entity-relationship-entity tuple T: T i = (E i → r i → E i+1 ) Among them, T i is E i -r i -E i+1 tuple, E i is the i-th entity, r i is the i-th relationship, E i+1 is the (i + 1)-th entity; match the entities e i and e j in the entity set after step S2 with all tuples T. There is a directed relationship between the entities within the tuple. The relationship r i between the entities is obtained by using a two-way matching algorithm, as shown in the following formula: r i = match[T i ,(e i ,e j )] The second stage: For a group of adjacent entities in the field of bridge maintenance where there is text information in between and to identify the relationship between them, first perform max pooling on the non-entity text between entities and to obtain a text feature vector If there is no text between the two identified entities, set to 0 to obtain the vector representation of the relationship. Each entity pair gets two relationship representations, as shown in the following formula: Among them, and are the recognized entity feature vectors; Input and These two relationships into a fully connected network, and then use the sigmoid function for activation. This process is expressed as follows: Among them, corresponding to the probabilities of different relationships in the relationship catalog R in the field of bridge maintenance, the entity relationship represented by the position with the largest probability value is the entity of this group matched relationship r i ; S4. Establish an entity-relationship-entity triple based on the matching relationship between the entities in the entity pair obtained in step S3, and then perform alignment ablation on multiple entities in all the entity-relationship-entity triples to obtain the bridge maintenance and management knowledge graph.

2. The automatic construction method of the bridge maintenance knowledge graph based on deep learning according to claim 1, wherein Step S1 includes the following sub-steps: S101. Clean and normalize the text in the field of bridge maintenance and management; S102. According to the entity catalog in the field of bridge maintenance and management, use the three-digit sequence annotation method to annotate the start part, middle part and non-entity part of the entities in the text of the bridge maintenance and management field.

3. The automatic construction method of the bridge maintenance knowledge graph based on deep learning according to claim 2, characterized in that Step S101 includes the following sub-steps: S101-1. Use the Jieba tool library and custom dictionary to segment the text in the field of bridge maintenance and management; S101-2. Convert the English expressions in the text information into Chinese expressions, and at the same time remove punctuation marks.

4. The automatic construction method of the bridge maintenance knowledge graph based on deep learning according to claim 2, characterized in that Step S102 includes the following sub-steps: S102-1. Establish an entity catalog in the field of bridge maintenance and management: Define the entities of the knowledge graph in the field of bridge maintenance and management, and extract the entities in the text information according to the defined entities to create an entity catalog E in the field of bridge maintenance and management: Among them, E i is a defined entity, i = {1, 2, …, n}, where n is the number of entities; S102-2. According to the entity catalog E in the field of bridge maintenance and management, use an entity annotation method that combines automatic annotation with the same field and expert correction to annotate the start, end and non-entity parts of the entities related to bridge maintenance and management, and obtain an annotation database in the field of bridge maintenance and management.

5. The automatic construction method of the bridge maintenance knowledge graph based on deep learning according to claim 1, wherein Step S2 includes the following sub-steps: S201. Use the Chinese transfer learning pre-trained model BERT to mask some words in the text sequence that has been segmented in S101-1 and cleaned in S101-2, and then add a special marker [CLS] at the beginning of the sequence, and separate sentences with the marker [SEP]; Input the sequence vector into the bidirectional Transformer encoder for feature extraction; Obtain a sequence vector containing semantic features, and the structure of the Transformer encoder is as follows: Among them, Q, K, and V are word vector matrices, and d k is the embedding dimension; S202. Use the multi-head attention mechanism to project Q, K, V through multiple different linear transformations, and finally splice the different Attention results together, as shown in formulas (3) and (4) below: MultiHead(Q,K,V)=Concat(head1,…,head n )W O head i = Attention(QW i Q ,KW i K ,VW i V ) Among them, W is the weight matrix, and the model can obtain the position information in different spaces; S203. The Transformer encoder adds positional encoding before data preprocessing and sums it with the input vector data to obtain the relative positions of each word in the sentence. The fully connected feed-forward network of the Transformer encoder includes two layers of fully connected networks: the activation function of the first layer is ReLU, and the second layer is a linear activation function; the fully connected feed-forward network FFN is expressed as the following formula: FFN(Z) = max(0, ZW1 + b1)W2 + b2 Among them, Z is the output of the multi-head attention mechanism, W1 and b1 are the weights and bias vectors of the first layer of the fully connected network respectively, and W2 and b2 are the weights and bias vectors of the second layer of the fully connected network respectively; S204. BiLSTM is used to capture the bidirectional semantic dependencies of the text context information in the field of bridge maintenance; LSTM includes a forget gate, an input gate, an output gate, and a memory Cell structure; both the input gate and the forget gate filter out the useless information for entity recognition and pass the useful information to the next moment; the output of the entire structure is obtained by multiplying the output of the memory Cell by the output of the output gate; the sequence is input into the LSTM model, and the output is: i t = σ(W xi x t + W hi h t-1 + W ci c t-1 + b i ) z t = tanh(W xc x t + W hc h t-1 + b c ) f t = σ(W xf x t + W hf h t-1 + W cf c t-1 + b f ) c t = f t c t-1 + i t z t o t = tanh(W xo x t + W ho h t-1 + W co c t + b0) h t = o t tanh(c t ) where, σ is the activation function, W is the weight matrix, b is the bias vector, and z t is the content to be added, c t is the updated state at time t, i t , f t , o t are the output results of the input gate, forget gate, and output gate respectively, and h t is the output of the entire LSTM cell at time t; BiLSTM takes the forward and backward LSTM for each word sequence respectively, and then combines the outputs at the same moment; for each moment, corresponding to the forward and backward information, the actual output is as follows: Among them, is the forward output of the LSTM, is the backward output of the LSTM; S205. Calculate the output score of BiLSTM according to the following formula: P = ∑p t Among them, P is the output score matrix of the BiLSTM. The size of P is n×k, where k is the number of words and n is the number of entities. P ij represents the score of the j-th entity of the i-th word; S206. The conditional random field CRF obtains an optimal prediction sequence through the relationship of adjacent entity tags to make up for the shortcoming that BiLSTM is only good at processing long-distance text information but cannot handle the dependencies between adjacent entity tags. For the bridge maintenance text X = (x1, x2, …, x n ) and the predicted bridge maintenance text Y = (y1, y2, …, y n ), its scoring function is as follows: where A represents the transition score matrix, and A ij represents the score of entity i transitioning to entity j; Obtain the probability of generating the predicted bridge maintenance text Y according to the following formula: Take the logarithm of the formula to obtain the likelihood function of the predicted text sequence: In the formula, Y represents the predicted labeled text sequence, and Yx represents all possible labeled text sequences; the output sequence with the maximum score is obtained after decoding: Among them, Y* is the entity in the field of bridge maintenance after being recognized by the model.

6. The method for automatically constructing a knowledge graph for bridge maintenance and management based on deep learning according to claim 5, wherein Step S3 includes the following sub-steps: S301. According to the entity relationships in the field of bridge maintenance, establish the entity relationship directory R in the field of bridge maintenance: Among them, r i is the entity relationship in the maintenance field, i = {1, 2, …, m}, and the entity relationship r i is used as a defined relationship group and integrated into the knowledge graph automatic construction model; S302. Using a two-stage relationship recognition algorithm, extract entity pairs (e i , e j ) from the bridge maintenance domain entities obtained in step S2, and identify the relationship between entity e i and e j .

7. The automatic construction method of the bridge maintenance knowledge graph based on deep learning according to claim 6, characterized in that Step S4 includes the following sub-steps: S401. Establish entity-relationship-entity triples: The automatically recognized entities (e i , e j ) and the automatically matched relationship r i are used to establish a directed relationship, forming the basic unit of the bridge maintenance knowledge graph - entity-relationship-entity triples (e i , r i , e j ); S402. Complete the bridge maintenance knowledge graph: Use cosine similarity to align and ablate all entities, as shown in the following formula: wherein, e i and e k are respectively entities in two different triples of (e i , r i , e i+1 ) and (e k , r k , e k+1 ); S403. According to the cosine value cos(θ) between entities e i and e k , and combined with a preset value, perform ablation alignment on e i ,e k entities to become one entity e i , so that the two triples (e i ,r i ,e i+1 ) and (e k ,r k ,e k+1 ) are fused to form a new graph structure (e i+1 ,r i ,e i ,r k ,e k+1 ), perform iterative loops, align and fuse all triples, and construct a unified and complete bridge maintenance graph; S404. Visualize the bridge maintenance knowledge graph: Use all entity-relationship-entity triples and use the Neo4j graph database to visualize the bridge maintenance knowledge graph.

8. The automatic construction method of the bridge maintenance knowledge graph based on deep learning according to claim 6, characterized in that In step S301, the r i is an entity relationship in the field of management and maintenance, i = {1, 2, …, m}, m = 6; The specific entity relationships are defined as: r1 - location of the component, r2 - diseases generated by the component, r3 - location of the disease, r4 - category of disease traits, r5 - numerical value of disease traits, r6 - recommended measures for diseases.

9. The automatic construction method of the bridge maintenance knowledge graph based on deep learning according to claim 4, characterized in that, In step S102: the E i is a defined entity, i = {1, 2, …, n}, n = 8, and the specific entity directory is defined as: E1 - bridge component, E2 - bridge component part, E3 - disease category, E4 - disease location, E5 - disease quantity, E6 - disease character category, E7 - disease character value, E8 - maintenance measure.