A method and device for evaluating construction accident reports
By building accident trees and knowledge graphs to evaluate the accuracy and depth of accident reports in construction industry, the problem of difficulty in automatic assessment of accident reports in the existing technology is solved, and the quality and safety of reports are improved.
Patent Information
- Application Number
- CN202410984522.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-22
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2044-07-22
AI Technical Summary
It is difficult for the existing technology to achieve automatic quality assessment of construction accident reports, resulting in uneven reporting quality, which may cause incorrect risk assessment and preventive measures, increasing the risk of re-accidents.
Use the logical relationship in the accident tree to evaluate the accuracy and analysis depth of accident reports, and build a knowledge graph and accident tree to evaluate the accuracy and analysis depth of the report by obtaining the entities and relationships between entities in the accident report.
It improves the quality of construction accident reports, reduces long-term safety risks, provides a basis for reporting optimization, and ensures the accuracy and in-depth analysis of causal logic.
Smart Images

Figure CN119066196B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of construction industry safety accident analysis, and particularly to a method and device for evaluating construction industry accident reports. Background Art
[0002] Construction safety issues are important issues affecting life and property, and accident reports are important tools for accident analysis and prevention. The quality of accident reports is directly related to the accuracy of accident cause analysis and the effectiveness of subsequent preventive measures, and will also affect subsequent enterprise training and education, risk management and improvement, safety culture construction, etc. based on accident reports.
[0003] However, since report writers often face challenges such as a large amount of information, numerous professional terms, and insufficient personal experience, it increases the difficulty of report writing and the possibility of errors, and there is a situation where it is difficult to objectively restore the facts, resulting in uneven quality of existing accident reports, which may lead to incorrect risk assessments and preventive measures, thereby increasing the risk of similar accidents occurring again.
[0004] Therefore, how to achieve automatic evaluation and optimization of the quality of construction industry accident reports is an important aspect of improving construction safety. In the prior art, Patent CN114580978A provides an environmental impact assessment report quality inspection system and method, and Patent CN116050872A provides a method for evaluating the quality of scientific and technological reports based on the analytic hierarchy process, neither of which can achieve automatic evaluation of the quality of construction industry accident reports. Summary of the Invention
[0005] The embodiments of the present invention provide a method and device for evaluating construction industry accident reports, which use the logical relationships in the fault tree to evaluate the accuracy and analysis depth of accident reports, so as to evaluate report quality.
[0006] In a first aspect, the embodiments of the present invention provide a method for evaluating construction industry accident reports, including:
[0007] Obtain a construction industry accident report to be evaluated;
[0008] Extract multiple entities and the relationships between entities in the accident report, where the entity categories include accident types and accident causes, and the relationships between entities include causal relationships;
[0009] Evaluate the accuracy and analysis depth of the accident report according to the levels and connection paths of two entities with causal relationships in the fault tree;
[0010] Wherein, the fault tree takes the accident type entity as the root node, and the accident cause entity as the intermediate node and leaf node.
[0011] Second aspect, embodiments of the present invention provide an electronic device, which includes:
[0012] One or more processors;
[0013] A memory for storing one or more programs,
[0014] When the one or more programs are executed by the one or more processors, the one or more processors implement the construction accident report assessment method described in any embodiment.
[0015] Embodiments of the present invention provide an assessment method for construction accident reports. By using the accident type as the root node, the accident causes as the intermediate nodes and leaf nodes of the accident tree, the causal factors and causal logic of most construction accidents are reflected, providing a knowledge reference for the causal logic of construction accidents; then mapping the content and matching the structure of the report to be evaluated with the accident tree, and judging the accuracy and analysis depth of the report by analyzing the level and connection path of the report entity in the accident tree, providing a basis for report optimization, effectively improving the quality of construction accident reports and reducing long-term safety risks. Description of the Drawings
[0016] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 Is a flowchart of an assessment method for construction accident reports provided by an embodiment of the present invention;
[0018] Figure 2 Is a flowchart of another assessment method for construction accident reports provided by an embodiment of the present invention;
[0019] Figure 3 Is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments
[0020] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0021] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0022] In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0023] Figure 1 is a flowchart of a method for evaluating construction accident reports provided by an embodiment of the present invention. This method is applicable to the situation of evaluating the quality of construction accident reports, especially for evaluating the accuracy and analysis depth of the reports. This method is executed by an electronic device, such as Figure 1 shown, and specifically includes the following steps:
[0024] S110. Obtain the construction accident report to be evaluated.
[0025] This report can be provided by the user, for example, uploaded by the user through a Web interface as the object of evaluation.
[0026] S120. Extract multiple entities and the relationships between entities in the accident report. Among them, the entity categories include accident types and accident causes, and the relationships between entities include causal relationships.
[0027] The entities here refer to the main concepts and elements in the accident report, including accident types, accident risks, occurrence locations, operation links, responsible parties, accident prevention, as well as rectification measures and accident causes. Among them, the accident cause category can be further divided into categories such as unsafe operation behaviors, unsafe management behaviors, unsafe states of objects, personal states, and organizational safety cultures.
[0028] It should be noted that there is a difference between an entity and an entity category, while the relationship between entities is consistent with the relationship between the categories to which the entities belong. Exemplarily, assuming there are two entities, "installer" and "weak safety awareness", in a report, and their categories are "responsible entity" and "personal status" respectively, the entity is different from the entity category; while the relationship between the entity "installer" and "weak safety awareness" is "has", and the relationship between the entity categories "responsible entity" and "personal status" is also "has", and the two are the same. Among all the above entity categories, accident type and accident cause are the essential entity categories in accident reports, and the relationship between these two entity categories, as well as the relationship between these two types of entities, is a causal relationship.
[0029] In a specific embodiment, a set of entity categories and a set of relationships between entity categories in the field of construction accidents can be predefined, with the aim of establishing a conceptual system in the field of construction accident report analysis, clarifying the main concepts / elements and their structural relationships. In this embodiment, these two sets are referred to as the knowledge ontology model in the field of construction accidents, providing a basis for the automatic extraction of entities and entity relationships. Optionally, the predefined set of entity categories is shown in Table 1:
[0030] Table 1 Set of Entity Categories
[0031]
[0032]
[0033] Correspondingly, the set of possible relationships between the entity categories in Table 1 is shown in Table 2. This set can be obtained based on the analysis of the text features of accident reports and expert opinions, covering one or more relationships for each possible combination of entity categories. Table 2 only shows a part of them. Among them, relationship names such as "attributed to", "affect", "be affected by", "cause", "be responsible for", "be implemented by", "implement", "be caused by", etc. can all be classified as "causal relationships".
[0034] Table 2 Set of Relationships between Entity Categories
[0035]
[0036]
[0037]
[0038]
[0039]
[0040] Optionally, Protégé can be used to define the above ontology model. Protégé is open-source ontology editing software that provides a powerful ontology mechanism and can perform advanced construction of ontology modules, including parent-child classification through methods such as "nested classes" and "synonyms". It also has a built-in Reasoner for automated reasoning. In addition, it uses an easy-to-use graphical user interface, making the construction, editing, and management of ontologies simple and clear. In summary, Protégé is a powerful and user-friendly visual ontology editing tool and an efficient tool for knowledge management.
[0041] Based on the above knowledge ontology, in this embodiment, entities, entity categories, and relationships between entities in existing construction accident reports and / or regulation documents can be pre-annotated, and then the BERT-Bi-LSTM-CRF model can be trained using the annotated dataset so that the trained model can be applied in S120 to automatically extract entities and relationships between entities in the report to be evaluated. To illustrate this process, the BERT-Bi-LSTM-CRF model is introduced first. This model is a deep learning structure for sequence labeling tasks and can be used to perform complex text processing tasks. It combines three parts: BERT (Bidirectional Encoder Representations from Transformers), BiLSTM (Bidirectional Long Short-Term Memory), and CRF (Conditional Random Field). Among them, BERT is a Transformer-based model for pre-training in natural language processing. By reading the entire sentence (instead of just unidirectionally), it can understand the meaning of words in different contexts, which makes BERT particularly effective in dealing with the diversity and complexity of word meanings. BiLSTM is a special type of recurrent neural network for processing sequence data. Different from the traditional unidirectional LSTM (Long Short-Term Memory), BiLSTM processes both the past (forward) and future (backward) information of the sequence, using forward and backward LSTMs to capture the left and right contexts in the word sequence. CRF is a statistical modeling method for structured prediction. It considers the dependencies between annotation results and predicts the sequence of output labels. In the named entity recognition task, the CRF layer is used to decode the output of the BiLSTM layer and determine the best label sequence. Combining the above three technologies, the BERT-BiLSTM-CRF model can effectively handle complex natural language processing tasks, especially in scenarios that require context understanding and sequence label prediction.
[0042] In a specific embodiment, to reduce the burden of annotating a large amount of accident report data, transfer learning and semi-supervised learning methods can be used to train the BERT-BiLSTM-CRF model so that it can automatically extract entities from accident reports. Specifically, the training process may include the following steps:
[0043] Step 1: Annotate existing construction accident reports and regulatory documents according to the predefined set of entity categories and the set of relationships between entity categories in the construction accident domain. The evaluation of the quality of accident reports depends not only on common sense knowledge but also on reasoning knowledge. Therefore, in this embodiment, a large amount of construction accident reports and regulatory documents are jointly used as the training data for the entity extraction model to comprehensively learn the entities and relationships in accident reports. Among them, accident report data contains rich information, including key information such as accident types and hazards, operation links and occurrence locations, accident causes and treatment measures, etc. These data can be used to discover hidden patterns and associations to improve the safety management of the construction site; institutional document data contains detailed information such as specific requirements for each operation link, job responsibilities of each responsible entity, and enterprise safety culture construction. These information effectively supplement the accident report data and can improve the model's learning ability of construction industry knowledge.
[0044] Optionally, 20% of the accident report data and 20% of the regulatory documents are respectively selected for annotation to ensure that the selected accident report data covers various accident scenarios, entity categories, and relationships, and the selected regulatory documents cover different types, entity categories, and relationships. The annotation content is shown in Tables 1 and 2. During the annotation process, the category of each entity can be annotated first according to the set of entity categories in Table 1. Assume that the text Ti consists of a series of words W1, W2,..., Wn, and the entity in the annotated text is Ei=(si, ei, "Ci"), where Ei represents the entity, si represents the starting position of the entity in the text, ei represents the ending position, and Ci represents its entity category. Exemplarily, if the accident report states that "due to the inadequate safety training of the manager, the workers have weak safety awareness of high-altitude operations", the annotation result of the entity category is:
[0045] E1=(6, 9, "SafetyCulture")
[0046] E2=(22, 25, "PersonalState")
[0047] Then, according to the category of each entity and the set of relationships between entity categories shown in Table 2, the relationships between entities are annotated. Exemplarily, the relationship between "SafetyCulture" and "PersonalState" in Table 2 is "influences", so the relationship between the above entities E1 and E2 is:
[0048] R1 = (E1, E2, “influences”)
[0049] Step 2: Use the labeled data to train the BERT - BiLSTM - CRF model. During the training process, each paragraph of text in the construction accident report or regulation document is sequentially input into the BERT - BiLSTM - CRF model. Calculate the loss function according to the model state and continuously update the model parameters to make the model output gradually approach the labeled entities and the relationships between entities in each paragraph of text.
[0050] Next, the internal data processing process of the model after a paragraph of text is input into the BERT - BiLSTM - CRF model is introduced. Optionally, the model first uses the tokenizer of BERT to split the text into Tokens and adds special markers (B - Entity, I - Entity) based on the Tokens to indicate the boundaries of entities. Use BERT for embedding to generate word vector representations of the text (including entities and context). That is, for each Token Xi, BERT outputs an embedding vector Di. The embedding representation of the text obtained through the BERT layer provides a basis for capturing local and global relationships between entities in subsequent steps.
[0051] Then, the output of BERT (i.e., the embedding vector Di of each Token) is used as the input of BiLSTM. BiLSTM captures context information from left to right and from right to left through its forward and backward LSTM layers respectively. This helps to understand the roles and interrelationships of entities in the entire document or a long text segment and further captures the sequential dependencies between Tokens. For the i - th Token in the input sequence, the output of BiLSTM can be expressed as where and are the hidden states of the forward LSTM and the backward LSTM at position i respectively.
[0052] Optionally, in BiLSTM, to enable the model to focus on the most important parts for the current task, an attention layer is added to help the model focus on keywords or phrases of key entities and relationships. This weighted information focusing helps to improve the performance of the model, especially when dealing with long texts containing a large amount of irrelevant information. The self - attention mechanism dynamically assigns the “attention” of each Token by analyzing and evaluating the interactions of each Token in the input sequence and their contribution degrees to the current task. Let the output of BiLSTM be T = [[H1, H2,..., H n , where H iis the hidden state vector of the i-th Token in the sequence. The self-attention mechanism first evaluates the relationships between Tokens by calculating the interactions between each Token in the sequence and all other Tokens. For each Token H in the sequence i , the model learns to generate three types of vectors: query vector (Query, Q), key vector (Key, K), and value vector (Value, V). Next, the model obtains the attention scores by calculating the dot products of each query vector (Q) and key vector (K) with all key vectors, and evaluates the influence degree of each Token on other Tokens in the sequence. If the scores are regarded as a kind of "relationship strength", then the higher the score, the more important the relationship between the two Tokens. The score calculation formula is:
[0053] Attention(Q,K)=QK T
[0054] Secondly, the Softmax function is used to normalize the scores to obtain the attention weights of each Token to other Tokens. The normalization formula is:
[0055]
[0056] where d k is the dimension of the key vector, and this scaling factor helps prevent the problem of gradient vanishing caused by the dot product being too large.
[0057] Finally, the model uses the value vector (Value, V) for weighted summation to obtain the final representation of each Token, which is the output of the BiLSTM. This representation not only contains the information of the Token itself, but also integrates the information of other Tokens in the sequence, especially those Tokens that are closely related to it. The formula for weighted summation is:
[0058]
[0059] The CRF layer receives the output of the BiLSTM with attention weighting and predicts the entity category label of each Token through the global optimal decoding strategy to ensure that the sequence labels are optimal as a whole.
[0060] In addition, a relation classifier is added to predict the relation types between entities by using the extracted entity and text features. First, the representation of each entity pair (i.e., the attention-weighted BiLSTM output) and the representation of their common context (the average of the BiLSTM outputs of the Tokens between the entities) are merged into a feature vector through concatenation. Second, a fully connected layer is used as the classifier, taking the feature vector as the input of the relation classifier. The relation classifier is trained using the training data (feature vectors and relation labels) prepared in the previous steps, and outputs the possible relation types between entity pairs.
[0061] Next, we look inside the model to illustrate the overall training process of the model. Optionally, first, based on a large-scale pre-trained language model, the pre-trained model is fine-tuned using the labeled construction accident report data. Large-scale pre-trained language models have strong generalization learning capabilities, but their direct application to professional fields may not yield good results. This is because there are differences in the distribution and pragmatic features between the general corpus used in pre-training and the professional field data, and it is difficult for the model to take into account the subtle semantics and usages in the professional field. The fine-tuning process is a method of re-training the model parameters on the basis of the pre-trained model using the labeled corpus in the professional field, which may include the following steps: First, load the weights of the BERT model pre-trained on the general language task and fix the parameters of the BERT layer. Second, unlock and train the task-related layers (BiLSTM, attention layer, CRF, and relation classifier), define the loss functions for entity recognition and relation classification, and train on the labeled data. Subsequently, gradually unfreeze the BERT parameters for training and updating. The CRF loss for entity recognition and the cross-entropy loss for relation classification are used as the optimization objectives. The model minimizes the total loss function through the gradient descent method to adjust its parameters to adapt to the specific requirements of construction accident reports. The loss function L is defined as:
[0062]
[0063] where x i is the input of the model, y i is the corresponding labeled entity or relation, and p(y i |x i ) is the probability predicted by the model.
[0064] Through the above fine-tuning, the parameters of the model can gradually adapt to and learn the specific pragmatic features of professional field report texts, such as industry terms, event description methods, etc., enabling the model to better process and understand construction report texts while maintaining its general understanding ability, and generating more accurate analysis results. This helps to improve the efficiency of the model in the application of the construction professional field.
[0065] The fine-tuned BERT-BiLSTM-CRF model is further applied to the unlabeled construction accident report data for predicting entity categories and entity relationships. The pseudo-labels generated during this process are used to expand the initial labeled dataset, forming a more abundant and diverse training dataset. The pseudo-labels are merged with the initial labeled data to form the augmented labeled data, and a semi-supervised learning strategy is adopted to iteratively use the augmented labeled data to further train the model. Each round of training includes the utilization and fine-tuning of pseudo-labels (the specific process is the same as the above process). The core of this strategy is to make the model more robust by gradually introducing more information from the labeled data, while improving its accuracy in identifying various entities and relationships in construction accident reports. In each round of iteration, the model not only learns to identify new entity and relationship types, but also adjusts its own parameters to better adapt to the characteristics and complexity of the dataset. After the iteration ends, the model performance is evaluated on an independent test set to ensure its good generalization performance on unseen data.
[0066] To facilitate the distinction of each training stage, the BERT-BiLSTM-CRF model trained at this time is referred to as the first model in this embodiment. After a piece of text is input into the first model, the model can output the key entities, entity categories, and relationships between entities in the text according to the logical relationships in the accident report.
[0067] Furthermore, considering that the language style, structure, and content of the rules and regulations documents are different from those of accident reports, the model can also be fine-tuned on the basis of the first model using the labeled rules and regulations documents to adapt to the specific language environment of the regulatory documents. After fine-tuning, pseudo-labels are generated again for the unlabeled rules and regulations documents. The pseudo-labels are merged with the initial labeled data, and a semi-supervised learning strategy is adopted to iteratively use the augmented labeled data to further train the model. The trained model is used as the final BERT-BiLSTM-CRF model, which is referred to as the second model in this embodiment. Rules and regulations documents usually contain detailed information on various aspects such as safety operations, risk management, accident prevention measures, and emergency responses. By deeply extracting entities and relationships from these documents, various safety specifications, operating procedures, risk factors, and recommended preventive measures can be identified. This information is crucial for supplementing and enriching the data obtained from accident reports, especially in covering the diversity of accident causes. After a piece of text is input into the second model, the model can comprehensively consider the logical relationships in construction project reports and rules and regulations, and output the key entities, entity categories, and relationships between entities in the text.
[0068] In practical applications, both models can perform entity extraction on any piece of text related to construction safety. Therefore, in S120, either the first model or the second model can be used to extract entities and the relationships between entities in the report to be evaluated, and no specific restrictions are imposed in this embodiment.
[0069] S130. Evaluate the accuracy and analysis depth of the accident report according to the levels and connection paths of two entities with a causal relationship in the accident tree, where the accident tree takes the accident type entity as the root node, and the accident cause entities as the intermediate nodes and leaf nodes.
[0070] An accident tree is a model or data body used to analyze accident causes. In this embodiment, after entity extraction is completed, with reference to a pre-constructed accident tree, the extracted causal logic is compared with the causal logic in the accident tree to evaluate the accuracy and depth of the causal analysis in the report. To illustrate this step, the construction process of the accident tree is introduced first. In a specific embodiment, the accident tree can be constructed using knowledge graph technology, which specifically includes the following steps:
[0071] Step 1. Extract entities and the relationships between entities from existing construction accident reports and regulatory documents. Optionally, either the first model or the second model in S120 can be used to extract entities and entity relationships from existing construction accident reports and store them in the entity set and the entity relationship set. Since there may be a problem in the extraction results that different entity representations actually refer to the same entity, the cosine similarity can be further used to calculate the similarity between entities, the K-means algorithm can be used for entity clustering, and the entity with the highest frequency in each cluster can be selected as the standardized representation of the entity in this class. Optionally, the clustering results and the standardized entities can be submitted to professionals for manual verification to ensure correctness and consistency, and adjustments and optimizations can be made when necessary. Finally, according to the standardized results, the entities in the entity set and the relationship set are updated to ensure that each entity is the standardized entity representation. The cosine similarity calculation formula is as follows:
[0072]
[0073] where is the vector representation of entity e.
[0074] At the same time, the second model mentioned in S120 can be used to extract entities and entity relationships from existing regulatory documents, perform entity standardization by calculating similarity and clustering, and add the standardized entities to the entity set and the relationship set.
[0075] Step 2: According to the categories of the extracted entities, match the relationships between the extracted entities with the set of relationships between the entity categories in the predefined construction accident domain. Optionally, form multiple triples with the entities and entity relationships identified in Step 1. Each triple includes two entities and the relationship between the two entities. For each triple, first determine the categories to which the two entities within the group belong, and then perform a Standard Query Language (SPARQL) query on the above ontology model to obtain the possible association relationships between the categories to which the two entities belong. Then, compare the possible association relationships with the relationship between the two entities stored in the triple. If the semantics are consistent, the relationship is considered a match. Optionally, use the string matching algorithm, the Levenshtein distance algorithm, to calculate the text similarity between the two relationships. When the similarity exceeds a predetermined threshold, it is confirmed as a matching relationship with consistent semantics. Among them, the Levenshtein distance L(a, b) is a function that measures the difference between two strings a and b, and is defined as the minimum number of single-character edits (insertions, deletions, or replacements) required to transform string a into string b. The similarity can be calculated using the following formula:
[0076]
[0077] In the formula, len(a) and len(b) represent the lengths of strings a and b respectively.
[0078] Exemplarily, the categories to which the two entities in a triple belong are "personal status" and "organizational safety culture" respectively. As can be seen from Table 2, the possible association relationships between these two categories include "reflect" and "influence". Then, calculate the text similarity between the relationship between the entities stored in the triple and "reflect" and "influence" respectively. If the text similarity between this relationship and "reflect" exceeds the set threshold, it is considered that this relationship matches "reflect", that is, this triple matches the entry {personal status, organizational safety culture, reflect} in the ontology model.
[0079] Step 3: Construct a knowledge graph for the construction accident domain based on the entities and relationships between entities that can be matched. Optionally, use the graph database management system Neo4j to convert all the triples obtained in Step 2 that can match a certain entry in the ontology model into a graph structure. In this structure, the entities are converted into nodes in the graph, and the relationships between the entities are converted into edges in the graph. Quality assessment, indexing, and optimization can be performed on the knowledge graph to support efficient data retrieval and analysis.
[0080] Step 4: For the triples that fail to match any entry in the ontology model, based on the occurrence frequencies of the entities and relationships in the triples and expert judgment, add the triples with occurrence frequencies greater than the set threshold and the triples considered important by the experts to the above knowledge graph to improve the knowledge system.
[0081] Step 5: Traverse and search for paths from the accident type entity node to the cause entity node in the knowledge graph. Output the fault tree by means of operations such as node and path construction, and repeated node / path merging. Combine probability calculation and industry expert scoring to supplement probability values for intermediate and bottom-level event nodes, and finally generate a fault tree with the accident type entity as the root node, and the accident cause entities as the intermediate nodes and leaf nodes. Specifically, first execute a SPARQL query on the knowledge graph to retrieve all entities and relationships of specific accident types and accident cause entities in the knowledge graph, and form a subgraph associated with the specific accident type and accident cause entities. Since the standard query language of SPARQL mainly targets graph pattern matching, all cause entities directly associated with the accident type can be retrieved first. For each found cause entity, execute the query again to find other entities associated with it until the end node is reached or specific conditions are met. Then, apply the depth-first search graph traversal algorithm in the subgraph to traverse and record all path information from the accident type node to the cause entity node. Again, for each path, construct a fault tree with the accident type entity as the root node; add nodes to the tree in the order of the path, each node representing an entity on the path, and the connection between nodes representing a causal relationship; for multiple paths, based on the Levenshtein string matching algorithm, traverse the nodes of the fault tree to search for duplicate nodes, and merge the corresponding subtrees or subsequent paths of the duplicate nodes onto the existing nodes; finally, use Graphviz to draw and display the fault tree. Finally, add probability information to each node in the fault tree. Optionally, for entity relationships obtained from statistical data in accident reports, calculate the triggering probability between entities, and its calculation formula is as follows:
[0082]
[0083] In the formula, Count(E1→E2) is the number of relationships where entity 1 causes entity 2, and Count(_→E2) is the number of all relationships that cause entity 2; the calculated probability P(E2|E1) represents the proportion caused by entity 1 among the events where entity 2 occurs.
[0084] For entity relationships supplemented by regulatory documents and not mentioned in accident reports, invite domain experts to score their triggering probabilities with reference to the probability data obtained from accident report statistics. Store the probability information in the form of key-value pairs in each tree structure node, where the key is the relationship identifier between entities and the value is the corresponding probability. Thus, the construction of the fault tree is completed.
[0085] It can be seen that the above knowledge graph and fault tree present different entities and relationships in large-scale construction accident reports and regulatory documents through a graph or tree structure, revealing the complex causal logic behind accident safety management. Among them, the logical relationships in the knowledge graph are more comprehensive and complex. The fault tree focuses on the causal logic between accident type entities and accident cause entities. In this embodiment, the fault tree will be used as the knowledge framework to evaluate the quality of the accident report to be evaluated.
[0086] Generally speaking, the evaluation of the integrity, accuracy, and analysis depth of the accident report can be comprehensively obtained from two aspects: whether the entity category is missing and whether the entity and its relationship are missing. Among them, to determine whether the entity and its relationship are missing, it is necessary to compare the causal logic in the report with the causal logic in the fault tree, and analyze the hierarchy and connection path of the entities with causal relationships in the report in the fault tree. S130 corresponds to this process. In a specific embodiment, the report evaluation combined with the fault tree may include the following steps:
[0087] Step 1: Generate a causal chain of the accident report according to every two entities with direct causal relationships in the accident report. As mentioned above, what is obtained in S120 are multiple triples. If the relationship in a triple is a causal relationship, then the two entities in this triple are called two entities with direct causal relationships. Further, if there are two pairs of entities with direct causal relationships in the accident report to be evaluated, and the cause of one is the result of the other, then the two pairs of entities can be concatenated into a causal chain. Exemplarily, if A causes B and B causes C, then a causal chain of A—>B—>C can be obtained, where the entity at the front end is the cause and the entity at the back end is the result.
[0088] Step 2: Determine the connection path of each pair of entities in the causal chain in the fault tree. Specifically, for each causal chain, execute the SPARQL query and the Levenshtein string matching algorithm on the fault tree to locate the corresponding nodes and edges of the causal chain in the tree. For each retrieved path, record the nodes and edges of the path and the relationships between them. Optionally, first locate the node corresponding to the first event in the causal chain in the fault tree. For example, use the string matching algorithm to take the node most similar to this event in the fault tree as the node corresponding to this event; then start from this node and search for the node most similar to the next event in the causal chain in the fault tree through the depth-first search method; when a matching adjacent node is found, construct a mapping between the causal chain and the fault tree to obtain the path of the entire causal chain or part of the causal chain in the fault tree, thereby obtaining the connection path of any pair of entities in the causal chain in the fault tree. During this process, if no node in the fault tree can correspond to the current event, then it will backtrack to the previous node of the fault tree and try other possible matches. Repeat the above process until all causal chains are searched and marked.
[0089] Step 3: Compare the connection relationship of each pair of entities in the causal chain with their link paths in the fault tree from different angles, verify the matching degree between the two, so as to evaluate the accuracy and analysis depth of the report. The present embodiment provides the following three verification methods:
[0090] Method 1: Verify whether there are missing intermediate events in the causal chain to evaluate the analysis integrity and accuracy of the report. Specifically, judge whether there are intermediate nodes not mentioned in the accident report in the connection path of two entities in the fault tree in the causal chain. If so, record it as a potential inconsistent relationship, which indicates that there may be negligence or deviation in the accident analysis of the report. Then, evaluate the accuracy of the accident report according to the network centrality and / or occurrence probability of the intermediate node in the fault tree, where the higher the network centrality and / or occurrence probability, the lower the accuracy.
[0091] Optionally, let the network centrality of the node be Z1 and the occurrence probability of the node be Z2. The calculation method of the network centrality is:
[0092] Node network centrality Z1 = degree centrality + neighbor centrality
[0093] Degree centrality = the number of edges directly connected by this node to other nodes / the number of all edges in the network
[0094] Neighbor centrality = the average short path length of this node / the average short path length of the network. The average short path length = L / (n - 1)
[0095] Wherein, L is the sum of the shortest path lengths from this node to each other node in the network, and n is the number of nodes in the network; the degree centrality represents the influence of this node on other nodes, and the neighbor centrality represents the influence of this node on neighboring nodes.
[0096] Node comprehensive score S_z = 0.5 × Z1 + 0.5 × Z2
[0097] If the score S_z(A) of node A is greater than a set ratio (such as 120%) of the average score of all node comprehensive scores S_z, then it is considered that this node is a node with "high network centrality and / or high accident probability" that may be missing in the accident report. Record the number of missing nodes n_z, and deduct two points for each missing node. Then, the calculation formula for the analysis integrity and accuracy of the accident report is:
[0098] score_z = 20 - 2 × n_z
[0099] Method 2: Verify the levels of the bottom events in the accident tree within the causal chain to evaluate the analysis depth of the report, especially focusing on the exploration of factors such as personal status, unsafe management, and organizational safety culture. Specifically, calculate the shortest path from the top event to the leaf nodes in the accident tree for the first cause entity in the causal chain. The level of this path is called the shortest chain path level. The more levels in the shortest chain path, it indicates that there are multiple deeper causal entities for the bottom events in the causal chain, which the accident report fails to consider, suggesting insufficient analysis depth. Especially when the report ignores deep factors such as personal status, unsafe management, and organizational safety culture, it is also marked as having insufficient analysis depth.
[0100] Optionally, record the number of levels of the shortest chain path as n_d, and a deduction score S_d can be assigned according to the number of levels:
[0101]
[0102] Meanwhile, define weight values for accident cause categories, and subtract corresponding scores for reports that ignore this part:
[0103] Personal status j1, weight1 = 2 points
[0104] Unsafe management j2, weight2 = 6 points
[0105] Organizational safety culture j3, weight3 = 4 points
[0106] Therefore, the formula for calculating the analysis depth of the accident report is:
[0107] score_d = 20 - (j1×weight1 + j2×weight2 + j3×weight3 + S_d)
[0108] Where j n represents the existence status of each cause category. j n = 0 indicates existence, and j n = 1 indicates absence.
[0109] Method 3: Verify the differences between the causal chain path and the accident tree path to evaluate the accuracy of the report. Specifically, for two entities in the causal chain, compare their connection paths in the causal chain with those in the accident tree. If the two paths are significantly different, it indicates that there may be major problems with the report logic, and it is marked as an incorrect potential identification relationship type.
[0110] Optionally, the calculation process of the path difference degree includes: recording the relationship path of the entity pair A-B in the report description as path T1, and among all the paths of the matching entity pair A-B in the fault tree, the path with the most duplicate nodes with path T1 is recorded as T2; calculating the number of common nodes in T1 and T2, denoted as N_T; calculating the sum of the number of different nodes in T1 and T2, denoted as N_BT, then the path difference degree S_t = N_BT / (N_T + N_BT). Correspondingly, the calculation formula for the recognition accuracy of the relationship type of the accident report is:
[0111]
[0112] It should be noted that the above three verification methods can be executed simultaneously, or separately, or partially according to needs. This embodiment does not make specific restrictions. In addition, if extra nodes or relationships that do not belong to the fault tree are detected in the report, experts can further analyze to determine the reasons for the omissions or mismatches in the report, and decide whether to update the fault tree to include new information, or correct the report to reflect known accident patterns.
[0113] S140. Evaluate the comprehensive quality of the accident report according to the accuracy, analysis depth, and the integrity of the entity categories in the accident report.
[0114] In addition to accuracy and analysis depth, report integrity is also an important aspect of quality assessment. In the analysis depth of the accident report in S130, the impacts of accident cause categories such as personal status, unsafe management, and organizational safety culture on report integrity have been described. In addition, the accident report must also include the accident type and causal accidents in the entity category. If missing, it indicates that the report is unqualified; the remaining entity categories, including accident risk, occurrence location, operation link, responsible entity, and accident prevention and rectification measures, also affect the quality of the report to varying degrees.
[0115] In a specific embodiment, the full score of the integrity of the accident report content and structure can be set as scoremax_c = 40 points, and a weight value is defined for each important entity category:
[0116] Accident type i1, weight1 = 40 points
[0117] Accident cause i2, weight2 = 40 points
[0118] Responsible entity i3, weight3 = 12 points
[0119] Accident risk i4, weight4 = 8 points
[0120] Occurrence location i5, weight5 = 8 points
[0121] Operation step i6, weight6 = 6 points
[0122] Accident prevention and rectification measures i7, weight7 = 6 points
[0123] When a required entity category is detected to be missing, subtract the weight value corresponding to this category from the total score. If multiple categories are missing, deductions are made separately for each missing category. If the accident type or cause category is missing, the integrity score is 0, and the report is directly regarded as unqualified without further deductions. To sum up, the calculation formula for the integrity of the accident report content and structure is as follows:
[0124]
[0125] Among them, i n represents the existence status of each entity category, and i n = 0 indicates existence, and i n = 1 indicates missing.
[0126] Finally, integrate the integrity of the entity categories, the accuracy and depth of the accident cause analysis to form a comprehensive evaluation of the accident report quality. Exemplarily, under the above scoring method, the comprehensive score score_f (full score 100 points) of the report quality can be expressed as:
[0127] score_f = score_c + score_z + score_d + score_t
[0128] At the same time, a detailed feedback report can also be generated for the recorded deduction points and deduction bases, pointing out specific inconsistencies and possible optimization directions, guiding the report writer on how to fill in the logical gaps and how to enhance the causal logic coherence in the report.
[0129] Figure 2 is a flowchart of another accident report evaluation method provided by an embodiment of the present invention, including the main steps in the above optional embodiments, and indicating the data flow of each step according to the time line (as shown by the arrows). In addition, in the figure, the overall method is divided into modules such as data extraction, knowledge graph construction, accident tree construction, and accident report quality self-evaluation according to the logic and close relationship between each operation. Among them, each module can be replaced by other equivalent modules or methods, and all belong to the protection scope of this embodiment.
[0130] In summary, this embodiment provides an evaluation method for construction accident reports, which comprehensively utilizes text materials such as accident reports and regulations documents to construct a knowledge graph, covering common sense knowledge and reasoning knowledge from different perspectives and text styles, and providing a complete knowledge framework; then, an accident tree is constructed based on the association relationship between accidents and accident causes in the knowledge graph, reflecting the causal elements and causal logic of most construction accidents, and providing a knowledge reference for the causal logic of construction accidents; finally, the report to be evaluated is subjected to content mapping and structural matching with the accident tree, and through content integrity inspection, intermediate event inspection, bottom event inspection, and path consistency inspection, accident reports with high integrity and logic are screened out, and possible improvement directions are provided in terms of content, analysis logic, etc., effectively improving the quality of construction accident reports and reducing long-term safety risks.
[0131] Figure 3 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present invention, as Figure 3 shown, the device includes a processor 60, a memory 61, an input device 62, and an output device 63; the number of processors 60 in the device may be one or more, Figure 3 here, one processor 60 is taken as an example; the processor 60, the memory 61, the input device 62, and the output device 63 in the device may be connected through a bus or other means, Figure 3 here, connection through a bus is taken as an example.
[0132] The memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the construction accident report evaluation method in the embodiment of the present invention. The processor 60 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 61, that is, implements the above-mentioned construction accident report evaluation method.
[0133] The memory 61 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory 61 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 61 may further include a memory remotely set relative to the processor 60, and these remote memories can be connected to the device through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.
[0134] The input device 62 can be used to receive input digital or character information and generate key signal inputs related to the user settings and function controls of the device. The output device 63 can include display devices such as a display screen.
[0135] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A method for evaluating construction accident reports, characterized in that, Including: Obtain a construction accident report to be evaluated; Extract multiple entities and the relationships between entities in the accident report, where the entity categories include accident types and accident causes, and the relationships between entities include causal relationships; Identify whether there are intermediate nodes in the accident tree for two entities with a causal relationship that are not mentioned in the accident report; if so, evaluate the accuracy of the accident report according to the network centrality and / or occurrence probability of the intermediate node in the accident tree, where the higher the network centrality and / or occurrence probability, the lower the accuracy; Generate a causal chain of the accident report based on the entities with causal relationships in the accident report; evaluate the accuracy of the accident report according to the difference between the paths of the two entities in the causal chain and in the accident tree, where the greater the difference, the lower the accuracy; evaluate the analysis depth of the accident report according to the shortest path from the first cause entity in the causal chain to the leaf node in the accident tree as the top event, where the more levels in the shortest path, the less sufficient the analysis depth; Wherein, the accident tree has an accident type entity as the root node, and accident cause entities as intermediate nodes and leaf nodes.
2. The method according to claim 1, wherein Before extracting multiple entities and the relationships between entities in the accident report, it further includes: Input existing construction accident reports and regulatory documents into the BERT-Bi-LSTM-CRF model for training, so that the trained model can extract multiple entities and the relationships between entities in the input text.
3. The method according to claim 2, wherein The inputting of existing construction accident reports and regulatory documents into the BERT-Bi-LSTM-CRF model for training includes: Annotate the categories of each entity in the existing construction accident reports and regulatory documents according to the set of entity categories in the predefined construction accident field; Annotate the relationships between entities according to the categories of each entity and the set of relationships between entity categories in the predefined construction accident field for model training.
4. The method according to claim 1, wherein Before identifying whether there are intermediate nodes in the accident tree for two entities with a causal relationship that are not mentioned in the accident report, it further includes: Extract multiple entities and the relationships between entities in the existing construction accident reports and regulatory documents; Match the relationships between entities with the set of relationships between entity categories in the predefined construction accident field according to the categories of each entity; Construct a knowledge graph of the construction accident field based on the entities and the relationships between entities that can be matched; Traverse and search for the path from the accident type entity to the accident cause entity in the knowledge graph, and generate an accident tree with the accident type entity as the root node and accident cause entities as intermediate nodes and leaf nodes.
5. The method according to claim 1, characterized in that, It further includes: Evaluate the comprehensive quality of the accident report according to the accuracy and analysis depth, and the integrity of the entity categories in the accident report.
6. The method according to claim 5, characterized in that, The evaluation of the comprehensive quality of the accident report according to the accuracy and analysis depth, and the integrity of the entity categories in the accident report includes at least one of the following: If the accident type entity or the accident cause entity is missing among the multiple entities, evaluate that the accident report is incomplete; If any one of the entities of responsible subject, accident risk, occurrence location, operation link, accident prevention and rectification measures is missing from the multiple entities, it is evaluated that the integrity of the accident report is insufficient; If any one of the entities of personal status, unsafe management and organizational safety culture is missing from the accident cause entities, it is evaluated that the integrity of the accident report is insufficient.
7. An electronic device, characterized in that, Including: One or more processors; A memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the construction accident report evaluation method according to any one of claims 1-6.
Citation Information
Patent Citations
System and method for performing automated system management
US7120559B1