A knowledge graph construction method based on a semantic network
By using a semantic network-based safety knowledge recognizer and analysis method, a safety knowledge graph is constructed, which solves the problems of low efficiency and unclear relationships in the construction of safety production management knowledge graphs in existing technologies. It enables intuitive presentation and quantitative analysis of accidents and their causes, supporting safety decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHIPONT (BEIJING) RES INST OF SAFETY PROD
- Filing Date
- 2026-01-19
- Publication Date
- 2026-04-14
AI Technical Summary
Existing safety production management knowledge graphs are inefficient to build, prone to errors, and unable to clearly present the relationship between accidents and their causes, making it difficult to trace the root cause of accidents and support accurate decision-making.
A semantic network-based approach is adopted to obtain a set of word vectors through a security knowledge recognizer, construct a basic security knowledge graph, analyze the scale coefficients and proportion coefficients of accidents and causes, perform confidence analysis and correction, label edge scales, and form a security knowledge graph.
It enables the efficient construction of a safety knowledge graph, clearly presents the relationship between accidents and their causes, quantifies the severity of accidents and the impact of their causes, supports safety risk identification and prevention planning, and provides high-quality knowledge support.
Smart Images

Figure CN121525816B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of knowledge graph construction technology, and in particular to a knowledge graph construction method based on semantic networks. Background Technology
[0002] In the field of safety production management, traditional safety knowledge management largely relies on manual organization, storing accident records and other information in documents, tables, or simple databases. A few knowledge graph construction solutions also focus solely on entity listing and depend on manual annotation and association. However, these technologies have significant shortcomings. The data is mostly unstructured or semi-structured, requiring extensive manual processing, which is inefficient and prone to errors. Furthermore, they cannot clearly present the relationship between accidents and their causes, making it difficult to trace the root causes of accidents and support accurate decision-making. In addition, manual annotation and association are not only time-consuming but also susceptible to bias due to experience, and data errors are difficult to correct, thus limiting their practicality. Summary of the Invention
[0003] This invention addresses the technical problems in existing safety production management knowledge graph construction, such as low efficiency and susceptibility to errors, inability to clearly present the relationship between accidents and causes, difficulty in tracing the root cause of accidents, and difficulty in supporting accurate decision-making. It provides a knowledge graph construction method based on semantic networks to solve these problems.
[0004] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:
[0005] In a first aspect, the present invention provides a method for constructing a knowledge graph based on a semantic network, comprising: acquiring a production dataset of the knowledge graph to be constructed; using a security knowledge recognizer to perform security knowledge recognition and obtain a set of security knowledge word vectors, wherein the set of security knowledge word vectors includes a set of accident type word vectors and a set of cause type word vectors; constructing a basic security knowledge graph based on the semantic network and the set of security knowledge word vectors, wherein the accident type word vector set and the set of cause type word vectors are used as entities and the accident causes are used as edges; analyzing the accident scale coefficient of each accident type word vector to obtain a set of scale coefficients, analyzing the proportion coefficient of each cause type word vector to obtain a set of proportion coefficients; analyzing and obtaining the accident coefficient set and the cause coefficient set based on the training features of the security knowledge recognizer, correcting the set of proportion coefficients, processing to obtain the accident annotation coefficient set and the cause annotation coefficient set, and annotating the edge scale of each accident type word vector and the cause type word vector to obtain a security knowledge graph, wherein confidence analysis is performed and the set of proportion coefficients is corrected.
[0006] Optionally, the production dataset to be constructed for the knowledge graph is obtained, and a safety knowledge recognizer is used to identify safety knowledge and obtain a set of safety knowledge word vectors. This includes: pre-training the safety knowledge recognizer; obtaining record data of accidents that occurred during the production process within a preset time range in the past, as the production dataset to be constructed for the knowledge graph; inputting the production dataset into the safety knowledge recognizer, and recognizing and outputting an accident type word vector set and a cause type word vector set, as the set of safety knowledge word vectors.
[0007] The pre-training of the safety knowledge recognizer includes: collecting a set of sample production data based on safety knowledge records over a historical period; extracting the accident type word vector and cause type word vector corresponding to each sample production data to obtain a set of sample accident type word vectors and a set of sample cause type word vectors; constructing a safety knowledge recognizer based on a natural language processing model; and using the set of sample production data, the set of sample accident type word vectors, and the set of sample cause type word vectors as training data to conduct supervised training and testing of the safety knowledge recognizer. Pre-training is completed when the test meets the recognition performance requirements.
[0008] Optionally, based on a semantic network, a basic security knowledge graph is constructed according to the security knowledge word vector set, including: forming multiple sets of individual knowledge by using the accident type word vector set and the cause type word vector set as entities and the accident cause as edges; and constructing the basic security knowledge graph based on the multiple sets of individual knowledge.
[0009] Optionally, the accident scale coefficient of each accident type word vector is analyzed to obtain a set of scale coefficients, and the proportion coefficient of each cause type word vector is analyzed to obtain a set of proportion coefficients. This includes: inputting each accident type word vector into an accident scale database, outputting the accident scale coefficient of the corresponding accident type, and obtaining a set of scale coefficients, wherein the accident scale database is constructed using multiple sample accident type word vectors and multiple labeled sample accident scale coefficients; and calculating the proportion of each cause type word vector in the set of cause type word vectors to obtain a set of proportion coefficients.
[0010] Optionally, based on the training features of the safety knowledge recognizer, an accident coefficient set and a cause coefficient set are obtained by analysis, including: acquiring the training data of the safety knowledge recognizer; extracting data including the accident type word vector set from the training data, and calculating the proportion of data corresponding to each accident type word vector to obtain the accident coefficient set; extracting data including the cause type word vector set from the training data, and calculating the proportion of data corresponding to each cause type word vector to obtain the cause coefficient set.
[0011] Optionally, the set of proportion coefficients is modified to obtain a set of accident labeling coefficients and a set of cause labeling coefficients. The side-scales of each accident type word vector and cause type word vector are labeled to obtain a safety knowledge graph. This includes: calculating the similarity between the set of proportion coefficients and the set of cause coefficients to obtain a set of cause confidence; using the set of cause confidence, modifying the set of proportion coefficients to obtain a set of cause labeling coefficients; calculating the set of accident labeling coefficients based on the set of scale coefficients and the set of accident coefficients; and using the set of cause labeling coefficients and the set of cause labeling coefficients to label the side-scales of each accident type word vector and cause type word vector to obtain a safety knowledge graph.
[0012] Optionally, the cause annotation coefficient set and cause annotation coefficient set are used to annotate the edge scales of each accident type word vector and cause type word vector to obtain a safety knowledge graph. This includes: calculating the annotation coefficient of each group of individual knowledge based on the cause annotation coefficient set and cause annotation coefficient set to obtain an annotation coefficient set; adjusting the preset edge scale based on the ratio of each annotation coefficient to the mean of the annotation coefficient set to obtain an edge scale set; and performing edge scale annotation within the basic safety knowledge graph to obtain a safety knowledge graph.
[0013] By implementing this invention, it is possible to obtain the production dataset of the knowledge graph to be constructed, use a security knowledge recognizer to perform security knowledge recognition, and obtain a set of security knowledge word vectors. The set of security knowledge word vectors includes a set of accident type word vectors and a set of cause type word vectors. Using accident records in the production process as the data source, it is directly related to the core requirements of the security knowledge graph, which can avoid interference from irrelevant data and improve the efficiency of subsequent construction.
[0014] By implementing this invention, a basic safety knowledge graph can be constructed based on a semantic network and the set of safety knowledge word vectors. The graph is constructed with accident type word vector sets and cause type word vector sets as entities and accident causes as edges. Through the classic "entity-edge" graph structure, the relationship between accidents and causes is presented intuitively, which conforms to the core logic of safety knowledge analysis, namely, clarifying what accidents are caused by what causes, which facilitates the subsequent interpretation of safety relationships.
[0015] By implementing this invention, it is possible to analyze the accident scale coefficient of each accident type word vector to obtain a set of scale coefficients, and analyze the proportion coefficient of each cause type word vector to obtain a set of proportion coefficients. The accident scale coefficient quantifies the fuzzy concept of accident size, and can intuitively distinguish the severity of different accident types, providing data support for safety priority judgment. Meanwhile, the cause proportion coefficient clearly presents the frequency of accidents caused by various causes, helping to identify high-frequency causes and providing a basis for key directions of safety prevention and control.
[0016] By implementing this invention, it is possible to analyze and obtain a set of accident coefficients and a set of cause coefficients based on the training features of the safety knowledge recognizer, correct the set of proportion coefficients, process and obtain a set of accident annotation coefficients and a set of cause annotation coefficients, and annotate the edge scales of the word vectors of each accident type and cause type to obtain a safety knowledge graph. By introducing the training features of the recognizer to correct the original coefficients, deviations in the data collection or recognition process can be offset, and the accuracy of each coefficient can be improved. The edge scale annotation presents the correlation strength between accidents and causes in a visual way, allowing for intuitive judgment of core safety correlations without complex calculations.
[0017] In summary, by implementing this invention, the visualization of the safety knowledge graph can be enhanced, clearly presenting the logical relationship between "accident type - cause type". Furthermore, by using the scale coefficient, cause proportion coefficient, and corrected annotation coefficient, the severity of the accident, the weight of the cause's influence, and the strength of the relationship between the two can be quantified, providing high-quality and highly available knowledge support for safety risk identification, accident cause tracing, key prevention and control planning, and subsequent optimization of safety analysis models. Attached Figure Description
[0018] Figure 1 A flowchart illustrating a knowledge graph construction method based on semantic networks provided by this invention;
[0019] Figure 2 This invention provides a schematic diagram of the process for constructing a security knowledge graph based on semantic networks. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0022] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.
[0023] Example 1, as Figure 1 As shown, this embodiment of the invention provides a method for constructing a knowledge graph based on a semantic network, including:
[0024] S100 acquires the production dataset of the knowledge graph to be constructed, uses a security knowledge recognizer to identify security knowledge, and obtains a set of security knowledge word vectors, wherein the set of security knowledge word vectors includes a set of accident type word vectors and a set of cause type word vectors;
[0025] S200 is based on a semantic network and constructs a basic security knowledge graph based on the security knowledge word vector set. The graph is constructed with accident type word vector set and cause type word vector set as entities and accident causes as edges.
[0026] S300 analyzes the accident scale coefficient of each accident type word vector to obtain a set of scale coefficients, and analyzes the proportion coefficient of each cause type word vector to obtain a set of proportion coefficients.
[0027] S400 analyzes and obtains an accident coefficient set and a cause coefficient set based on the training features of the safety knowledge recognizer, corrects the proportion coefficient set, processes and obtains an accident labeling coefficient set and a cause labeling coefficient set, and labels the side scales of each accident type word vector and cause type word vector to obtain a safety knowledge graph, wherein confidence analysis is performed and the proportion coefficient set is corrected.
[0028] In step S100 of this application embodiment, a production dataset of the knowledge graph to be constructed is obtained, and a security knowledge recognizer is used to perform security knowledge recognition to obtain a set of security knowledge word vectors, including:
[0029] Pre-train the security knowledge recognizer;
[0030] Obtain recorded data of accidents that occurred during the production process within a preset time range in the past, and use this as the production dataset for building the knowledge graph;
[0031] The production dataset is input into the safety knowledge recognizer, and the recognition output obtains a set of accident type word vectors and a set of cause type word vectors, which are used as the safety knowledge word vector set.
[0032] In this embodiment of the application, the purpose of step S100 is to provide standardized, high-quality core data for the subsequent construction of the safety knowledge graph, namely, information related to accident types and cause types. Through the process of "data acquisition - knowledge recognition - structured output", unstructured production accident records are transformed into entity data that can be directly used in the graph.
[0033] In step S100 of this application embodiment, the security knowledge recognizer is pre-trained, including:
[0034] Based on safety knowledge records from historical periods, a set of sample production data is collected. The accident type word vector and cause type word vector corresponding to each sample production data are extracted to obtain a set of sample accident type word vectors and a set of sample cause type word vectors.
[0035] A security knowledge recognizer is built based on a natural language processing model;
[0036] The safety knowledge recognizer is trained and tested in a supervised manner using the sample production data set, the sample accident type word vector set, and the sample cause type word vector set as training data. The pre-training is completed when the recognition performance requirements are met during testing.
[0037] In this embodiment, a security knowledge recognizer is pre-trained. The core objective is to create an automated knowledge extraction tool with accurate recognition capabilities, laying the foundation for efficient and accurate extraction of security knowledge from production data in the future.
[0038] First, sample data needs to be prepared to train the security knowledge recognizer.
[0039] Specifically, the data source should be "safety knowledge records within a historical period". A set of sample production data related to safety should be collected during the production process. The data should cover various accident scenarios such as equipment failure accidents and operational error accidents to ensure the diversity and representativeness of the sample production data set, and to support the generalization ability of the safety knowledge recognizer.
[0040] Then, for each piece of sample production data collected, the corresponding "accident type word vector" and "cause type word vector" are extracted manually or using basic tools. For example, from the sample data "In Month X of 2023, a fire accident occurred on the production line due to equipment overload," "fire accident" is extracted as the accident type word vector, and "equipment overload" is extracted as the cause type word vector. This ultimately forms a set of sample accident type word vectors and a set of sample cause type word vectors, providing corresponding "input-output" labels for the subsequent supervised training of the safety knowledge recognizer.
[0041] Considering the task types of security knowledge recognizers, security knowledge recognizers can be built based on the fine-tuning method of the BERT (Bidirectional Encoder Representations from Transformers) pre-trained model.
[0042] Specifically, the BERT-Base model needs to be selected as the basic architecture of the security knowledge recognizer. The model contains 12 Transformer encoders, and each Transformer contains 12 self-attention heads.
[0043] The hidden layer dimension of the safety knowledge recognizer is set to 768, and the intermediate layer dimension of the Feed-Forward network is 3072. The vocabulary size adopts the original BERT vocabulary of 30522 words, and is supplemented with safety-specific terms such as "equipment overload", "operational violation", and "fire accident", expanding the vocabulary size to 30800 words. The maximum length of the input sequence is set to 128 to adapt to the common length of accident record text in production data and avoid the decrease in training efficiency due to excessively long texts. The output layer adopts a dual-output branch structure, corresponding to the tasks of "accident type word vector extraction" and "cause type word vector extraction" respectively. The output dimension of each output branch is consistent with the word vector dimension, set to 768.
[0044] For the parameter settings of the security knowledge recognizer, the optimizer used is AdamW, with an initial learning rate of 0.00002. The learning rate decay strategy adopts linear decay, completing the warm-up within 10% of the total training rounds, and then linearly reducing the learning rate with the number of training rounds. The batch size is set to 32. For regularization parameters, the Dropout probability is set to 0.1, and the L2 regularization weight is set to 0.0001. The loss function is the cross-entropy loss function. For dual-output branches, the losses of the two branches are added together with a 1:1 weight as the total loss function to optimize the model parameters.
[0045] The training samples for the safety knowledge recognizer are the aforementioned sample production data set, along with the corresponding "sample accident type word vector set" and "sample cause type word vector set". Specifically, 15,000 sample production data entries need to be collected. For each sample production data entry, the corresponding accident type word vector and cause type word vector are extracted, ultimately forming a "sample accident type word vector set" and "sample cause type word vector set" containing 15,000 samples. The above samples are divided into a training set, a validation set, and a test set in an 8:1:1 ratio for training the safety knowledge recognizer.
[0046] For training the security knowledge recognizer, the total number of training rounds is set to 50. During training, after each round, the performance of the security knowledge recognizer is evaluated using a validation set. If the performance of the security knowledge recognizer does not improve on the validation set for five consecutive rounds, an early stopping mechanism is triggered to terminate training prematurely and prevent overfitting of the security knowledge recognizer on the training set. If the early stopping mechanism is not triggered, parameter updates are stopped after 50 rounds of training. When the security knowledge recognizer achieves an accuracy rate of ≥92% for "accident type word vector extraction", ≥90% for "cause type word vector extraction", and ≥88% for "accident type word vector extraction" and ≥86% for "cause type word vector extraction", the security knowledge recognizer is deemed to meet the recognition performance requirements, training converges, and pre-training is completed.
[0047] Furthermore, it is necessary to obtain the recorded data of accidents that occurred during the production process within a preset time range in the past, as the production dataset for the knowledge graph to be constructed.
[0048] First, it is necessary to determine the "preset time range", such as the past 3 years or 5 years. The specific time range should be set according to the time dimension requirements of the map construction to ensure that the data not only covers enough accident scenarios, but also reflects the actual situation of recent production safety.
[0049] Then, all accident-related records during the production process within the preset time frame are collected, including but not limited to accident reports, equipment failure records, safety inspection logs, and accident handling ledgers. The data must contain clear information about the "accident occurrence" and "accident cause analysis" to ensure that the subsequent identifier can extract valid information. The collected records are then preliminarily screened to remove duplicates and incomplete data, forming the final production dataset for constructing the knowledge graph.
[0050] Finally, the production dataset is input into the safety knowledge recognizer, and the recognition output obtains a set of accident type word vectors and a set of cause type word vectors, which are used as the safety knowledge word vector set.
[0051] In step S200 of this application embodiment, a basic security knowledge graph is constructed based on a semantic network and the security knowledge word vector set, including:
[0052] Based on the accident type word vector set and the cause type word vector set as entities, and using accident causes as edges, multiple sets of individual knowledge are formed.
[0053] Based on multiple sets of individual knowledge, a basic security knowledge graph is constructed.
[0054] In this embodiment of the application, the core objective of step S200 is to transform the standardized safety knowledge extracted in S100, which includes accident type vectors and cause type word vectors, into a structured knowledge framework with clear logical connections, providing a basic carrier for the subsequent quantitative optimization of the graph.
[0055] Therefore, it is necessary to form multiple sets of entity knowledge by using the accident type word vector set and the cause type word vector set as entities and the accident cause as edges.
[0056] First, it is necessary to clarify the definitions of "entity" and "edge". Each word vector in the "accident type word vector set" and each word vector in the "cause type word vector set" output by S100 are respectively regarded as independent "entities" in the knowledge graph; the logical relationship "an accident is caused by a certain cause" is defined as the "edge" connecting the two types of entities, and the attribute of the "edge" is labeled as "accident cause".
[0057] Then, based on the actual relationships recorded in the production data, each accident type word vector is matched with a corresponding cause type word vector to construct a combination of "single accident entity - accident cause edge - single cause entity" to form a single set of "monotype knowledge". For example, if the production data records "a fire accident was caused by equipment overload", then the word vector of "fire accident" is used as the accident entity and the word vector of "equipment overload" is used as the cause entity. The two are connected by the "accident cause" edge to form a set of monotype knowledge: "fire accident word vector - accident cause edge - equipment overload word vector". Following this logic, all actual relationships between accidents and causes are traversed to form multiple sets of mutually independent monotype knowledge.
[0058] Furthermore, it is necessary to construct a basic security knowledge graph based on multiple sets of individual knowledge.
[0059] Specifically, it is necessary to sort out the relational logic of individual knowledge, that is, to classify all the formed individual knowledge. If multiple accident type word vectors correspond to the same cause type word vector, such as "fire accident" and "circuit damage accident" both being caused by "equipment overload", then the "equipment overload" cause entities in these two sets of individual knowledge are merged to form a relational structure of "multiple accident entities - accident cause edge - single cause entity". If a single accident type word vector corresponds to multiple cause type word vectors, such as "fire accident" being caused by "equipment overload" and "operation violation", then the "fire accident" entities in the corresponding individual knowledge are merged to form a relational structure of "single accident entity - accident cause edge - multiple cause entity".
[0060] Then, all the individual pieces of knowledge, after being sorted out, are integrated into a unified network structure according to the above-mentioned association logic, forming the basic safety knowledge graph. In this knowledge graph, all accident type word vectors and cause type word vectors serve as nodes, and "accident cause" serves as the edge connecting the nodes, intuitively presenting the network relationship between various accidents and causes, thus completing the construction of the basic safety knowledge graph.
[0061] In step S300 of this application embodiment, the accident scale coefficient of each accident type word vector is analyzed to obtain a scale coefficient set, and the proportion coefficient of each cause type word vector is analyzed to obtain a proportion coefficient set, including:
[0062] Each accident type word vector is input into the accident scale database, and the corresponding accident scale coefficient is output to obtain a set of scale coefficients. The accident scale database is constructed using multiple sample accident type word vectors and labeled multiple sample accident scale coefficients.
[0063] Calculate the proportion of each cause type word vector within the set of cause type word vectors to obtain a set of proportion coefficients.
[0064] In this embodiment of the application, the core objective of step S300 is to supplement the basic security knowledge graph with quantitative dimension information, transforming vague concepts such as "the size of the accident" into quantifiable coefficients, and providing data support for subsequent optimization of graph association strength and improvement of the practicality of security analysis.
[0065] First, the word vectors for each accident type need to be input into the accident scale database, and the accident scale coefficients for the corresponding accident types need to be output to obtain a set of scale coefficients.
[0066] The accident scale database needs to be constructed in advance. Its construction logic is to use word vectors of various sample accident types and labeled accident scale coefficients of multiple samples. That is, to collect various accident types from historical safety knowledge records, such as fire accidents, mechanical injury accidents, electric shock accidents, etc., and label the corresponding sample accident scale coefficients for each accident type. The accident scale coefficient setting needs to take into account the losses and impact range caused by the accident. For example, the labeling coefficient for major accidents is 0.8-1.0, for relatively large accidents it is 0.5-0.7, for general accidents it is 0.2-0.4, and for minor accidents it is 0-0.1, etc., and finally form an accident scale database covering multiple accident scenarios.
[0067] Then, each word vector in the "accident type word vector set" obtained in S100 is input into the accident scale database one by one. Through the matching logic built into the database, the accident scale coefficient corresponding to each accident type word vector is output. All the output coefficients are integrated to form a "scale coefficient set". Each scale coefficient in this set corresponds one-to-one with the accident type word vector, which intuitively reflects the severity of various accidents.
[0068] Furthermore, it is necessary to calculate the proportion of each cause type word vector within the set of cause type word vectors to obtain a set of proportion coefficients.
[0069] First, it is necessary to clarify the calculation base of the aforementioned proportion. Specifically, the "cause type word vector set" obtained in S100 should be used as the calculation object. This cause type word vector set contains all cause type word vectors related to the accident extracted from the production data. The total number of cause type word vectors is the sum of the number of all cause type word vectors.
[0070] For each cause-type word vector in the cause-type word vector set, such as the word vector for "equipment overload", count the number of times it appears in the entire cause-type word vector set, that is, the number of times the accident is recorded due to this cause; divide the "number of times a single cause appears" by the "total number of the set" to obtain the proportion coefficient of the cause-type word vector. For example, if "equipment overload" appears 300 times and the total number of cause-type word vectors in the set is 1000 times, then its proportion coefficient is 0.3.
[0071] The above percentage calculation is performed on each of the word vectors for all cause types. The percentage coefficients corresponding to each cause are compiled and summarized to form a "percentage coefficient set". Each percentage coefficient in this set directly reflects the frequency of the accident caused by the corresponding cause.
[0072] In step S400 of this application embodiment, based on the training characteristics of the security knowledge recognizer, the accident coefficient set and the cause coefficient set are analyzed and obtained, including:
[0073] Obtain the training data of the security knowledge recognizer;
[0074] Extract data including the accident type word vector set from the training data, and calculate the proportion of data corresponding to each accident type word vector to obtain the accident coefficient set;
[0075] Extract data including the set of word vectors for the cause type from the training data, and calculate the proportion of data corresponding to each word vector for the cause type to obtain the set of cause coefficients.
[0076] In this embodiment of the application, the objective of the above steps in step S400 is to extract quantitative features related to "accident type" and "cause type" from the training data of the safety knowledge recognizer, namely accident coefficient and cause coefficient, to provide a basis for the subsequent correction of the scale coefficient and proportion coefficient obtained in step S300, and ultimately improve the accuracy and reliability of the safety knowledge graph.
[0077] First, it is necessary to obtain the training data of the security knowledge recognizer. The "training data" here refers to the sample data used when the security knowledge recognizer is pre-trained. Specifically, it is a set of sample production data collected based on security knowledge records within a historical period, as well as the corresponding set of sample accident type word vectors and the set of sample cause type word vectors, which are consistent with the data source for training the security knowledge recognizer in S100.
[0078] Then, it is necessary to extract data including the accident type word vector set from the training data, and calculate the proportion of data corresponding to each accident type word vector to obtain the accident coefficient set.
[0079] The accident type word vector set refers to the set of accident type word vectors extracted from production data in S100 for constructing the knowledge graph, such as word vectors for "fire accident" and "equipment failure accident". All sample data containing these accident type word vectors need to be selected from the training data. For the selected training data, the frequency of occurrence of each accident type word vector used in the knowledge graph is counted. For example, if there are 500 samples in the training data containing the word vector for "fire accident", the proportion of the frequency of each accident type word vector to the total number of samples in the selected training data is calculated; this proportion is the accident coefficient for the corresponding accident type. All accident coefficients for all accident types are integrated to form an "accident coefficient set", where each accident coefficient corresponds one-to-one with an accident type word vector in the knowledge graph.
[0080] Then, it is necessary to extract data including the set of cause type word vectors from the training data, and calculate the proportion of data corresponding to each cause type word vector to obtain the set of cause coefficients.
[0081] The "cause type word vector set" refers to the set of cause type word vectors extracted from production data in S100 for constructing the knowledge graph, such as word vectors for "equipment overload" and "operational violation." All sample data containing these cause type word vectors need to be selected from the training data. For the selected training data, the frequency of each cause type word vector used in the knowledge graph is counted. For example, there are 300 samples in the training data containing the word vector for "equipment overload." The proportion of the frequency of each cause type word vector to the total number of samples in the selected training data is calculated; this proportion is the "cause coefficient" for the corresponding cause type. All cause coefficients for all cause types are integrated to form a "cause coefficient set," where each cause coefficient corresponds one-to-one with a cause type word vector in the knowledge graph.
[0082] like Figure 2 As shown, in step S400 of this embodiment, the process of constructing and obtaining a safety knowledge graph specifically involves: correcting the set of proportion coefficients, processing to obtain a set of accident labeling coefficients and a set of cause labeling coefficients, labeling the edge scales of each accident type word vector and cause type word vector, and obtaining a safety knowledge graph. The process of performing confidence analysis and correcting the set of proportion coefficients includes:
[0083] Calculate the similarity between the set of proportion coefficients and the set of cause coefficients to obtain the set of cause confidence scores;
[0084] The set of cause confidence scores is used to correct the set of proportion coefficients to obtain a set of cause labeling coefficients;
[0085] Based on the aforementioned set of scale coefficients and set of accident coefficients, the set of accident labeling coefficients is calculated.
[0086] Using the aforementioned set of cause annotation coefficients, the edge scales of each accident type word vector and cause type word vector are annotated to obtain a safety knowledge graph.
[0087] In step S400 of this embodiment, the cause annotation coefficient set and the cause annotation coefficient set are used to annotate the side scales of each accident type word vector and cause type word vector to obtain a safety knowledge graph, including:
[0088] Based on the set of cause annotation coefficients and the set of cause annotation coefficients, calculate the annotation coefficients for each group of individual knowledge to obtain the set of annotation coefficients;
[0089] Based on the ratio of each annotation coefficient to the mean of the annotation coefficient set, the preset edge scale is adjusted and calculated to obtain the edge scale set. Edge scale annotation is then performed within the basic security knowledge graph to obtain the security knowledge graph.
[0090] In this embodiment of the application, the goal of the above steps in step S400 is to correct the proportion coefficients and scale coefficients obtained in the early stage by introducing a security knowledge recognizer to train the accident coefficients and cause coefficients related to the features, and to convert the corrected coefficients into scale labels of the "edges" in the graph, so as to finally construct a security knowledge graph with accurate quantitative correlation.
[0091] First, it is necessary to calculate the similarity between the set of proportion coefficients and the set of cause coefficients to obtain the set of cause confidence.
[0092] Specifically, similarity calculation methods such as cosine similarity and Pearson correlation coefficient can be used to compare the "proportion coefficient set" obtained in S300 with the "cause coefficient set" obtained in the previous steps of S400 to obtain the similarity value corresponding to each cause. This value is the "cause confidence" - the closer the confidence is to 1, the more consistent the distribution of the cause is between the production data and the training data, and the higher the confidence of the proportion coefficient; conversely, the lower the confidence is.
[0093] Next, the proportion coefficient is adjusted using the cause confidence level as the weight. Optionally, if the confidence level is high (e.g., ≥0.8), the original proportion coefficient is directly retained; if the confidence level is low (e.g., <0.8), the proportion coefficient is adjusted in conjunction with the cause coefficient. For example, the adjusted coefficient is calculated using the formula "proportion coefficient × confidence level + cause coefficient × (1 - confidence level)". All the adjusted coefficients are then integrated to form a "cause labeling coefficient set".
[0094] Furthermore, based on the aforementioned set of scale coefficients and set of accident coefficients, it is necessary to calculate and obtain the set of accident labeling coefficients.
[0095] Optionally, the "scale coefficient set" reflecting the severity of the accident in S300 can be used as a basis, and the "accident coefficient set" reflecting the distribution characteristics of accidents in the training data calculated in the aforementioned steps of S400 can be introduced for weighted calculation. For example, the formula "accident labeling coefficient = scale coefficient × 0.7 + accident coefficient × 0.3" can be used for calculation. The specific weights in the formula can be adjusted according to the actual scenario, and the adjustment is based on balancing the severity of the accident itself and the reliability of the data. The scale coefficient of each accident type is corrected according to the above method, and finally the "accident labeling coefficient set" is formed. This accident labeling coefficient set reflects the actual impact of the accident and avoids the bias that may exist in single production data.
[0096] Furthermore, based on the aforementioned set of cause annotation coefficients, it is necessary to calculate the annotation coefficients for each group of individual knowledge to obtain the set of annotation coefficients.
[0097] Here, we need to revisit the definition of "individual knowledge" in S200. In the aforementioned steps, each group of individual knowledge is a combination of "accident type word vector - accident cause edge - cause type word vector". For each group of individual knowledge, the corresponding "accident annotation coefficient" and "cause annotation coefficient" are extracted. Then, the "annotation coefficient" of that group of individual knowledge is calculated by multiplication or weighted summation. For example, the annotation coefficient = accident annotation coefficient × cause annotation coefficient. This annotation coefficient directly reflects the comprehensive importance of the "accident-cause" relationship of that group of individual knowledge. The annotation coefficients of all groups of individual knowledge are integrated to form a "annotation coefficient set".
[0098] Furthermore, it is necessary to adjust and calculate the preset edge scale based on the ratio of each annotation coefficient to the mean of the annotation coefficient set, obtain the edge scale set, and perform edge scale annotation within the basic security knowledge graph to obtain the security knowledge graph.
[0099] First, a "preset edge scale" needs to be set, which is the initial uniform scale of the "edges" in the basic knowledge graph, such as a line with a width of 1px. Then, the mean of the annotation coefficient set is calculated, which is the average of the annotation coefficients of all individual knowledge items. Next, for each group of individual knowledge items, its "annotation coefficient" is divided by the "mean of the annotation coefficient set" to obtain the "scale adjustment coefficient," i.e., scale adjustment coefficient = annotation coefficient / mean of the annotation coefficient set. If the adjustment coefficient > 1, it indicates that the importance of this group's association is higher than the average level, and the edge scale needs to be increased, i.e., the lines need to be widened; if < 1, the edge scale needs to be decreased, i.e., the lines need to be narrowed.
[0100] Finally, the preset edge scale is calculated based on the scale adjustment coefficient. For example, the actual edge width = preset edge width × adjustment coefficient, resulting in the "edge scale" corresponding to each group of individual knowledge. This scale is then labeled on the corresponding "edge" in the basic security knowledge graph. After completing the scale labeling of all edges, the security knowledge graph with accurate quantitative relationships is officially constructed.
[0101] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0102] Those skilled in the art will understand that embodiments of the present invention can provide methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0103] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0104] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0105] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0106] Although preferred embodiments of the invention have been described, those skilled in the art, once they have learned the basic inventive concept, can make other changes and modifications to these embodiments.
[0107] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for constructing a knowledge graph based on semantic networks, characterized in that, The method includes: Obtain the production dataset of the knowledge graph to be constructed, use a security knowledge recognizer to identify security knowledge, and obtain a set of security knowledge word vectors, wherein the set of security knowledge word vectors includes a set of accident type word vectors and a set of cause type word vectors; Based on semantic networks, a basic security knowledge graph is constructed according to the security knowledge word vector set, wherein the accident type word vector set and the cause type word vector set are the entities, and the accident cause is the edge. Analyze the accident scale coefficient of each accident type word vector to obtain a set of scale coefficients, and analyze the proportion coefficient of each cause type word vector to obtain a set of proportion coefficients; Based on the training features of the safety knowledge recognizer, an accident coefficient set and a cause coefficient set are obtained through analysis. The proportion coefficient set is then corrected, and an accident labeling coefficient set and a cause labeling coefficient set are obtained. The side scales of the word vectors for each accident type and the word vectors for each cause type are labeled to obtain a safety knowledge graph. Confidence analysis is performed, and the proportion coefficient set is then corrected. The process involves obtaining a set of accident annotation coefficients and a set of cause annotation coefficients, and then annotating the edge scales of the word vectors for each accident type and cause type to obtain a safety knowledge graph, including: Calculate the similarity between the set of proportion coefficients and the set of cause coefficients to obtain the set of cause confidence scores; The set of cause confidence scores is used to correct the set of proportion coefficients to obtain a set of cause labeling coefficients; Based on the aforementioned set of scale coefficients and set of accident coefficients, the set of accident labeling coefficients is calculated. Using the aforementioned set of cause annotation coefficients, the edge scales of each accident type word vector and cause type word vector are annotated to obtain a safety knowledge graph; Specifically, the cause annotation coefficient set and the cause annotation coefficient set are used to annotate the side scales of each accident type word vector and cause type word vector to obtain a safety knowledge graph, including: Based on the set of cause annotation coefficients and the set of cause annotation coefficients, calculate the annotation coefficients for each group of individual knowledge to obtain the set of annotation coefficients; Based on the ratio of each annotation coefficient to the mean of the annotation coefficient set, the preset edge scale is adjusted and calculated to obtain the edge scale set. Edge scale annotation is then performed within the basic security knowledge graph to obtain the security knowledge graph.
2. The knowledge graph construction method based on semantic networks according to claim 1, characterized in that, Obtain the production dataset for the knowledge graph to be constructed, use a security knowledge recognizer to identify security knowledge, and obtain a set of security knowledge word vectors, including: Pre-train the security knowledge recognizer; Obtain recorded data of accidents that occurred during the production process within a preset time range in the past, and use this as the production dataset for building the knowledge graph; The production dataset is input into the safety knowledge recognizer, and the recognition output obtains a set of accident type word vectors and a set of cause type word vectors, which are used as the safety knowledge word vector set.
3. The knowledge graph construction method based on semantic networks according to claim 2, characterized in that, Pre-trained security knowledge recognizers include: Based on safety knowledge records from historical periods, a set of sample production data is collected. The accident type word vector and cause type word vector corresponding to each sample production data are extracted to obtain a set of sample accident type word vectors and a set of sample cause type word vectors. A security knowledge recognizer is built based on a natural language processing model; The safety knowledge recognizer is trained and tested in a supervised manner using the sample production data set, the sample accident type word vector set, and the sample cause type word vector set as training data. The pre-training is completed when the recognition performance requirements are met during testing.
4. The knowledge graph construction method based on semantic networks according to claim 1, characterized in that, Based on semantic networks and the aforementioned set of security knowledge word vectors, a basic security knowledge graph is constructed, including: Based on the accident type word vector set and the cause type word vector set as entities, and using accident causes as edges, multiple sets of individual knowledge are formed. Based on multiple sets of individual knowledge, a basic security knowledge graph is constructed.
5. The knowledge graph construction method based on semantic networks according to claim 1, characterized in that, Analyze the accident scale coefficient of each accident type word vector to obtain a set of scale coefficients, and analyze the proportion coefficient of each cause type word vector to obtain a set of proportion coefficients, including: Each accident type word vector is input into the accident scale database, and the corresponding accident scale coefficient is output to obtain a set of scale coefficients. The accident scale database is constructed using multiple sample accident type word vectors and labeled multiple sample accident scale coefficients. Calculate the proportion of each cause type word vector within the set of cause type word vectors to obtain a set of proportion coefficients.
6. The knowledge graph construction method based on semantic networks according to claim 1, characterized in that, Based on the training characteristics of the safety knowledge recognizer, the accident coefficient set and cause coefficient set are obtained through analysis, including: Obtain the training data of the security knowledge recognizer; Extract data including the accident type word vector set from the training data, and calculate the proportion of data corresponding to each accident type word vector to obtain the accident coefficient set; Extract data including the set of word vectors for the cause type from the training data, and calculate the proportion of data corresponding to each word vector for the cause type to obtain the set of cause coefficients.
Citation Information
Patent Citations
Accident reason analysis method and system based on hidden danger knowledge graph
CN115374174A
Knowledge graph construction and intelligent question and answer method and device based on deep learning
CN119691135A