Mine ventilation knowledge graph construction method based on large model
By constructing a mine ventilation knowledge graph based on a large model, the problems of data silos and insufficient knowledge expression in traditional mine ventilation systems are solved, the structured and intelligent management of mine ventilation knowledge is realized, and the system's intelligent perception and response capabilities are improved.
Patent Information
- Application Number
- CN202510597506.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional mine ventilation systems have data silos, information fragmentation, and insufficient knowledge expression capabilities. Multi-source heterogeneous data have not been uniformly modeled, resulting in the inability to structure and dynamically express mine ventilation knowledge, which limits the system's intelligent perception and active response capabilities.
A large-model-based mine ventilation knowledge graph construction method is adopted, including multi-source heterogeneous text cleaning, ontology model construction, entity, attribute and relationship extraction, entity alignment and graph database construction, to form a graph database structure with clear structure, semantic coherence and clear hierarchy.
It realizes the automatic structured expression, semantically consistent fusion and intelligent visual query of mine ventilation knowledge, providing an explainable, scalable and reasonable intelligent support platform for mine ventilation systems.
Smart Images

Figure CN120706527A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of mine ventilation monitoring, and in particular relates to a method for constructing a mine ventilation knowledge graph based on a large model. Background Art
[0002] Mine ventilation safety is directly related to miners' lives and coal mine operational efficiency. While traditional ventilation systems can achieve partial automation and sensor monitoring, they face challenges such as data silos, information fragmentation, and insufficient knowledge representation.
[0003] Currently, multi-source heterogeneous data has not yet been uniformly modeled, resulting in the inability to structure and dynamically express mine ventilation knowledge, which limits the system's intelligent perception and active response capabilities.
[0004] Traditional rule-based knowledge systems have bottlenecks in dealing with unexpected working conditions, knowledge updating, and semantic reasoning, and are also costly to build and complex to maintain. Summary of the Invention
[0005] In order to solve at least one of the above-mentioned technical problems existing in the prior art, the present invention provides a method for constructing a mine ventilation knowledge graph based on a large model.
[0006] The present invention is implemented by the following technical solution: a method for constructing a mine ventilation knowledge graph based on a large model, comprising the following steps:
[0007] S100: Obtain and clean multi-source heterogeneous text related to mine ventilation to obtain knowledge fragments in a unified structure in JSON format;
[0008] S200: Construct an ontology model and define a semantic structure; the ontology model is used to output in a parseable JSON / OWL format, including an entity type definition table, a relationship definition table, and an attribute structure table;
[0009] S300: Construct four types of knowledge extraction tasks to extract entities, attributes, and relationships;
[0010] S400: Build an entity alignment module to form a trusted alignment mapping table;
[0011] S500: Build a graph database structure with clear structure, semantic coherence, and clear hierarchy.
[0012] Preferably, the step of acquiring multi-source heterogeneous texts related to mine ventilation and cleaning them to obtain knowledge fragments in a unified structure in JSON format includes:
[0013] S101: Extract the content of ventilation-related pages from the document website; use structured path expressions to extract the title and body fields from the page, and construct standardized page structure data;
[0014] S102: Segment the acquired page content and extract structural information including chapter number, text body, applicable conditions, and reference clauses using expression templates; divide the text into rule description segments with the smallest semantic units;
[0015] S103: Constructing format standardization and anomaly detection rules to process the divided text content;
[0016] S104: Design an automatic segmentation mechanism based on semantic coherence and punctuation to break down long text into multiple logically independent content segments and append structure fields to each segment;
[0017] S105: Through field mapping, format normalization and content integration operations, the heterogeneous source content is converted into JSON format knowledge fragments with a unified structure.
[0018] Preferably, the steps of constructing an ontology model and defining a semantic structure include:
[0019] S201: Construct three core entity types, including physical entities, functional entities, and dynamic entities;
[0020] S202: Based on the semantic associations between entities, three core relationship types are constructed, including spatial topological relationships, functional control relationships, and dynamic influence relationships;
[0021] S203: construct three types of attribute fields, including static attributes, dynamic attributes, and constraint attributes;
[0022] S204: Combine the semantic structures in steps S201 to S203 to construct entity-relationship-attribute triple expressions to form a unified ontology template; each triple structure is represented as (e i ,r ij ,e j ) or (e i ,a k ,v k ), where e i ,e j is the entity node, r ij is the relationship type, a k is the attribute key, v k is the attribute value;
[0023] S205: Generate ontology model output. The ontology model output is in a parseable JSON / OWL format, including an entity type definition table, a relationship definition table, and an attribute structure table.
[0024] Preferably, the physical entities are physical objects including fans, dampers, ventilation tunnels, and gas sensors; the functional entities are abstract elements used to describe the regulation mechanism, including regulation strategies and ventilation control logic; and the dynamic entities are operating parameters with time-varying characteristics, including wind speed, pressure, and gas concentration.
[0025] The spatial topological relationship represents the connectivity and adjacency between physical structures; the functional control relationship is used to describe the action path of the functional entity on the physical device or parameter; the dynamic influence relationship is used to characterize the causal relationship or constraint logic between dynamic variables;
[0026] The static attributes include the designed cross-sectional area, rated air volume, and equipment model; the dynamic attributes include the real-time wind speed and gas concentration; and the constraint attributes include the maximum allowable concentration and minimum wind speed requirements.
[0027] Preferably, the steps of constructing four types of knowledge extraction tasks and extracting entities, attributes and relationships include:
[0028] S301: Constructing four types of knowledge extraction tasks based on the diversity of knowledge types in ventilation scenarios; the knowledge extraction tasks include information extraction tasks, knowledge extraction tasks, entity and attribute extraction tasks, and relationship and rule extraction tasks;
[0029] S302: Build a unified prompt input template for each type of task, including the following structure: instruction, input, and output;
[0030] S303: Programmatically call the DeepSeek-R1 interface, send templated input in batches, and return structured output through the interface;
[0031] S304: Convert the model output into the following two structures: a triple structure: {subject, predicate, object}; a multi-attribute structure: {entity, attribute, value, type}; the model output is used for subsequent entity modeling and edge relationship generation of the graph;
[0032] S305: Set up a sampling evaluation process for the extraction results, record the model performance indicators, ensure that the extraction accuracy of entity attribute tasks is not less than 95%, and the accuracy of relationship and rule tasks is not less than 95%; finally output a JSONL file in a unified format for graph construction.
[0033] Preferably, the step of constructing an entity alignment module and forming a trusted alignment mapping table includes:
[0034] S401: Extract the subject and object fields from the two heterogeneous graphs respectively, and form two entity sets denoted as ε1={e 11 ,e 12 ,...,e 1m} and ε2={e 21 ,e 22 ,...,e 2n}, semantic encoding is performed on each entity name using the Chinese sentence vector model to obtain its semantic vector representation;
[0035] S402: Calculate the semantic similarity matrix S based on the semantic vector representation name , which is defined as:
[0036]
[0037] in, and Represent the vector representation of the entities in ε1 and ε2 respectively, and calculate their cosine similarity;
[0038] S403: Construct an entity adjacency graph, extract the adjacency set N(e) of each entity, and construct a structural similarity matrix S struct , which is defined as:
[0039] S struct (i,j)=|N(e 1i )∩N(e 2j )| / |N(e 1i )∩N(e 2j )|
[0040] Among them, N(e) represents the set of adjacent entities of entity e, which is used to measure structural similarity;
[0041] S404: Introduce a fusion strategy, linearly weight the semantic similarity and structural similarity, and define the fusion similarity matrix:
[0042] S fused =λ·S name +(1-λ)·S struct ,λ=0.6
[0043] Among them, λ is the weight of semantic information, which defaults to 0.6;
[0044] S405: For each entity in ε1, select the entity with the highest fusion similarity in ε2 as its matching item, build a set of candidate entity pairs, and record its fusion score for subsequent screening;
[0045] S406: Input the candidate entity pairs into the structured judgment template and call the language model to judge whether the entity pairs are semantically consistent; the model output is "yes" or "no", indicating whether it constitutes a valid alignment; finally, all entity pairs judged as "yes" are summarized to form a trusted alignment mapping table.
[0046] Preferably, the step of constructing a graph database structure with a clear structure, coherent semantics, and clear hierarchy includes:
[0047] S501: Based on the entity alignment results in step S406, the subject and object fields in the different source graphs are uniformly replaced, and the entity fields of all triples are replaced with the main standard entity in the standard procedure graph;
[0048] S502: Extracting technical measure entities and their entry contents from the structured results of Task 1 in the webpage corpus, converting them into triple expressions by semantic mapping with the entities in the regulation graph, and incorporating them into a unified structure set;
[0049] S503: Applying a hash algorithm based on a subject-verb-object structure to the fused triple set to remove redundancy and perform semantic disambiguation processing;
[0050] S504: Based on the entity type system defined in the above extraction, an entity attribute field system is established, and its static attributes, dynamic attributes, and constraint attributes are extracted respectively;
[0051] S505: Convert the above structured entities, attributes, and relationship contents into node and edge formats recognizable by the graph database; the node information includes name, category, and attribute value, and the edge information includes relationship type and logical path annotation; the entity type field is automatically mapped to the node label in Neo4j;
[0052] S506: After the graph is constructed, the system automatically generates semantic aggregation nodes based on the "belongs to" relationship and connects them to related nodes, ultimately forming a graph database structure with a clear structure, semantic coherence, and clear hierarchy.
[0053] Compared with the prior art, the present invention has the following beneficial effects:
[0054] The present invention is a method for constructing a knowledge graph for mine ventilation based on a large model, which includes the steps of data collection, ontology modeling, knowledge extraction, semantic alignment, and graph construction. Data collection forms standardized input through the automatic collection and cleaning of multi-source heterogeneous texts; ontology modeling provides a unified semantic structure for knowledge organization; knowledge extraction combines the capabilities of the large model to achieve accurate extraction of entities, attributes, and relationships; semantic alignment ensures consistency and standardization of terms across sources; and graph construction utilizes a graph database to achieve visualization and continuous updating of structured knowledge. This method can realize the automatic structured expression, semantic consistency fusion, and intelligent visual query of knowledge in the field of mine ventilation, providing an interpretable, scalable, and reasonable intelligent support platform for mine ventilation systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0056] Figure 1 Method steps for constructing a knowledge graph for mine ventilation based on a large model;
[0057] Figure 2 Construct a flow chart for the mine ventilation knowledge graph;
[0058] Figure 3 It is a flowchart of multi-source data collection and processing;
[0059] Figure 4 This is the structural diagram of the mine ventilation model;
[0060] Figure 5 Labeling diagrams for knowledge extraction tasks;
[0061] Figure 6 This is a schematic diagram of the Prompt template;
[0062] Figure 7 It is the distribution map of the semantic similarity matrix between book and website graph entities;
[0063] Figure 8 It is the similarity matrix distribution diagram of the entity structure between books and website graphs;
[0064] Figure 9 It is the t-NSE two-dimensional distribution map of entity vectors of books and websites;
[0065] Figure 10 This is the mine ventilation knowledge map (part). DETAILED DESCRIPTION
[0066] The technical solutions in the embodiments of the present invention are clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other implementations derived by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.
[0067] It should be noted that the structures, proportions, sizes, etc. illustrated in the drawings of this specification are only used to match the contents disclosed in the specification for people familiar with this technology to understand and read, and are not used to limit the conditions under which the present invention can be implemented. Therefore, they have no substantive technical significance. Any structural modification, change in proportional relationship or adjustment of size should fall within the scope of the technical content disclosed in the present invention without affecting the efficacy and purpose that can be achieved by the present invention. It should be noted that in this specification, relational terms such as first and second are only used to distinguish one entity from several other entities, and do not necessarily require or imply any actual relationship or order between these entities.
[0068] The present invention provides an embodiment:
[0069] The present invention proposes a method for constructing a mine ventilation knowledge graph based on a large model. The method includes five core steps: data collection, ontology modeling, structured knowledge extraction, semantic alignment, and graph storage. In this embodiment, a complete knowledge graph is constructed based on coal mine ventilation safety clauses, website data, and book data. The flow chart is shown in Figure 1 and Figure 2 The present invention adopts the following technical solutions:
[0070] S100: Build a multi-source heterogeneous data collection and cleaning module, which specifically includes the following steps:
[0071] S101: Web page data collection:
[0072] Build a data collection module based on a web crawler framework. By configuring request scheduling and page parsing rules, extract the content of ventilation pages on technical documentation websites. Set classification rules, use structured path expressions to extract fields such as titles and body text from pages, and build standardized page structure data.
[0073] In this embodiment, a distributed web crawler system is constructed, with classification rules and crawl depth set for typical technology websites, and appropriate time intervals set to control the frequency of collection. The crawler program automatically extracts the technical title, description field, and paragraph body of the page, and converts the web page structure into page-level JSON data format.
[0074] S102: Standard document analysis:
[0075] The acquired regulatory text and the main text from professional literature are segmented, and structural information such as chapter numbering, text body, applicable conditions, and reference clauses are extracted using expression templates. The extracted results are uniformly encoded as knowledge item fragments with independent semantic unit attributes. The text is then divided into rule description fragments with the smallest semantic units for subsequent structure extraction.
[0076] S103: Format cleaning and exception filtering:
[0077] Format standardization and anomaly detection rules are constructed to process the captured text content, including tag removal, abnormal character removal, and duplicate content identification. A threshold is set to identify extremely short or structurally missing samples, which are then filtered and marked as noise. In this embodiment, this includes removing redundant content such as HTML tags, tabs, image identifiers, and extra spaces. An anomaly detection function is constructed to filter out text segments with fewer than 10 characters, delete completely duplicate content, and retain the most semantically complete version.
[0078] S104: Automatic segmentation and structure annotation:
[0079] An automatic segmentation mechanism based on semantic coherence and punctuation is designed to break down long texts into multiple logically independent content fragments. Structural fields such as source type, content location index, and text category identifier are appended to each fragment to form a unified data structure.
[0080] In this example, we build automatic segmentation logic through semantic and punctuation analysis to segment long, technical and regulatory texts into structured segments. We add structural annotations to each segment, including information such as the data source type (webpage / regulation / textbook), chapter index, and classification tags, and output the result as a structured corpus in JSON format.
[0081] S105: Standardized knowledge fragment generation:
[0082] Through field mapping, format normalization, and content integration, heterogeneous source content is ultimately transformed into knowledge fragments in a unified JSON format. This data structure possesses stable semantic boundaries and structural consistency, providing a reliable data foundation for subsequent entity extraction, relationship building, and graph generation.
[0083] In this embodiment, the processed text is uniformly archived and abstracted into a computable semantic structure. The output fields include: "source_type", "section_id", "text", "label", etc., providing a standard input format for subsequent modules. Figure 3 Shows the schematic structure of the processed fragment.
[0084] S200: Ontology model design and semantic structure definition, specifically including the following steps:
[0085] S201: Entity type definition:
[0086] Using an expert-led approach and combining the composition and functions of the ventilation system, three core entity types are set: physical entities: including physical objects such as fans, dampers, ventilation tunnels, and gas sensors; functional entities: including abstract elements used to describe the regulation mechanism, such as control strategies and ventilation control logic; dynamic entities: referring to operating parameters with time-varying characteristics such as wind speed, pressure, and gas concentration.
[0087] S202: Relationship type definition:
[0088] Based on the semantic associations between entities, three core relationship types are designed: spatial topological relationship: representing the connectivity and adjacency between physical structures, such as "air gates connect air lanes" and "coal mining faces are adjacent to goafs"; functional control relationship: describing the action path of functional entities on physical equipment or parameters, such as "ventilators control air volume" and "adjusting air gates affects wind speed"; dynamic influence relationship: depicting the causal relationship or constraint logic between dynamic variables, such as "wind speed changes lead to gas accumulation."
[0089] S203: Attribute system settings:
[0090] Three types of attribute fields are defined for the entities in the ontology: static attributes: such as designed cross-sectional area, installation location, rated parameters, equipment model, etc.; dynamic attributes: such as real-time wind speed (in m / s) and gas concentration (in %), which are collected by sensors; constraint attributes: such as maximum allowable concentration and minimum wind speed requirements, such as the maximum wind speed threshold ≥8.0m / s and the concentration must not exceed 1.5%, etc., which are used to build control rules.
[0091] S204: Unified modeling of ontology structure:
[0092] Combining the semantic structures from S201 to S203, we construct entity-relationship-attribute triples to form a unified ontology template. Each triple structure is represented as (e i ,r ij ,e j ) or (e i ,a k ,v k ), where e i ,e j is the entity node, r ij is the relationship type, a k is the attribute key, v k is the attribute value.
[0093] In this embodiment, the following structure examples are constructed: physical entity triple: (local ventilator, installed in, return air channel); attribute structure example: (gas sensor, concentration threshold, ≤1.5%); dynamic relationship structure: (wind speed reduction, trigger, gas accumulation);
[0094] The above structure is uniformly modeled as a JSON ontology template, with fields including entity_type, relation_type, attribute_key, attribute_unit, constraint_value, etc.
[0095] S205: Generate ontology model output
[0096] The ontology model is output in a parsable JSON / OWL format, including entity type definition table, relationship definition table and attribute structure table, providing standardized semantic input for modules such as graph extraction, reasoning and alignment. Figure 4 Shows the semantic network structure diagram between entities, relationships and attributes.
[0097] S300: Knowledge extraction process design, specifically including the following steps:
[0098] S301: Extraction task system design:
[0099] Based on the diversity of knowledge types in ventilation scenarios, four types of knowledge extraction tasks are constructed, namely: Task 1: Information extraction, oriented to web page texts, identifies elements such as equipment, objects, and technical descriptions in ventilation measures, and identifies fields such as technical measure titles, equipment names, and measure descriptions; Task 2: Knowledge extraction, oriented to textbooks and standard books, extracts knowledge about ventilation principles and causal reasoning, and identifies entities such as "local ventilation fan" and their corresponding attributes such as "wind speed ≥ 1.5 m / s" from regulatory texts, with attribute types marked as dynamic attributes; Task 3: Entity and attribute extraction, identify physical, functional, dynamic entities and their static, dynamic and constraint attributes, identify entities such as "local ventilation fan" and their corresponding attributes such as "wind speed ≥ 1.5m / s" from the regulations, and mark the attribute type as dynamic attribute; Task 4: Relationship and rule extraction, identify the topology, control and causal relationship between entities, and the "if-then" rule structure, identify trigger rules such as "gas concentration is too high → cut off power", and extract relationship types such as TRIGGERS or CONTROLS. The task system structure is as follows: Figure 5 shown.
[0100] S302: Prompt template construction and input design:
[0101] A unified prompt input template is built for each task type. The structure includes: instruction: explicitly indicating the current task objective; input: the raw text corpus to be processed; output: the structured format of the target, required to be a standard triple or attribute object. This template structure adapts to the input structure of the large model API and supports switching between tasks.
[0102] In this embodiment, each task design corresponds to a prompt template, which includes an instruction (such as "Please extract entities and their attribute types from the following text"), an input original sentence, and an output JSON structure result. Take Task 3 as an example:
[0103] {"instruction":"Extract entities, attributes and their types contained in the sentence",
[0104] "input": "The wind speed of the local ventilation fan shall not be less than 1.5m / s.",
[0105] "output":{"entity":"local ventilator","attribute":"wind speed","type":"dynamic attribute","comparison relation":"≥","value":"1.5","unit":"m / s"}}.
[0106] Template structure such as Figure 6 shown.
[0107] S303: Extract the calling process implementation:
[0108] The DeepSeek-R1 interface is called programmatically, sending templated inputs in batches and returning structured outputs through the interface. Each task is configured with sample and temperature parameters to control generation stability and accuracy.
[0109] In this embodiment, Python is used to call the DeepSeek-R1 interface to submit structured Prompt data. Task three is set to use 16 example samples to optimize the extraction of context prompts, and Task four uses 8 samples to control the logical relationship recognition effect, and automatically returns structured results.
[0110] S304: Structuring and categorizing the extraction results:
[0111] The model output is uniformly converted into the following two structures: triple structure: {subject, predicate, object};
[0112] Multi-attribute structure: {entity, attribute, value, type}. The results are archived and saved as a .jsonl file for subsequent entity modeling and edge relationship generation in the graph.
[0113] S305: Extraction quality assessment and storage:
[0114] A sampling evaluation process is set up for the extraction results, and model performance indicators are recorded to ensure that the accuracy of entity attribute extraction tasks is no less than 95%, and the accuracy of relationship and rule extraction tasks is no less than 95%. Finally, a unified JSONL file is output for graph construction.
[0115] In this example, from a manual sampling of 500 examples, the accuracy of entity attribute extraction in Task 3 was 95.0%, and the accuracy of relational rule extraction in Task 4 was 99.1%. These results meet the requirements for high-quality ventilation map extraction and serve as a reference for model deployment quality verification.
[0116] S400: Entity alignment module construction, specifically including the following steps:
[0117] S401: Entity set extraction and semantic vector encoding:
[0118] Extract the subject and object fields from the two heterogeneous graphs respectively to form two entity sets denoted as ε1 = {e 11 ,e 12 ,...,e 1m} and ε2={e 21 ,e 22 ,...,e 2n}, semantic encoding is performed on each entity name using the Chinese sentence vector model to obtain its semantic vector representation.
[0119] In this embodiment, the subject and object fields in the triples are extracted from the standard procedure map and the website class map respectively, and two entity sets are constructed, which contain about 5700 entities in total. The entity names are encoded using the text2vec-large model to obtain semantic vector representation. The semantic similarity values of 32,000 pairs of entities are preliminarily calculated. The distribution diagram of the semantic and structural similarity matrix of the book and website map entities is shown in the figure below. Figure 7 and Figure 8 As shown, the t-NSE two-dimensional distribution diagram of the entity vectors of books and websites is as follows Figure 9 shown.
[0120] S402: Semantic similarity matrix construction:
[0121] Based on the vector representation, calculate the semantic similarity matrix S name , which is defined as:
[0122]
[0123] in, and Represent the vector representation of the entities in ε1 and ε2 respectively, and calculate their cosine similarity.
[0124] In this embodiment, the semantic similarity matrix S is calculated name , calculate the cosine similarity for all entity pairs. Statistics show that there are more than 9,000 entity pairs with semantic similarity greater than 0.6, providing a basis for subsequent screening.
[0125] S403: Structural similarity matrix construction:
[0126] Construct an entity adjacency graph, extract the adjacency set N(e) of each entity, and construct a structural similarity matrix S struct , which is defined as:
[0127]
[0128] Among them, N(e) represents the set of adjacent entities of entity e, which is used to measure structural similarity.
[0129] In this embodiment, an adjacency graph is constructed for all entities in the two graphs, the adjacency set N(e) of each entity is extracted, and the structural similarity matrix S is calculated using Jaccard similarity. struct Due to differences in graph structures, about 70% of entity pairs have a structural similarity of 0, and there are less than 1,000 entity pairs with a structural similarity greater than 0.3.
[0130] S404: Fusion similarity calculation:
[0131] A fusion strategy is introduced to linearly weight the semantic similarity and structural similarity, and the fusion similarity matrix is defined as:
[0132] S fused =λ·S name +(1-λ)·S struct ,λ=0.6
[0133] Among them, λ is the weight of semantic information, which defaults to 0.6.
[0134] In this embodiment, a fusion strategy is adopted:
[0135] S fused =0.6·S name +0.4·S struct
[0136] After calculation, for each entity, the candidate entity with the highest fusion score is selected from the target graph, and about 9,000 preliminary candidate entity pairs are obtained.
[0137] S405: Greedy matching generates candidate entity pairs:
[0138] For each entity in ε1, the entity with the highest fusion similarity in ε2 is selected as its matching item, a set of candidate entity pairs is constructed, and its fusion score is recorded for subsequent screening.
[0139] In this embodiment, each candidate entity pair is constructed as a judgment prompt, for example:
[0140] Entity 1: Local ventilation fan; Entity 2: Local area fan. Please determine whether these two entities are semantically consistent. Answer only "yes" or "no".
[0141] The DeepSeek-R1 interface is called to perform semantic consistency judgment, and a total of 1,637 pairs of entity pairs with "yes" are output.
[0142] S406: Large model semantic judgment verification:
[0143] Candidate entity pairs are input into a structured judgment template, and a language model (such as DeepSeek-R1) is used to determine whether the entity pairs are semantically consistent. The model outputs a "yes" or "no" response, indicating whether a valid alignment is achieved. Finally, all entity pairs judged as "yes" are summarized to form a trusted alignment mapping table for use in graph fusion and reasoning.
[0144] In this example, manual spot checks were performed on entity pairs with a "yes" judgment result, with a 100% accuracy rate. This ultimately resulted in a trusted entity mapping table for use in downstream graph fusion and reasoning task modules, significantly improving entity normalization and cross-graph connectivity.
[0145] S500: Graph database construction and graph fusion visualization, specifically including the following steps:
[0146] S501: Unified entity naming and triple normalization:
[0147] Based on the entity alignment results obtained in step S406, the subject and object fields in the different source graphs are uniformly replaced. The entity fields of all triples are replaced with the main standard entity in the standard procedure graph to ensure the consistent expression of multi-source semantics in a unified space.
[0148] In this embodiment, the synonymous entities in the book map and the website map are uniformly mapped into the standard ventilation terminology vocabulary. For example, "main return air channel" and "main return air channel" are uniformly classified into the "main return air channel" node to realize the semantic merging of multi-source entities.
[0149] S502: Cross-source knowledge fusion and mapping structure generation:
[0150] The technical measure entity (measure_entity) and its entry content (measure) from the structured results of Task 1 in the web page corpus are extracted, and through semantic mapping with the procedure graph entity, they are converted into triple expressions and incorporated into a unified structure set.
[0151] In this embodiment, "technical measures items" and their corresponding descriptive information, such as "ventilation switching principles during fan maintenance", are extracted from the web page extraction structure, and a mapping relationship is established between them and control behaviors such as "fan switching strategy" in the existing triples to complete the semantic chain.
[0152] S503: Structure deduplication and entry information completion:
[0153] A hashing algorithm based on the subject-verb-object structure is applied to the fused triples to remove redundancy and perform semantic disambiguation. A subject-object matching mechanism is introduced to complement the measure_entity and its measure description in the fused graph, improving the graph's item coverage and content readability.
[0154] In this example, the following triples appear repeatedly in different source graphs:
[0155] (Local ventilation fan)—[wind speed ≥ 1.5 m / s]→(safety regulations);
[0156] (Local ventilation equipment)—[wind speed ≥ 1.5 m / s]→(operation requirements);
[0157] After normalization, the data are uniformly converted into: (local ventilation fan)-[wind speed ≥ 1.5m / s]→(safety specification); and the hash comparison algorithm is used to remove duplicates.
[0158] S504: Multi-type entity attribute modeling and specification cleaning:
[0159] Based on the entity type system defined in the previous extraction, an entity attribute field system is established, and its static attributes, dynamic attributes, and constraint attributes are extracted separately. All attribute values are normalized to remove redundant unit symbols, unstructured numbers, and abnormal values, retaining the semantically clear and length-standardized field content.
[0160] In this example, the extracted fields for "wind speed should be no less than 1.5 m / s" are standardized as {entity: local ventilator, attribute: wind speed, operator: ≥, value: 1.5, unit: m / s}. Auxiliary words such as "should" and "no less than" are removed from the text to achieve standardized parameter format.
[0161] S505: Graph database format conversion and node edge generation:
[0162] Convert the structured entities, attributes, and relationships described above into a node and edge format recognizable by graph databases. Node information includes name, category, and attribute values, while edge information includes relationship type and logical path annotations. The entity type field is automatically mapped to the node label in Neo4j.
[0163] In this embodiment, the aforementioned entities and attributes are mapped to nodes and edges in a graph database, such as {node: [local ventilator] {type: RealObject}, node: [wind speed ≥ 1.5 m / s] {type: Condition}, edge: [HAS_CONDITION]). In the visualization, each node type is distinguished by a different shape or grayscale tone to improve structural readability.
[0164] S506: Graph construction and visualization integration:
[0165] Once the graph is constructed, the system automatically generates semantically aggregated nodes based on "belongs to" relationships (e.g., "technical measures" aggregates all entities and behaviors involved) and connects them to related nodes. This ultimately creates a well-structured, semantically coherent, and layered graph database, enabling visualization of entity associations, path reasoning, and dynamic updates.
[0166] In this embodiment, after the graph is constructed, a Neo4j browser is used for visualization. Figure 10 shown.
[0167] Through the above examples, the application process of the method of the present invention on real procedure data is fully demonstrated, indicating that it has engineering implementation capabilities and logical closed-loop properties.
[0168] The foregoing description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be readily conceived by a person skilled in the art within the technical scope disclosed herein should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A method for constructing a mine ventilation knowledge graph based on a large model, characterized in that , including the following steps: S100: Obtain and clean multi-source heterogeneous text related to mine ventilation to obtain knowledge fragments in a unified structure in JSON format; S200: Construct an ontology model and define a semantic structure; the ontology model is used to output in a parseable JSON / OWL format, including an entity type definition table, a relationship definition table, and an attribute structure table; S300: Construct four types of knowledge extraction tasks to extract entities, attributes, and relationships; S400: Build an entity alignment module to form a trusted alignment mapping table; S500: Build a graph database structure with clear structure, semantic coherence, and clear hierarchy.
2. The method for constructing a mine ventilation knowledge graph based on a large model according to claim 1, characterized in that: The steps of obtaining multi-source heterogeneous texts related to mine ventilation and cleaning them to obtain knowledge fragments in a unified structure in JSON format include: S101: Extract the content of ventilation-related pages from the document website; use structured path expressions to extract the title and body fields from the page, and construct standardized page structure data; S102: Segment the acquired page content and extract structural information including chapter number, text body, applicable conditions, and reference clauses using expression templates; divide the text into rule description segments with the smallest semantic units; S103: Constructing format standardization and anomaly detection rules to process the divided text content; S104: Design an automatic segmentation mechanism based on semantic coherence and punctuation to break down long text into multiple logically independent content segments and append structure fields to each segment; S105: Through field mapping, format normalization and content integration operations, the heterogeneous source content is converted into JSON format knowledge fragments with a unified structure.
3. The method for constructing a mine ventilation knowledge graph based on a large model according to claim 2 is characterized in that: The steps of constructing the ontology model and defining the semantic structure include: S201: Construct three core entity types, including physical entities, functional entities, and dynamic entities; S202: Based on the semantic associations between entities, three core relationship types are constructed, including spatial topological relationships, functional control relationships, and dynamic influence relationships; S203: construct three types of attribute fields, including static attributes, dynamic attributes, and constraint attributes; S204: Combine the semantic structures in steps S201 to S203 to construct entity-relationship-attribute triple expressions to form a unified ontology template; each triple structure is represented as (e i ,r ij ,e j ) or (e i ,a k ,v k ), where e i ,e j is the entity node, r ij is the relationship type, a k is the attribute key, v k is the attribute value; S205: Generate ontology model output. The ontology model output is in a parseable JSON / OWL format, including an entity type definition table, a relationship definition table, and an attribute structure table.
4. The method for constructing a mine ventilation knowledge graph based on a large model according to claim 3 is characterized in that: The physical entities are physical objects including fans, dampers, ventilation tunnels, and gas sensors; the functional entities are abstract elements used to describe the regulation mechanism, including regulation strategies and ventilation control logic; the dynamic entities are operating parameters with time-varying characteristics, including wind speed, pressure, and gas concentration; The spatial topological relationship represents the connectivity and adjacency between physical structures; the functional control relationship is used to describe the action path of the functional entity on the physical device or parameter; the dynamic influence relationship is used to characterize the causal relationship or constraint logic between dynamic variables; The static attributes include the designed cross-sectional area, rated air volume, and equipment model; the dynamic attributes include the real-time wind speed and gas concentration; and the constraint attributes include the maximum allowable concentration and minimum wind speed requirements.
5. The method for constructing a mine ventilation knowledge graph based on a large model according to claim 4 is characterized in that: The steps of constructing four types of knowledge extraction tasks and extracting entities, attributes and relationships include: S301: Constructing four types of knowledge extraction tasks based on the diversity of knowledge types in ventilation scenarios; the knowledge extraction tasks include information extraction tasks, knowledge extraction tasks, entity and attribute extraction tasks, and relationship and rule extraction tasks; S302: Build a unified prompt input template for each type of task, including the following structure: instruction, input, and output; S303: Programmatically call the DeepSeek-R1 interface, send templated input in batches, and return structured output through the interface; S304: Convert the model output into the following two structures: a triple structure: {subject, predicate, object}; a multi-attribute structure: {entity, attribute, value, type}; the model output is used for subsequent entity modeling and edge relationship generation of the graph; S305: Set up a sampling evaluation process for the extraction results, record the model performance indicators, ensure that the extraction accuracy of entity attribute tasks is not less than 95%, and the accuracy of relationship and rule tasks is not less than 95%; finally output a JSONL file in a unified format for graph construction.
6. The method for constructing a mine ventilation knowledge graph based on a large model according to claim 5 is characterized in that: The step of constructing an entity alignment module and forming a trusted alignment mapping table includes: S401: Extract the subject and object fields from the two heterogeneous graphs respectively, and form two entity sets denoted as ε1={e 11 ,e 12 ,...,e 1m } and ε2={e 21 ,e 22 ,...,e 2n }, semantic encoding is performed on each entity name using the Chinese sentence vector model to obtain its semantic vector representation; S402: Calculate the semantic similarity matrix S based on the semantic vector representation name , which is defined as: in, and Represent the vector representation of the entities in ε1 and ε2 respectively, and calculate their cosine similarity; S403: Construct an entity adjacency graph, extract the adjacency set N(e) of each entity, and construct a structural similarity matrix S struct , which is defined as: S struct (i,j)=|N(e 1i )∩N(e 2j )| / |N(e 1i )∩N(e 2j )| Among them, N(e) represents the set of adjacent entities of entity e, which is used to measure structural similarity; S404: Introduce a fusion strategy, linearly weight the semantic similarity and structural similarity, and define the fusion similarity matrix: S fused =λ·S name +(1-λ)·S struct ,λ=0.6 Among them, λ is the weight of semantic information, which defaults to 0.6; S405: For each entity in ε1, select the entity with the highest fusion similarity in ε2 as its matching item, build a set of candidate entity pairs, and record its fusion score for subsequent screening; S406: Input the candidate entity pairs into the structured judgment template and call the language model to judge whether the entity pairs are semantically consistent; the model output is "yes" or "no", indicating whether it constitutes a valid alignment; finally, all entity pairs judged as "yes" are summarized to form a trusted alignment mapping table.
7. The method for constructing a mine ventilation knowledge graph based on a large model according to claim 6, characterized in that: The steps of constructing a graph database structure with a clear structure, coherent semantics, and clear hierarchy include: S501: Based on the entity alignment results in step S406, the subject and object fields in the different source graphs are uniformly replaced, and the entity fields of all triples are replaced with the main standard entity in the standard procedure graph; S502: Extracting technical measure entities and their entry contents from the structured results of Task 1 in the webpage corpus, converting them into triple expressions by semantic mapping with the entities in the regulation graph, and incorporating them into a unified structure set; S503: Applying a hash algorithm based on a subject-verb-object structure to the fused triple set to remove redundancy and perform semantic disambiguation processing; S504: Based on the entity type system defined in the above extraction, an entity attribute field system is established, and its static attributes, dynamic attributes, and constraint attributes are extracted respectively; S505: Convert the above structured entities, attributes, and relationship contents into node and edge formats recognizable by the graph database; the node information includes name, category, and attribute value, and the edge information includes relationship type and logical path annotation; the entity type field is automatically mapped to the node label in Neo4j; S506: After the graph is constructed, the system automatically generates semantic aggregation nodes based on the "belongs to" relationship and connects them to related nodes, ultimately forming a graph database structure with a clear structure, semantic coherence, and clear hierarchy.
Citation Information
Cited By
Semantic modeling method and system for expressway construction safety training
CN121681798A
Multi-level well drilling knowledge base intelligent construction method and device and computer equipment
CN121902941A
Multi-level drilling knowledge base intelligent construction method and device and computer equipment
CN121902941B
Multi-source heterogeneous data fusion and knowledge graph automatic construction system
CN121981232A
A multi-source heterogeneous data fusion and knowledge graph automatic construction system
CN121981232B