Knowledge fusion based method and system for collaborative construction of knowledge in the field of hydraulic and environmental geology
By performing structured extraction and semantic deviation identification on engineering data in the field of hydrogeology and environmental geology, the problems of information redundancy and structural conflict in traditional methods are solved, and efficient dynamic collaborative management and consistent expression of multi-source data are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 四川省第十地质大队
- Filing Date
- 2026-02-28
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional methods for collaboratively constructing hydrogeological and environmental knowledge rely on manual compilation and expert experience, which cannot adapt to the dynamic characteristics of frequent updates to engineering data and concurrent input from multiple entities. They lack semantic alignment capabilities, resulting in information redundancy and difficulty in identifying structural conflicts. Furthermore, the lack of dynamic alignment and revision mechanisms affects the accuracy and timeliness of multi-source knowledge collaboration.
By acquiring engineering data from landslide hazard assessment zones and groundwater dynamic monitoring sections, we extract lithological descriptions, spatial boundaries, and hydrological conditions. We construct semantic segments and perform structured grouping, identify semantic deviations, split conflicting segments, analyze term frequency and consistency trends, generate collaborative construction aggregation identifiers, screen standard semantic content, and establish fusion structural units.
It improves the efficiency of orderly organization and management of multi-source data, enhances semantic standardization and expression consistency, realizes dynamic collaborative management of heterogeneous geological data throughout the entire process, and strengthens the accuracy and timeliness of knowledge fusion.
Smart Images

Figure CN122132578A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of knowledge graph technology, and in particular to a method and system for collaborative construction of knowledge in the field of hydrogeology and environmental geology based on knowledge fusion. Background Technology
[0002] The field of knowledge graph technology mainly involves core aspects such as data semantic modeling, knowledge extraction, knowledge fusion, knowledge representation, and management for specific professional domains. Its overall approach includes using natural language processing, information extraction, entity alignment, and relational reasoning to extract semantic information from multi-source heterogeneous data, constructing a structured, queryable, and reasonable knowledge network to achieve systematic organization and intelligent management of professional domain knowledge. Traditional methods for collaboratively constructing hydrogeological and environmental geological knowledge address fragmented geological knowledge distributed across different industry stakeholders in fields such as hydrogeology, engineering geology, and environmental geology. These methods rely on manual summarization and expert experience-based classification, integrating existing knowledge through manual data entry, static knowledge template maintenance, or static document integration. Content extraction and archiving are based on geological reports, exploration documents, monitoring records, and other literature. This approach primarily relies on targeted knowledge integration and manual review by domain experts to achieve initial knowledge integration, lacking the ability to semantically align heterogeneous data and dynamic knowledge collaboration strategies among different stakeholders.
[0003] Traditional knowledge integration relies on manual summarization and expert classification, with information processing efficiency limited by the speed of manual operations, failing to adapt to the dynamic characteristics of frequent updates to engineering data and concurrent input from multiple entities. Knowledge sources are primarily presented in static document forms such as geological reports and monitoring records, resulting in coarse content granularity, unclear semantic structure, and a lack of refined semantic tagging systems for comparison. This leads to problems such as information redundancy and difficulty in identifying structural conflicts during subsequent integration. Different participating entities submit content with varying expression styles, and current methods lack mechanisms to identify semantic deviations between expressions, resulting in low structural consistency and standardization of expression, and risks of logical gaps and contextual disconnects in the integrated content. During the integration phase, knowledge archiving is completed solely through static templates or targeted summarization, lacking the ability to track semantic changes and identify evolutionary paths. This makes it difficult to effectively manage the divergences and evolutionary trends exhibited by different knowledge construction paths, leading to poor traceability and weak collaborative construction capabilities. Regarding knowledge collaboration, traditional methods lack dynamic alignment and revision mechanisms, making it difficult to reach consensus on the same geological construction goal under different entity expressions, affecting the accuracy and timeliness of multi-source knowledge collaboration. Summary of the Invention
[0004] To address the technical problems existing in the prior art, this invention provides a knowledge collaborative construction method in the field of hydrogeology and environmental geology based on knowledge fusion, including the following steps:
[0005] S1: Obtain engineering data from landslide hazard assessment zones and groundwater dynamic monitoring sections, extract lithological descriptions, spatial boundaries and hydrological conditions submitted by participating entities, construct semantic fragments, engineering attribute items and applicability entries, and group them in a structured manner according to entity number to form a set of geological knowledge fragments;
[0006] S2: Extract keyword fragments, spatial descriptions and semantic tags from the geological knowledge fragment set for the same surface water-groundwater conversion zone division unit, and generate semantic deviation parameter sets by comparing term overlap, boundary intersection and tag differences;
[0007] S3: Identify the construction content with semantic conflicts based on the semantic deviation parameter set, split the conflicting segments from the original trajectory, combine the source subject, semantic label and construction scope to generate the current path, and establish a semantic divergence construction trajectory set;
[0008] S4: Construct a trajectory set based on the semantic divergence, perform trend analysis on term frequency, applicability coverage and semantic consistency, and mark objects that show convergence as aggregable objects to generate a collaborative construction aggregation identifier set.
[0009] S5: Extract representative semantic fragments, spatial ranges and construction tags based on the collaborative construction aggregation identifier set, filter standard semantic content, and generate construction fusion structural units by combining source subject and path information.
[0010] As a further aspect of the present invention, the geological knowledge fragment set includes lithological semantic fragments, spatial boundary information items, hydrological condition description items, engineering attribute item set, applicability item set, and subject number index item; the semantic deviation parameter set includes term overlap index, spatial boundary intersection range, semantic tag difference item, difference performance record item, and deviation type identifier item; the semantic divergence construction trajectory set includes semantic conflict fragment units, branch trajectory identifier items, source subject number identifier items, semantic tag type index item, construction range constraint item, and trajectory association marker item; the collaborative construction aggregation identifier set includes term usage frequency sequence, applicability coverage matrix item, semantic consistency trend item, aggregateable object list item, and aggregation judgment marker item; the construction fusion structure unit includes standard semantic content item, representative spatial range template item, construction tag specification item, contributing subject weight table item, path version identifier item, and fusion version structure framework item.
[0011] As a further aspect of the present invention, the specific steps of S1 are as follows:
[0012] S101: Obtain engineering data from landslide hazard assessment zones and groundwater dynamic monitoring sections, extract lithological descriptions, hydrological conditions and spatial boundary coordinates corresponding to the participating entity numbers, index and organize lithological particle size, water level change range and section coordinates according to the entity number, and generate a set of lithological and hydrological boundary indicators.
[0013] S102: Based on the lithological and hydrological boundary index set, extract the structural hierarchy sequence and spatial boundary combination form, classify them according to the stratigraphic order and boundary extension characteristics, divide the content with the same combination form into semantic fragment types, establish a corresponding relationship with the main number, and generate a semantic fragment index sequence.
[0014] S103: Based on the semantic fragment index sequence, combined with the spatial boundary coordinates and water level change range, the distribution range of the semantic fragments within the monitoring section is classified, and engineering attribute items and applicability items are constructed according to the boundary range and hydrological grouping relationship to generate a geological knowledge fragment set.
[0015] As a further aspect of the present invention, the specific steps of S2 are as follows:
[0016] S201: Based on the construction content of the main body in the geological knowledge fragment set regarding the same surface water-groundwater conversion zone division unit, extract keyword fragments, spatial description range and semantic tag information, call the stratigraphic name in the keyword fragment and the coordinate boundary in the spatial description range, sort out the construction content under the main body number, and generate a semantic construction feature set;
[0017] S202: Based on the semantic construction feature set, call the keyword fragments and the content pointed to by the semantic tags, determine the degree of word repetition in the same construction unit, divide the word consistency level according to the repetition rate, and perform difference classification processing in combination with the mapping relationship between the content pointed to by the semantic tags to generate a semantic overlap coefficient.
[0018] S203: For the semantic overlap coefficient, combined with the set of boundary coordinates in the spatial description range, determine the number of intersection segments and the overlap ratio of spatial boundaries within the same building unit, call the combination relationship between the number of intersections and the semantic overlap coefficient value, merge and express the degree of difference, and generate a semantic deviation parameter set.
[0019] As a further aspect of the present invention, during the generation of the semantic construction feature set, the stratigraphic names extracted from the keyword fragments are subjected to standardized comparison processing. The standardized comparison processing is limited to the scope of the construction content corresponding to the same subject number, and stratigraphic names with a matching degree lower than a preset comparison threshold are marked as inconsistent terms. The term consistency level is limited to no less than three discrete levels, and the calculation range of the repetition rate is limited to keyword fragments belonging to the same surface water-groundwater conversion zone division unit in the semantic construction feature set. During the generation of the semantic overlap coefficient, the mapping relationship between semantic tags pointing to content is limited to a one-to-one or many-to-one mapping form, and only semantic tags that are successfully mapped participate in the calculation of the semantic overlap coefficient. During the generation of the semantic deviation parameter set, the boundary coordinate set is limited to a set of closed polygons under the same coordinate reference system, and is combined and expressed based on the interval division results of the number of intersection segments and the overlap ratio.
[0020] As a further aspect of the present invention, the specific steps of S3 are as follows:
[0021] S301: Based on the difference information in the semantic deviation parameter group, call the subject number, semantic tag and construction range boundary in the construction content, compare and identify the semantic representation under the same conversion zone construction unit, filter out the fragment content with differences exceeding the identification threshold, and establish a semantic conflict fragment index set.
[0022] S302: Based on the semantic conflict fragment index set, extract the corresponding fragment position in the original main body trajectory, call the fragment start and end positions and corresponding numbers to perform splitting operation, peel the marked fragments from the main body trajectory, and record the source main body, semantic tag type and construction boundary range to generate a split fragment trajectory parameter group;
[0023] S303: Based on the segmented trajectory parameter group, using the main body number as the identification field, combined with semantic labels and boundary ranges, construct multiple trajectory paths with corresponding numbers, set the pointing identifiers between trajectories, generate a set of branch paths with associated number identifiers, and establish semantic divergence to construct the trajectory set.
[0024] As a further aspect of the present invention, during the establishment of the semantic conflict fragment index set, the difference magnitude exceeding the identification threshold is defined as the construction content whose corresponding value in the semantic deviation parameter group falls within a preset difference range, and the comparison identification is limited to the semantic expression range of construction units with the same subject number and belonging to the same surface water-groundwater conversion zone construction unit; in the splitting operation, the fragment start and end positions are defined as the sequence position identifiers continuously recorded in the original subject trajectory, and the splitting operation is defined as being performed only on fragments that completely cover the boundary of the construction range; the source subject, semantic label type, and construction boundary range recorded in the split fragment trajectory parameter group are defined as corresponding one-to-one with the semantic conflict fragment index set; during the generation of the branch path set, the pointing identifier is defined as establishing a one-way association relationship between trajectory paths corresponding to different subject numbers, and the subject number is used as a unique index field for differentiation.
[0025] As a further aspect of the present invention, the specific steps of S4 are as follows:
[0026] S401: Based on the semantic divergence, construct the semantic revision trajectory of the same rock mass structure stability control point in the trajectory set, extract the term record sequence in the construction path, call the number of times the term appears in each path, calculate the term usage frequency distribution value in the path, and generate the term usage frequency coefficient.
[0027] S402: Based on the term usage frequency coefficient, combined with the applicability label range and semantic revision content in the constructed path, determine the applicability coverage ratio of the path within the same control point range, extract the degree of association between the label and the path paragraph, and obtain the semantic coverage consistency trend.
[0028] S403: In response to the semantic coverage consistency trend, identify the range of changes in term usage frequency and the direction of changes in coverage ratio, determine whether paths show a gradual convergence during the revision process, filter the set of path numbers that reach the aggregation judgment threshold, and establish a collaborative construction aggregation identifier set.
[0029] As a further aspect of the present invention, the specific steps of S5 are as follows:
[0030] S501: Based on the path information recorded in the collaborative construction aggregation identifier set, extract the semantic fragments, spatial description range and construction tag content under the corresponding construction path, identify the completeness of term expression, boundary extension method and tag usage position in the path, filter the content that is representative of the path, and generate a representative item group of the aggregation path.
[0031] S502: Based on the aggregated path representative item group, the semantic fragment expression form, spatial description boundary range and tag setting field are compared item by item to determine the word structure repetition, boundary coordinate intersection ratio and tag association number consistency, filter content fragments with matching relationship and obtain standard semantic content items;
[0032] S503: Based on the standard semantic content item, call the source subject number and corresponding path identification information to uniformly identify the corresponding position of the content item in the path, and establish a content structure combination sequence with the set of numbers as the index field to generate a construction fusion structure unit.
[0033] A knowledge collaborative construction system for the field of hydrogeology and environmental geology based on knowledge fusion includes:
[0034] The knowledge fragment extraction module is used to perform S1: obtain engineering data formed in landslide hazard assessment zones and groundwater dynamic monitoring sections, extract lithological descriptions, spatial boundary information and hydrological condition descriptions submitted by participating entities, group semantic fragments, engineering attribute items and applicability items, and organize them in a structured manner according to the entity number to generate a set of geological knowledge fragments;
[0035] The difference parameter identification module is used to perform S2: based on the construction content of the main body of the geological knowledge fragment set about the same surface water-groundwater conversion zone division unit, extract keyword fragments, spatial description range and semantic tag information, identify the degree of overlap of terms, the intersection range of spatial description boundaries and the differences between the content pointed to by tags by comparison, integrate the difference performance, and generate a semantic deviation parameter group;
[0036] The divergence trajectory construction module is used to execute S3: based on the difference information in the semantic deviation parameter group, identify the construction content with semantic conflict, split the segments with obvious divergence from the original main trajectory, and generate the current construction path based on the source subject, semantic label type and construction scope information, construct multiple branch trajectories with associated identifiers, and generate a set of semantic divergence construction trajectories;
[0037] The aggregation path judgment module is used to execute S4: Based on the semantic revision trajectory of the same rock mass structure stability control point in the semantic divergence construction trajectory set, analyze the changing trend of the term usage frequency, applicability coverage and semantic consistency performance in the construction path. If the features show a gradual convergence state, mark the corresponding construction path as an aggregateable object and generate a collaborative construction aggregation identifier set.
[0038] The fusion structure generation module is used to execute S5: based on the path information recorded in the collaborative construction aggregation identifier set, extract representative semantic fragments, spatial description ranges and construction tag content from the associated construction paths, extract standard semantic content items after comparison and filtering, and establish a fusion version structure by combining the construction source subject and path identifier information to generate a construction fusion structure unit.
[0039] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0040] In this invention, by structurally extracting and classifying lithology, hydrology, and spatial information through subject numbering, the efficiency of orderly organization and management of multi-source data can be improved. By comparing keywords and boundaries to identify semantic differences, the ability to identify and quantify conflicts is enhanced. By combining the changing trends of terminology frequency, applicability, and consistency, aggregateable content can be identified, representative expressions can be extracted and integrated to construct a unified structure, thereby strengthening semantic standardization and expression consistency, and realizing dynamic collaborative management of heterogeneous geological data from identification to integration. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a schematic diagram of the steps of the present invention;
[0043] Figure 2 This is a detailed schematic diagram of S1 of the present invention;
[0044] Figure 3 This is a detailed schematic diagram of S2 of the present invention;
[0045] Figure 4 This is a detailed schematic diagram of S3 of the present invention;
[0046] Figure 5 This is a detailed schematic diagram of S4 of the present invention;
[0047] Figure 6 This is a detailed schematic diagram of S5 of the present invention;
[0048] Figure 7 This is a system module diagram of the present invention. Detailed Implementation
[0049] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0050] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0051] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0052] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0053] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0054] Please see Figure 1 This invention provides a knowledge collaborative construction method in the field of hydrogeology and environmental geology based on knowledge fusion, including the following steps:
[0055] S1: Obtain engineering data generated in landslide hazard assessment zones and groundwater dynamic monitoring sections, extract lithological descriptions, spatial boundary information and hydrological condition descriptions submitted by participating entities, group semantic fragments, engineering attribute items and applicability items, organize them in a structured manner according to the entity number, and generate a set of geological knowledge fragments;
[0056] S2: Based on the construction content of the main body of the geological knowledge fragment set regarding the division unit of the same surface water-groundwater conversion zone, extract keyword fragments, spatial description range and semantic tag information, identify the degree of overlap of terms, the intersection range of spatial description boundaries and the differences between the content pointed to by tags by comparison, integrate the differences, and generate semantic deviation parameter groups.
[0057] S3: Based on the difference information in the semantic deviation parameter group, identify the construction content with semantic conflict, split the segments with obvious divergence from the original main trajectory, and generate the current construction path according to the source subject, semantic label type and construction scope information, construct multiple branch trajectories with related identifiers, and generate a set of semantic divergence construction trajectories;
[0058] S4: Based on the semantic divergence, construct the semantic revision trajectory of the same rock mass structure stability control point in the trajectory set. Analyze the changing trends of the term usage frequency, applicability coverage and semantic consistency in the construction path. If the features show a gradual convergence, mark the corresponding construction path as an aggregable object and generate a collaborative construction aggregation identifier set.
[0059] S5: Based on the path information recorded in the collaborative construction aggregation identifier, extract representative semantic fragments, spatial description ranges and construction tag content from the associated construction paths, extract standard semantic content items after comparison and filtering, and establish a fusion version structure by combining the construction source subject and path identifier information to generate a construction fusion structure unit.
[0060] The geological knowledge fragment set includes lithological semantic fragments, spatial boundary information items, hydrological condition description items, engineering attribute items, applicability item sets, and subject number index items; the semantic deviation parameter set includes term overlap index, spatial boundary intersection range, semantic label difference items, difference performance record items, and deviation type identifier items; the semantic divergence construction trajectory set includes semantic conflict fragment units, branch trajectory identifier items, source subject number identifier items, semantic label type index items, construction range constraint items, and trajectory association marker items; the collaborative construction aggregation identifier set includes term usage frequency sequence, applicability coverage matrix items, semantic consistency trend items, aggregateable object list items, and aggregation judgment marker items; the construction fusion structure unit includes standard semantic content items, representative spatial range template items, construction label specification items, contributing subject weight table items, path version identifier items, and fusion version structure framework items.
[0061] Please see Figure 2 The specific steps of S1 are as follows:
[0062] S101: Obtain engineering data from landslide hazard assessment zones and groundwater dynamic monitoring sections, extract lithological descriptions, hydrological conditions and spatial boundary coordinates corresponding to the participating entity numbers, index and organize lithological particle size, water level change range and section coordinates according to the entity number, and generate a set of lithological and hydrological boundary indicators.
[0063] Engineering data packages are retrieved from the assessment zoning database and the monitoring section database. The fields in the data packages are in a fixed order: subject number, lithological description, hydrological condition information, and spatial boundary coordinates. The subject numbers are then checked, comparing each one against the registration list. Any missing or duplicate records are removed. For the retained records, lithological grain size quantification is performed, converting particle size distribution terms in the lithological description into grain size ranges. These ranges are mapped to the grain size boundaries and written to the index field. In the example, subject 02 describes sand, with a grain size range of 0.050 mm to 2.000 mm; subject 03 describes gravel, with a grain size range of 2.000 mm to 20.000 mm. Next, the hydrological condition information is processed. First, null values are removed from the daily sampling sequences of the observation wells, removing periods with two consecutive missing sampling intervals. Then, median filtering is applied with a window length of 7 to remove daily spikes. After filtering, the minimum and maximum water levels within the statistical period are used to form the water level variation range. In the example, Well 02 recorded a minimum water level of 7.60 meters and a maximum water level of 8.40 meters within 30 days, with the recorded interval being 7.60 meters to 8.40 meters. Next, the spatial boundary coordinates of the cross-section were processed. The coordinate points were read in closed sequence, and the closure error was calculated. This was verified by the planar distance between the start and end points, with an allowable error set at 0.020 meters. This setting was based on the fact that the closure errors of three repeated measurements at the same cross-section (0.012 meters, 0.016 meters, and 0.018 meters) were all less than 0.020 meters. Finally, an index row was created using the main body number. The grain size interval, water level change interval, and cross-section coordinate set were written into the same row and sorted by cross-section number to obtain the lithological and hydrological boundary index set.
[0064] S102: Based on the lithological and hydrological boundary index set, extract the structural hierarchy sequence and spatial boundary combination form, classify them according to the stratigraphic order and boundary extension characteristics, divide the content with the same combination form into semantic fragment types, establish a corresponding relationship with the main body number, and generate a semantic fragment index sequence.
[0065] The lithological and hydrological boundary index set is read, and the structural hierarchy sequence is extracted for each main record. The lithological description is first split into segments, and then a stratigraphic sequence list is generated based on their order of appearance. Consistency checks are performed during stratigraphic sequence writing, comparing stratigraphic terms with the cross-section reference table item by item; if a match is found, it is added to the list. Subsequently, the spatial boundary combination form is extracted, and the coordinate set of each boundary layer is projected onto the cross-section baseline direction to obtain the lateral coverage area. Simultaneously, the elevation points of the layer boundaries are read to obtain the elevation range. Next, the boundary extension characteristics are classified. First, the lateral length of the cross-section is divided into segments at 1.000-meter intervals. The proportion of boundary coverage segments to the total number of segments is calculated to obtain the continuity ratio. Then, the number of inflection points on the broken lines is counted; the inflection point criterion is that the angle between adjacent line segments is less than 150.000 degrees. Straight-line characteristics with a continuity ratio greater than or equal to 0.80 and fewer than or equal to 2 inflection points are classified as linear; characteristics with a continuity ratio between 0.50 and 0.80 and between 3 and 6 inflection points are classified as folded-line characteristics; and the rest are classified as fragmented-line characteristics. In the example, the silt layer has 37 segments and a total of 50 segments, with a continuity ratio of 0.74 and 4 inflection points, thus classifying it as folded-line characteristics. Then, the stratigraphic sequence and extension type are concatenated into a combined record, and all main records are classified according to the consistency of the combined form, grouping completely consistent combined forms into the same semantic fragment type. Fragment type numbers are generated incrementally according to the order of their first appearance, and a one-to-one correspondence is established between the main record number and the fragment type number, outputting a semantic fragment index sequence.
[0066] S103: Based on the semantic fragment index sequence, combined with the spatial boundary coordinates and water level change range, classify the distribution range of semantic fragments within the monitoring section, construct engineering attribute items and applicability items based on the boundary range and hydrological grouping relationship, and generate a geological knowledge fragment set;
[0067] The semantic fragment index sequence is read, and the spatial boundary coordinate set and water level change interval are retrieved back by fragment type number to classify the distribution range of semantic fragments within the cross section. First, the cross section boundary and fragment boundary are converted into polygonal regions. Then, the cross section polygons are rasterized into 0.500-meter grids. The total number of grids in the cross section and the number of grids covered by fragments are counted to obtain the coverage area percentage. In the example cross section, there are 2000 total grids, 860 fragment-covered grids, and a coverage area percentage of 0.43. Next, the water level change interval is quantified. The difference between the upper and lower boundaries of the interval is used to obtain the water level fluctuation amplitude, which is then grouped according to thresholds: amplitudes less than 0.50 meters are classified as micro-amplitude, amplitudes between 0.50 and 1.50 meters as medium-amplitude, and amplitudes greater than 1.50 meters as strong-amplitude. The thresholds are derived from the fluctuation amplitude statistics of 20 wells over 90 consecutive days, with the 30th percentile at 0.48 meters and the 70th percentile at 1.52 meters, rounded to 0.50 meters and 1.50 meters respectively. In the example, the water level range is 7.60 meters to 8.40 meters, with a range of 0.80 meters, and is classified as a medium-span group. Then, the coverage area percentage is divided into intervals of 0.00 to 0.20, 0.20 to 0.50, and 0.50 to 1.00. Percentages falling within these intervals are combined with hydrological groupings to generate engineering attribute items. Next, applicability entries are generated based on the engineering attribute items. Each entry field records the fragment type number, coverage interval, hydrological grouping, and source subject number. Finally, the engineering attribute items and applicability entries are aggregated by fragment type number to output a geological knowledge fragment set.
[0068] Please see Figure 3 The specific steps of S2 are as follows:
[0069] S201: Based on the construction content of the main body in the geological knowledge fragment set regarding the division unit of the same surface water-groundwater conversion zone, extract keyword fragments, spatial description range and semantic tag information, call the stratigraphic name in the keyword fragment and the coordinate boundary in the spatial description range, sort out the construction content under the main body number, and generate a semantic construction feature set;
[0070] The geological knowledge fragment set aggregates construction content records from different subjects according to construction unit number. For each record, keyword fragments, spatial description ranges, and semantic tags are extracted. Keyword fragment extraction first segments by punctuation, then removes stop words from each segment. The stop word list is stored permanently and maintained according to commonly used function words in engineering writing. Next, stratigraphic name recognition is performed in the term sequence, comparing each term with the stratigraphic name in the cross-section reference table. If a match is found, it is written to the stratigraphic keyword field. Simultaneously, lithological terms and boundary indicator terms are extracted and written to the keyword fragment set. Spatial description range extraction parses coordinate segments, splitting coordinate pairs into point sequences by commas. The number of points is checked to be no less than 4 and the closure error to be less than 0.020 meters. If these conditions are met, the points are written to the coordinate set field. Semantic tag extraction matches tag terms with the tag dictionary item by item. If a match is found, it is written to the semantic tag set. Then, construction content under the same subject number is organized by category, using the stratigraphic name as the category key to aggregate paragraphs with the same stratigraphic name. Finally, using the coordinate set as the constraint key, the paragraphs are assigned to the corresponding spatial range entries. In the example, Unit 01 identifies two types of strata: silt and sand. The silt layer corresponds to the area from point 01 to point 04, and the sand layer corresponds to the area from point 05 to point 08, forming two sub-records. Finally, the subject number, construction unit number, keyword fragment, coordinate set, and semantic label are written into the same structure and sorted by subject number to output the semantic construction feature set.
[0071] S202: Based on the semantic feature set, call the keyword fragments and the content pointed to by the semantic tags, determine the degree of word repetition in the same building unit, divide the word consistency level according to the repetition rate, and perform difference classification processing in combination with the mapping relationship between the content pointed to by the semantic tags to generate a semantic overlap coefficient.
[0072] The semantic feature set is read, and the degree of term repetition is judged for subject pairs within the same construction unit. First, the keyword fragment set of each subject is deduplicated to obtain a deduplicated term set. Then, the number of common terms and the number of union terms are counted for subject pairs. Subsequently, the term repetition rate is obtained by dividing the number of common terms by the number of union terms and kept to 0.01. In the example, subject 01 has 21 deduplicated terms, subject 02 has 19 deduplicated terms, 15 common terms, and 25 union terms, with a term repetition rate of 0.60. Next, consistency levels are divided according to the repetition rate. The threshold is determined by the manual annotation statistics of 40 subject pairs. The median of the adjacent level boundaries forms three threshold lines: 0.75, 0.55, and 0.35. A repetition rate greater than or equal to 0.75 is classified as level 1, between 0.55 and 0.75 as level 2, between 0.35 and 0.55 as level 3, and less than 0.35 as level 4. In the example, 0.60 is classified as level 2. Next, semantic tag mapping difference classification is performed. First, the number of common tag names is counted and compared with the total number of tags to obtain the tag overlap ratio. Then, the paragraph numbers of the content pointed to by the tags are checked. If the tag names are the same but the paragraph numbers they point to are different, it is marked as a pointing difference. Finally, a semantic overlap coefficient is generated. The calculation is performed by weighted summation, with a weight of 0.60 for the term repetition rate and 0.40 for the tag overlap ratio. The weights are derived from accuracy experiments under different combinations. The combination of 0.60 and 0.40 has an accuracy of 0.86. In the example, the term repetition rate is 0.60, the tag overlap ratio is 0.55, and the semantic overlap coefficient is 0.58.
[0073] S203: For the semantic overlap coefficient, combined with the set of boundary coordinates in the spatial description range, determine the number of intersection segments and the overlap ratio of spatial boundaries within the same building unit, call the combination relationship between the number of intersections and the semantic overlap coefficient value, merge and express the degree of difference, and generate a semantic deviation parameter set.
[0074] First, the coordinate sets of the two main entities are converted into polygonal regions, and the intersection of the two polygons is calculated to obtain the intersection region. Then, the intersection region is projected along the cross-sectional baseline, and the projection result is decomposed into continuous intervals and the number of intervals is counted to obtain the number of intersection segments. The spatial overlap ratio is obtained by dividing the area of the intersection region by the smaller of the areas of the two main entities. In the example, the number of intersection segments is 3, and the spatial overlap ratio is 0.66. Subsequently, the semantic overlap coefficient and the spatial overlap ratio are converted into difference scores. The conversion rule applies 1.00 to each ratio to obtain the difference score. In the example, the semantic overlap coefficient of 0.58 corresponds to a semantic difference score of 0.42, and the spatial overlap ratio of 0.66 corresponds to a spatial difference score of 0.34. Next, the number of intersection segments is categorized: 1 segment is categorized as a low segment, 2 to 3 segments as a medium segment, and 4 or more segments as a high segment. In the example, 3 segments are categorized as a medium segment. Finally, a semantic bias parameter set is generated. The semantic difference score and spatial difference score are weighted and summed with a weighting of 0.50 and 0.50 respectively to obtain the overall bias value. The weights are determined based on a recall-precision balance experiment, and the balanced combination keeps both above 0.80. In the example, the overall bias value is 0.38, and it is written into the semantic bias parameter set along with the middle segment.
[0075] Please see Figure 4 The specific steps of S3 are as follows:
[0076] S301: Based on the difference information in the semantic deviation parameter group, call the subject number, semantic label and construction range boundary in the construction content to compare and identify the semantic representation under the same transformation zone construction unit, filter out the fragment content with differences exceeding the identification threshold, and establish a semantic conflict fragment index set;
[0077] The semantic deviation parameter group, including the subject ID pairs, semantic tags, and construction boundaries, is read. Semantic representations within the same construction unit are compared, identified, and conflicting fragments are filtered. First, using the fragment type ID as the alignment key, keyword fragment sequences under the same fragment type ID are aligned sequentially. After alignment, the number of inconsistent terms is counted item by item, and then the number of inconsistencies is divided by the total number of alignable terms to obtain the term difference score. The same statistical analysis is performed on the semantic tags to obtain the tag difference score. Then, the comprehensive deviation value, term difference score, and tag difference score are combined into a fragment difference amplitude. The combination is performed using a weighted summation, with a weight of 0.40 for the comprehensive deviation value, 0.40 for the term difference score, and 0.20 for the tag difference score. These weights are derived from a validation set of 60 known conflicting fragments, achieving an accuracy of 0.88. Then, an identification threshold is set and filtering is performed. The threshold is determined through scanning experiments with a step size of 0.05, ranging from 0.20 to 0.70. A harmonic mean of 0.84 is reached when the threshold is set to 0.45. In the example, fragment type 05 has a comprehensive deviation value of 0.52, a term difference score of 0.55, and a tag difference score of 0.40. After weighted calculation, the fragment difference amplitude is 0.51, which exceeds 0.45. Therefore, fragment type 05 is written into the semantic conflict fragment index set, and the subject number pair, semantic tag, and construction boundary coordinate set are recorded simultaneously. The index set is archived according to the construction unit number.
[0078] S302: Based on the semantic conflict fragment index set, extract the corresponding fragment position in the original main trajectory, call the fragment start and end positions and corresponding numbers to perform splitting operation, peel the marked fragments from the main trajectory, and record the source subject, semantic label type and construction boundary range to generate split fragment trajectory parameter group;
[0079] The system reads the semantic conflict fragment index set, locates the start and end positions of fragments in the original main trajectory for each index, and performs splitting. The main trajectory is managed using both the constructed content text and paragraph number sequences. During positioning, fragment type numbers are first mapped to paragraph numbers, and then keyword fragments are used to locate the start and end word positions within the target paragraph. The start position is taken from the word number of the first occurrence of the starting keyword, and the end position is taken from the word number of the last occurrence of the ending keyword. The word numbers are then converted into character start and end positions. In the example, main body 02 is located in construction unit 01, fragment type 05, and the character start and end positions are from position 1240 to 1685. Subsequently, a splitting operation is performed, the main trajectory text is truncated, the text of this interval is stripped and written to the split fragment record, and a placeholder is written in the original position. The placeholder is generated by concatenating the construction unit number and the fragment type number to ensure consistent backfill positioning. The split fragment record synchronously writes the source main number, semantic tag type, construction boundary coordinate set, and records the original paragraph number and batch timestamp. The timestamp is taken from the system time when the processing batch started to avoid confusion caused by multiple batch splitting. Finally, the output segment trajectory parameter group is sorted by main body number and segment number to ensure that subsequent branch path construction can be directly reassembled in order.
[0080] S303: Based on the segmented trajectory parameter group, using the main body number as the identification field, combined with semantic labels and boundary range, construct multiple trajectory paths with corresponding numbers, set the pointing identifiers between trajectories, generate a set of branch paths with associated number identifiers, and establish semantic divergence to construct the trajectory set;
[0081] The system reads the segment trajectory parameter group, merges the segmented data by subject number, and constructs a branch path set. First, segments under the same subject number are sorted in ascending order by their original segment numbers to form a segmentation sequence. Then, candidate backfill segments are generated for each placeholder position. The generation rule for candidate backfill segments is constrained by the construction unit number. Segments from different subjects under the same construction unit are searched; segments with the same semantic label are included in the candidate set, while those with different semantic labels are kept separately as a difference candidate set. The candidate set is then sorted, using the segment difference magnitude obtained from S301 as the sorting key, arranged in ascending order. Next, a pointing identifier is generated, and the preceding and following placeholder identifiers are written as directed association records, with the candidate segment number list for that position written into the record field. Then, multiple paths are constructed. For each placeholder position, different paths are selected sequentially according to the candidate sorting, with path numbers generated in ascending order. The candidate segment number selected for each placeholder position is recorded. In the example, placeholder identifier unit 01 has 3 candidate segments for segment 05. Therefore, path 01 selects candidate segment 23, path 02 selects candidate segment 31, and path 03 selects candidate segment 42. The remaining positions share the same text sequence and the same pointing identifier. Finally, the semantic divergence trajectory set is output. The trajectory set fields retain the main number, path number, pointing identifier set, and backfill selection sequence.
[0082] Please see Figure 5 The specific steps of S4 are as follows:
[0083] S401: Based on semantic divergence, construct semantic revision trajectories for the same rock mass structure stability control point in the trajectory set, extract term record sequences in the construction path, call the number of times terms appear in each path, calculate the term usage frequency distribution value in the path, and generate term usage frequency coefficients;
[0084] The semantic divergence construction trajectory under the same control point is read, and term record sequences are extracted for each path, and term usage frequency coefficients are calculated. During term extraction, the path text is segmented by periods, and then matched against matching terms in an engineering terminology table. Matching terms are written into the term record sequence, along with sentence numbers and character position ranges. The total number of occurrences of each term for that path is then counted, followed by the frequency of each term. The frequency value is obtained by dividing the term's occurrence count by the total number of occurrences and rounded down to 0.01. In the example, path 01 has a total term occurrence count of 50, with seepage channels appearing 8 times, rock mass structural surfaces appearing 12 times, and aquitards appearing 5 times, corresponding to frequency values of 0.16, 0.24, and 0.10, respectively. Subsequently, term usage frequency coefficients are generated to describe the concentration of frequency distribution. During calculation, the frequency value sequence is sorted, and the maximum and second-largest frequency values are summed. Concentration levels are categorized using thresholds of 0.40 and 0.25. Values greater than or equal to 0.40 are considered high concentration, values between 0.25 and 0.40 are considered medium concentration, and values less than 0.25 are considered low concentration. The thresholds are determined statistically based on 18 historical revision paths, using the 70th percentile (0.41) and 30th percentile (0.26) of the largest and second-largest frequency values, respectively, to fix 0.40 and 0.25. In the example, the sum of the largest (0.24) and the second-largest (0.16) is 0.40, which is considered high concentration. The frequency value sequence and categorization results are then written into the output.
[0085] S402: Based on the term usage frequency coefficient, combined with the scope of applicability tags and semantic revisions in the constructed path, determine the applicability coverage ratio of the path within the same control point, extract the degree of association between tags and path paragraphs, and obtain the semantic coverage consistency trend.
[0086] The process involves reading terminology usage frequency coefficients, applicability label ranges, and revision content to calculate the applicability coverage ratio for the same control point and generate a coverage consistency trend. First, the control point boundary range is rasterized to a 0.500-meter grid to obtain the total number of control point grids. Then, the label names in the path text are located and mapped to the pointing segment type number. The segment boundary coordinate set is used to generate the segment region and mapped to the grid. The number of grids covered by the tag-pointed segment is counted. The coverage ratio is obtained by dividing the number of covered grids by the total number of grids. In the example, the total number of control point grids is 1200, and the path 01 label-pointed segment covers 780 grids, resulting in a coverage ratio of 0.65. Next, the correlation between the label and the path segment is calculated, measured by the number of sentence intervals. A strong correlation is defined as the interval between the first occurrence of a period in the label and the first occurrence of a period in the pointed segment being less than or equal to 1; a medium correlation is defined as 2 to 3; and a weak correlation is defined as 4 or greater. In the example, the interval is 2 sentences, classifying it as a medium correlation. Next, the coverage ratio sequences of multiple paths under the same control point are summarized. The difference between the maximum and minimum coverage ratios is calculated and categorized according to a threshold. A difference less than or equal to 0.10 is considered coverage convergence, between 0.10 and 0.25 is considered coverage differentiation, and greater than or equal to 0.25 is considered coverage dispersion. The thresholds are derived from historical versions of 12 control points, with a median difference of 0.12 and an upper quartile distance of 0.24, thus fixing 0.10 and 0.25. In the example, the coverage ratios are 0.65, 0.60, and 0.58, with a difference of 0.07, which is considered coverage convergence. The coverage ratio, association level, and coverage trend are then output.
[0087] S403: In response to the semantic coverage consistency trend, identify the range of changes in term usage frequency and the direction of changes in coverage ratio, determine whether paths show a gradual convergence during the revision process, filter the set of path numbers that reach the aggregation judgment threshold, and establish a collaborative construction aggregation identifier set;
[0088] The system reads the coverage consistency trend and term frequency concentration, performs path convergence discrimination, and filters aggregation identifiers. First, it calculates the term frequency variation range. For multiple paths under the same control point, it calculates the difference between the largest and second-largest frequency values. A difference less than or equal to 0.05 is considered frequency stable, between 0.05 and 0.12 is considered frequency changing, and greater than or equal to 0.12 is considered frequency fluctuating. The coverage direction is determined using the coverage convergence, coverage differentiation, and coverage dispersion labels obtained from S402. Convergence discrimination follows fixed rules: coverage convergence with stable frequencies is considered convergence; for coverage differentiation, the association level difference is checked to ensure it does not exceed one level before being considered convergence; coverage dispersion is directly marked as non-convergence. Subsequently, it calculates a comprehensive score and filters the path number set. The comprehensive score is composed of the average coverage ratio, the coverage ratio difference, and the frequency difference. The calculation first assigns a score to the average coverage ratio, then deducts points from the coverage ratio difference and frequency difference, with the deduction increasing as the difference increases. Finally, the scores are combined to obtain the comprehensive score. The aggregation decision threshold is determined by a threshold scan based on the manual confirmation results of six control points. The threshold is set in increments of 0.05 within the range of 0.60 to 0.85, with the highest manual consistency rate of 0.90 achieved at a threshold of 0.75. In the example, the average coverage ratio is 0.61, the coverage ratio difference is 0.07, the frequency difference is 0.03, and the merged comprehensive score is 0.79. Since the score exceeds 0.75, paths 01, 02, and 03 are written into the collaborative construction aggregation identifier set, and the control point number, path number set, and score details are recorded.
[0089] Please see Figure 6 The specific steps of S5 are as follows:
[0090] S501: Based on the path information recorded in the collaborative construction aggregation identifier, extract the semantic fragments, spatial description range and construction tag content under the corresponding construction path, identify the completeness of term expression, boundary extension method and tag usage position in the path, filter the content that is representative of the path, and generate a representative item group of the aggregation path.
[0091] The collaborative construction aggregation identifier set is read, and semantic fragments, spatial description ranges, and constructed label content are extracted according to the selected paths, and representative items are selected. First, placeholder identifiers and backfill selection sequences are used to locate backfill fragments, and fragment text, fragment type number, and source entity number are extracted to form candidate items. The spatial description range directly references the boundary coordinate set corresponding to the fragment type number, and verifies that the closure error is less than 0.020 meters and the number of points is no less than 4. The constructed label content is obtained by finding the label name in the path text and recording its sentence number to form a label position sequence. Then, the terminology completeness is calculated, with the completeness taken as the percentage of essential terms hit. The list of essential terms is provided by the control point establishment records. In the example, control point 01 has 3 essential terms, path 01 hits 3 terms (completeness 1.00), and path 02 hits 2 terms (completeness 0.67). The boundary extension method calls the classification results of S102. The label's position distribution value is calculated using location. The total number of sentences in the path is used as a baseline. Each label sentence number is divided by the total number of sentences to obtain a relative position value. The distribution of relative position values falling into three segments—0.00 to 0.33, 0.33 to 0.67, and 0.67 to 1.00—is then analyzed. Representative item selection follows fixed conditions: terminology completeness greater than or equal to 0.80, the boundary extension method is consistent with the majority type of the selected path set, and the proportion of the first and middle segments of the label position distribution is greater than or equal to 0.60. In the example, path 01 has a completeness of 1.00, is a majority type (folded type), and the proportion of the first and middle segments is 0.70. Therefore, path 01 is included in the aggregated path representative item group.
[0092] S502: Based on the aggregated path representative item group, the semantic fragment expression form, spatial description boundary range and tag setting field are compared item by item to determine the repetition of term structure, the intersection ratio of boundary coordinates and the consistency of tag association number, and content fragments with matching relationship are filtered to obtain standard semantic content items;
[0093] First, the repetition rate of the semantic segment expression is calculated. Each representative item is segmented and deduplicated. Then, the number of common terms and the number of terms in the union of representative items are counted. The repetition ratio is obtained by dividing the number of common terms by the number of terms in the union. The repetition range is fixed as follows: high repetition 0.70 to 1.00, medium repetition 0.50 to 0.70, and low repetition less than 0.50. The threshold is based on statistics of 24 manually confirmed synonymous segments, with a confirmation rate of 0.92 for the high repetition range. Next, the intersection ratio is calculated for the spatial boundaries. The intersection area of the two boundary polygons is obtained, and the intersection ratio is obtained by dividing the intersection area by the smaller of the two polygon areas. An intersection ratio greater than or equal to 0.60 is considered a boundary match, between 0.40 and 0.60 is considered a boundary alignment, and less than 0.40 is considered a boundary mismatch. The thresholds are based on statistical analysis of eight cross-sections. The minimum intersection ratio for repeated plotting within the same range is 0.62, and the maximum intersection ratio for adjacent ranges is 0.58. Therefore, thresholds of 0.60 and 0.40 are fixed. Next, the consistency of tag association numbers is judged. Tag names are mapped to tag dictionary numbers, and the ratio of the number of intersections to the number of unions of tag numbers is calculated. A ratio greater than or equal to 0.70 is considered consistent. The filtering conditions are then used for parallel judgment. If the repetition falls into the medium or high repetition range, the boundary judgment is boundary matching, and the tag consistency is consistent. When the conditions are met, a common expression skeleton is extracted. Skeleton extraction is performed according to the longest common term sequence, and the consistent boundary coordinate set and tag number set are written into the standard item. In the example, path 01 and path 03 have a repetition of 0.72, an intersection ratio of 0.66, and a tag consistency of 0.75, generating the standard semantic content item 01.
[0094] S503: Based on the standard semantic content items, call the source subject number and corresponding path identification information to uniformly identify the corresponding position of the content items in the path, and establish a content structure combination sequence with the set of numbers as the index field to generate the construction fusion structure unit;
[0095] The system reads standard semantic content items, applies unified identification to selected paths, and generates fusion structure units. First, it locates the backfill position for each path using placeholder identifiers. For each position, it matches and verifies the fragment type number with the boundary coordinate set. The matching condition uses a dual threshold: a term structure repetition rate of 0.70 or higher and a boundary intersection ratio of 0.60 or higher are considered a match. If a match is found, the placeholder identifier at that position is replaced with the standard semantic content item number identifier, and this is written to the path identifier log. The log fields record the control point number, path number, original placeholder identifier, standard content item number, repetition rate, and intersection ratio. Then, a content structure combination sequence is established. The standard content item numbers for each path under the same control point are arranged in the order they appear in the path text to form a number sequence. Multiple path number sequences are then merged and deduplicated. The sorting criterion is the path with the highest comprehensive score. When order conflicts occur, an order difference entry is written, and both numbering orders for the conflicting positions are retained. Finally, the fusion structure unit is generated, and the output fields include the control point number, combination sequence, boundary coordinate set corresponding to each standard content item, tag number set, source subject number set, and path number set. The test data under the aggregation judgment threshold of 0.75 is written into the control description. The aggregation consistency rate of the combination with threshold of 0.75 is 0.90, which is higher than that of the combination with threshold of 0.70 (0.83). The number of order difference items is reduced from 12 to 5, a decrease of 58.33%.
[0096] Please see Figure 7 A knowledge collaborative construction system for the field of hydrogeology and environmental geology based on knowledge fusion includes:
[0097] The knowledge fragment extraction module is used to perform S1: obtain engineering data formed in landslide hazard assessment zones and groundwater dynamic monitoring sections, extract lithological descriptions, spatial boundary information and hydrological condition descriptions submitted by participating entities, group semantic fragments, engineering attribute items and applicability items, and organize them in a structured manner according to the entity number to generate a set of geological knowledge fragments;
[0098] The difference parameter identification module is used to perform S2: Based on the construction content of the main body of the geological knowledge fragment set about the same surface water-groundwater conversion zone division unit, extract keyword fragments, spatial description range and semantic tag information, identify the degree of overlap of terms, the intersection range of spatial description boundaries and the differences between the content pointed to by tags through comparison, integrate the difference performance, and generate semantic deviation parameter group;
[0099] The divergence trajectory construction module is used to execute S3: based on the difference information in the semantic deviation parameter group, it identifies the construction content with semantic conflicts, splits the segments with obvious divergence from the original main trajectory, and generates the current construction path based on the source subject, semantic label type and construction scope information, constructs multiple branch trajectories with associated identifiers, and generates a set of semantic divergence construction trajectories;
[0100] The aggregation path judgment module is used to execute S4: based on the semantic divergence, construct the semantic revision trajectory of the same rock mass structure stability control point in the trajectory set, analyze the changing trend of the term usage frequency, applicability coverage and semantic consistency performance in the construction path, and if the features show a gradual convergence state, mark the corresponding construction path as an aggregateable object and generate a collaborative construction aggregation identifier set;
[0101] The fusion structure generation module is used to execute S5: based on the path information recorded in the collaborative construction aggregation identifier set, it extracts representative semantic fragments, spatial description ranges and construction tag content from the associated construction paths, extracts standard semantic content items after comparison and filtering, and establishes a fusion version structure by combining the construction source subject and path identifier information to generate a construction fusion structure unit.
[0102] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A knowledge collaborative construction method in the field of hydrogeology and environmental geology based on knowledge fusion, characterized in that: Includes the following steps: S1: Obtain engineering data from landslide hazard assessment zones and groundwater dynamic monitoring sections, extract lithological descriptions, spatial boundaries and hydrological conditions submitted by participating entities, construct semantic fragments, engineering attribute items and applicability entries, and group them in a structured manner according to entity number to form a set of geological knowledge fragments; S2: Extract keyword fragments, spatial descriptions and semantic tags from the geological knowledge fragment set for the same surface water-groundwater conversion zone division unit, and generate semantic deviation parameter sets by comparing term overlap, boundary intersection and tag differences; S3: Identify the construction content with semantic conflicts based on the semantic deviation parameter set, split the conflicting segments from the original trajectory, combine the source subject, semantic label and construction scope to generate the current path, and establish a semantic divergence construction trajectory set; S4: Construct a trajectory set based on the semantic divergence, perform trend analysis on term frequency, applicability coverage and semantic consistency, and mark objects that show convergence as aggregable objects to generate a collaborative construction aggregation identifier set. S5: Extract representative semantic fragments, spatial ranges and construction tags based on the collaborative construction aggregation identifier set, filter standard semantic content, and generate construction fusion structural units by combining source subject and path information.
2. The knowledge collaborative construction method in the field of hydrogeology and environmental geology based on knowledge fusion according to claim 1, characterized in that, The geological knowledge fragment set includes lithological semantic fragments, spatial boundary information items, hydrological condition description items, engineering attribute item set, applicability item set, and subject number index item; the semantic deviation parameter set includes term overlap index, spatial boundary intersection range, semantic tag difference item, difference performance record item, and deviation type identifier item; the semantic divergence construction trajectory set includes semantic conflict fragment unit, branch trajectory identifier item, source subject number identifier item, semantic tag type index item, construction range constraint item, and trajectory association marker item; the collaborative construction aggregation identifier set includes term usage frequency sequence, applicability coverage matrix item, semantic consistency trend item, aggregateable object list item, and aggregation judgment marker item; the construction fusion structure unit includes standard semantic content item, representative spatial range template item, construction tag specification item, contributing subject weight table item, path version identifier item, and fusion version structure framework item.
3. The knowledge collaborative construction method in the field of hydrogeology and environmental geology based on knowledge fusion according to claim 1, characterized in that, The specific steps of S1 are as follows: S101: Obtain engineering data from landslide hazard assessment zones and groundwater dynamic monitoring sections, extract lithological descriptions, hydrological conditions and spatial boundary coordinates corresponding to the participating entity numbers, index and organize lithological particle size, water level change range and section coordinates according to the entity number, and generate a set of lithological and hydrological boundary indicators. S102: Based on the lithological and hydrological boundary index set, extract the structural hierarchy sequence and spatial boundary combination form, classify them according to the stratigraphic order and boundary extension characteristics, divide the content with the same combination form into semantic fragment types, establish a corresponding relationship with the main number, and generate a semantic fragment index sequence. S103: Based on the semantic fragment index sequence, combined with the spatial boundary coordinates and water level change range, the distribution range of the semantic fragments within the monitoring section is classified, and engineering attribute items and applicability items are constructed according to the boundary range and hydrological grouping relationship to generate a geological knowledge fragment set.
4. The knowledge collaborative construction method in the field of hydrogeology and environmental geology based on knowledge fusion according to claim 3, characterized in that, The specific steps of S2 are as follows: S201: Based on the construction content of the main body in the geological knowledge fragment set regarding the same surface water-groundwater conversion zone division unit, extract keyword fragments, spatial description range and semantic tag information, call the stratigraphic name in the keyword fragment and the coordinate boundary in the spatial description range, sort out the construction content under the main body number, and generate a semantic construction feature set; S202: Based on the semantic construction feature set, call the keyword fragments and the content pointed to by the semantic tags, determine the degree of word repetition in the same construction unit, divide the word consistency level according to the repetition rate, and perform difference classification processing in combination with the mapping relationship between the content pointed to by the semantic tags to generate a semantic overlap coefficient. S203: For the semantic overlap coefficient, combined with the set of boundary coordinates in the spatial description range, determine the number of intersection segments and the overlap ratio of spatial boundaries within the same building unit, call the combination relationship between the number of intersections and the semantic overlap coefficient value, merge and express the degree of difference, and generate a semantic deviation parameter set.
5. The knowledge collaborative construction method in the field of hydrogeology and environmental geology based on knowledge fusion according to claim 4, characterized in that, In the process of generating the semantic construction feature set, the stratigraphic names extracted from the keyword fragments are subjected to standardized comparison processing. The standardized comparison processing is limited to the scope of the construction content corresponding to the same subject number, and the stratigraphic names with a matching degree lower than the preset comparison threshold are marked as inconsistent terms. The term consistency level is limited to no less than three discrete levels, and the calculation range of the repetition rate is limited to keyword fragments in the semantic construction feature set that belong to the same surface water-groundwater conversion zone division unit. In the process of generating the semantic overlap coefficient, the mapping relationship between the semantic tags pointing to the content is limited to a one-to-one or many-to-one mapping form, and only the semantic tags that are successfully mapped participate in the calculation of the semantic overlap coefficient. In the process of generating the semantic deviation parameter set, the boundary coordinate set is limited to a set of closed polygons under the same coordinate reference system, and is combined and expressed based on the interval division results of the number of intersection segments and the overlap ratio.
6. The knowledge collaborative construction method in the field of hydrogeology and environmental geology based on knowledge fusion according to claim 4, characterized in that, The specific steps for S3 are as follows: S301: Based on the difference information in the semantic deviation parameter group, call the subject number, semantic tag and construction range boundary in the construction content, compare and identify the semantic representation under the same conversion zone construction unit, filter out the fragment content with differences exceeding the identification threshold, and establish a semantic conflict fragment index set. S302: Based on the semantic conflict fragment index set, extract the corresponding fragment position in the original main body trajectory, call the fragment start and end positions and corresponding numbers to perform splitting operation, peel the marked fragments from the main body trajectory, and record the source main body, semantic tag type and construction boundary range to generate a split fragment trajectory parameter group; S303: Based on the segmented trajectory parameter group, using the main body number as the identification field, combined with semantic labels and boundary ranges, construct multiple trajectory paths with corresponding numbers, set the pointing identifiers between trajectories, generate a set of branch paths with associated number identifiers, and establish semantic divergence to construct the trajectory set.
7. The knowledge collaborative construction method in the field of hydrogeology and environmental geology based on knowledge fusion according to claim 6, characterized in that, During the establishment of the semantic conflict fragment index set, the difference magnitude exceeding the identification threshold is defined as the construction content whose corresponding value in the semantic deviation parameter group falls within the preset difference interval range, and the comparison identification is limited to the semantic expression range of construction units with the same subject number and belonging to the same surface water-groundwater conversion zone construction unit; in the splitting operation, the fragment start and end positions are defined as the sequence position identifiers continuously recorded in the original subject trajectory, and the splitting operation is defined as being performed only on fragments that completely cover the boundary of the construction range; the source subject, semantic label type and construction boundary range recorded in the split fragment trajectory parameter group are defined as corresponding one-to-one with the semantic conflict fragment index set; during the generation of the branch path set, the pointing identifier is defined as establishing a one-way association relationship between trajectory paths corresponding to different subject numbers, and the subject number is used as a unique index field for differentiation.
8. The knowledge collaborative construction method in the field of hydrogeology and environmental geology based on knowledge fusion according to claim 6, characterized in that, The specific steps of S4 are as follows: S401: Based on the semantic divergence, construct the semantic revision trajectory of the same rock mass structure stability control point in the trajectory set, extract the term record sequence in the construction path, call the number of times the term appears in each path, calculate the term usage frequency distribution value in the path, and generate the term usage frequency coefficient. S402: Based on the term usage frequency coefficient, combined with the applicability label range and semantic revision content in the constructed path, determine the applicability coverage ratio of the path within the same control point range, extract the degree of association between the label and the path paragraph, and obtain the semantic coverage consistency trend. S403: In response to the semantic coverage consistency trend, identify the range of changes in term usage frequency and the direction of changes in coverage ratio, determine whether paths show a gradual convergence during the revision process, filter the set of path numbers that reach the aggregation judgment threshold, and establish a collaborative construction aggregation identifier set.
9. The knowledge collaborative construction method in the field of hydrogeology and environmental geology based on knowledge fusion according to claim 8, characterized in that, The specific steps of S5 are as follows: S501: Based on the path information recorded in the collaborative construction aggregation identifier set, extract the semantic fragments, spatial description range and construction tag content under the corresponding construction path, identify the completeness of term expression, boundary extension method and tag usage position in the path, filter the content that is representative of the path, and generate a representative item group of the aggregation path. S502: Based on the aggregated path representative item group, the semantic fragment expression form, spatial description boundary range and tag setting field are compared item by item to determine the word structure repetition, boundary coordinate intersection ratio and tag association number consistency, filter content fragments with matching relationship and obtain standard semantic content items; S503: Based on the standard semantic content item, call the source subject number and corresponding path identification information to uniformly identify the corresponding position of the content item in the path, and establish a content structure combination sequence with the set of numbers as the index field to generate a construction fusion structure unit.
10. A knowledge collaborative construction system for the field of hydrogeology and environmental geology based on knowledge fusion, characterized in that: The system is used to implement the knowledge collaborative construction method in the field of hydrogeology and environmental geology based on knowledge fusion as described in any one of claims 1-9, and the system includes: The knowledge fragment extraction module is used to perform S1: obtain engineering data formed in landslide hazard assessment zones and groundwater dynamic monitoring sections, extract lithological descriptions, spatial boundary information and hydrological condition descriptions submitted by participating entities, group semantic fragments, engineering attribute items and applicability items, and organize them in a structured manner according to the entity number to generate a set of geological knowledge fragments; The difference parameter identification module is used to perform S2: based on the construction content of the main body of the geological knowledge fragment set about the same surface water-groundwater conversion zone division unit, extract keyword fragments, spatial description range and semantic tag information, identify the degree of overlap of terms, the intersection range of spatial description boundaries and the differences between the content pointed to by tags by comparison, integrate the difference performance, and generate a semantic deviation parameter group; The divergence trajectory construction module is used to execute S3: based on the difference information in the semantic deviation parameter group, identify the construction content with semantic conflict, split the segments with obvious divergence from the original main trajectory, and generate the current construction path based on the source subject, semantic label type and construction scope information, construct multiple branch trajectories with associated identifiers, and generate a set of semantic divergence construction trajectories; The aggregation path judgment module is used to execute S4: Based on the semantic revision trajectory of the same rock mass structure stability control point in the semantic divergence construction trajectory set, analyze the changing trend of the term usage frequency, applicability coverage and semantic consistency performance in the construction path. If the features show a gradual convergence state, mark the corresponding construction path as an aggregateable object and generate a collaborative construction aggregation identifier set. The fusion structure generation module is used to execute S5: based on the path information recorded in the collaborative construction aggregation identifier set, extract representative semantic fragments, spatial description ranges and construction tag content from the associated construction paths, extract standard semantic content items after comparison and filtering, and establish a fusion version structure by combining the construction source subject and path identifier information to generate a construction fusion structure unit.