Unconventional water source problem tracing method and system using knowledge graph

By using knowledge graphs to construct problem-related transmission chains in the tracing of unconventional water source issues and calculating the influence attenuation coefficient of semantic related edges, the problem of inaccurate tracing in existing technologies is solved, achieving efficient and accurate tracing results and improving the stability and efficiency of water source development and utilization.

CN122114942APending Publication Date: 2026-05-29INNER MONGOLIA AGRICULTURAL UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INNER MONGOLIA AGRICULTURAL UNIVERSITY
Filing Date
2026-01-13
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies rely on human experience and simple data analysis in tracing the source of unconventional water source problems, making it difficult to comprehensively and accurately identify the root causes of the problems. They also lack systematicity and standardization, resulting in inaccurate and in-depth tracing results.

Method used

A method for tracing the root causes of unconventional water source problems is constructed using knowledge graphs. By acquiring problem representation information and mapping it to a pre-built knowledge graph, a problem association transmission chain is constructed, the influence attenuation coefficient of semantic association edges is calculated, an attenuation coefficient matrix is ​​generated, effective transmission paths are screened, and the root cause of the problem is traced.

Benefits of technology

It improves the accuracy and efficiency of tracing the source of unconventional water source problems, enabling accurate identification of the root causes of problems from a holistic perspective, improving development and utilization efficiency and stability, and reducing operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114942A_ABST
    Figure CN122114942A_ABST
Patent Text Reader

Abstract

The application provides a non-conventional water source problem tracing method and system using a knowledge graph. First, the problem representation information of the non-conventional water source is obtained and mapped to a pre-constructed non-conventional water source knowledge graph containing water source feature units, process feature units, equipment feature units and environmental feature units. Based on the mapping result, a problem association conduction chain is constructed, and the feature units and association types in the extension process are recorded. The influence attenuation coefficient of each semantic association edge in the problem association conduction chain is calculated, and an attenuation coefficient matrix is generated. According to the attenuation coefficient matrix, the effective association conduction path is obtained through path screening. Finally, based on the effective path, the starting feature unit is tracked, the information is integrated to generate a tracing result, and the root cause of the non-conventional water source problem can be accurately and efficiently traced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of knowledge graph technology, and more specifically, to a method and system for tracing the source of unconventional water source problems using knowledge graphs. Background Technology

[0002] In the development and utilization of unconventional water sources, various problems often arise due to their complex characteristics, diverse development processes, and significant influence from environmental factors. These problems include declining water quality, obstructed development processes, and equipment malfunctions. Accurately and quickly tracing the root causes of these problems is crucial for ensuring a stable supply and efficient utilization of unconventional water sources.

[0003] Currently, source tracing for unconventional water source problems mainly relies on human experience and simple data analysis methods. Human experience is often limited by individual knowledge and experience, making it difficult to comprehensively and accurately identify the root cause of the problem, especially when dealing with complex issues involving multiple stages, which can easily lead to misjudgments or omissions. Simple data analysis methods typically only process single types of data, lacking the ability to explore and analyze the relationships between different types of data, and failing to grasp the overall mechanism of problem generation and transmission, resulting in inaccurate and superficial source tracing results. Furthermore, existing source tracing methods lack systematicity and standardization, making it difficult to establish standardized operating procedures, which hinders large-scale application. Summary of the Invention

[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a method for tracing the source of unconventional water source problems using knowledge graphs, the method comprising: Obtain problem representation information of unconventional water sources, and map the problem representation information to a pre-constructed unconventional water source knowledge graph. The problem representation information includes abnormal water source quality, abnormal development process, abnormal equipment operation, and abnormal environmental correlation. The unconventional water source knowledge graph includes water source feature units, process feature units, equipment feature units, and environmental feature units, and each feature unit is connected by semantic association edges. Based on the mapping results between the problem representation information and the unconventional water source knowledge graph, a problem association transmission chain is constructed. The problem association transmission chain starts from the feature unit corresponding to the mapping and extends along the semantic association edge to other feature units in the knowledge graph. The association types of the feature units and semantic association edges passed through during the extension process are recorded. Calculate the influence attenuation coefficient of each semantic association edge in the problem association transmission chain, and generate an attenuation coefficient matrix containing the influence attenuation coefficient of each semantic association edge. The influence attenuation coefficient is determined based on the transmission strength corresponding to the association type and the semantic distance between feature units. Based on the attenuation coefficient matrix, the problem-related transmission chain is filtered to retain the transmission paths whose attenuation coefficients meet the transmission requirements, and to remove the transmission paths whose attenuation coefficients do not meet the transmission requirements, thus obtaining the effective related transmission paths. Based on the effective correlation transmission path, the starting feature unit is traced back to the effective correlation transmission path and the information of the root cause feature unit to generate the source tracing result of the unconventional water source problem. The starting feature unit is the root cause feature unit of the problem.

[0005] Furthermore, embodiments of the present invention also provide a source tracing system for unconventional water source problems utilizing knowledge graphs, characterized by comprising: A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to execute the aforementioned unconventional water source tracing method utilizing knowledge graphs by executing the machine-executable instructions.

[0006] In another aspect, embodiments of the present invention also provide a computer program product, the computer program product including machine-executable instructions, the machine-executable instructions being stored in a computer-readable storage medium, a processor of a computer device reading the machine-executable instructions from the computer-readable storage medium, the processor executing the machine-executable instructions, causing the computer device to execute the above-described unconventional water source problem tracing method using knowledge graphs.

[0007] Based on the above, by acquiring the problem representation information of unconventional water sources and mapping it to a pre-constructed unconventional water source knowledge graph, and utilizing the semantic relationships between various feature units in the knowledge graph, a problem-related transmission chain can be constructed. This effectively sorts out the transmission path of the problem in the knowledge graph. The influence attenuation coefficient of each semantic association edge in the problem-related transmission chain is calculated and an attenuation coefficient matrix is ​​generated. Based on this attenuation coefficient matrix, the transmission paths are filtered, retaining those whose influence attenuation coefficients meet the transmission requirements and eliminating invalid paths. This effectively improves the accuracy and efficiency of source tracing and avoids interference from invalid paths. Finally, based on the effective associated transmission paths, the starting feature unit is traced, and relevant information is integrated to generate source tracing results. This can accurately identify the root cause of the problem from a holistic perspective, helping to improve the development and utilization efficiency and stability of unconventional water sources and reduce operation and maintenance costs. Attached Figure Description

[0008] Figure 1 This is a schematic diagram of the execution flow of the unconventional water source problem tracing method using knowledge graphs provided in an embodiment of the present invention.

[0009] Figure 2 This is a schematic diagram of exemplary hardware and software components of an unconventional water source problem tracing system utilizing knowledge graphs provided in an embodiment of the present invention. Detailed Implementation

[0010] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating a method for tracing the source of unconventional water source problems using knowledge graphs, as provided in an embodiment of the present invention. The following is a detailed description of this method for tracing the source of unconventional water source problems using knowledge graphs.

[0011] Step S110: Obtain problem representation information of unconventional water sources, and map the problem representation information to a pre-constructed unconventional water source knowledge graph. The problem representation information includes abnormal water source quality, abnormal development process, abnormal equipment operation, and abnormal environmental correlation. The unconventional water source knowledge graph includes water source feature units, process feature units, equipment feature units, and environmental feature units, and each feature unit is connected by semantic correlation edges.

[0012] In this embodiment, a reclaimed water reuse system in an industrial park is used as a unified application scenario. This system treats industrial wastewater and reuses it in processes such as cooling circulation within the park. First, it is necessary to obtain the problem characterization information of this reclaimed water reuse system and map it onto a knowledge graph. This process involves the collection and integration of various types of abnormal data.

[0013] Step S111: Collect abnormal data related to water source quality through an unconventional water source monitoring system. The abnormal data includes changes in water source composition, changes in water source turbidity, and changes in water source odor. Organize the collected abnormal data related to water source quality into abnormal water source quality characteristics.

[0014] In this reclaimed water reuse system, the unconventional water source monitoring system includes an array of water quality sensors distributed at various nodes of the treatment process. For example, a multi-parameter online water quality monitor installed at the outlet of the sand filter continuously collects data. When the concentration of suspended solids in the water exceeds the preset range and shows a continuous upward trend, it indicates a change in water turbidity. Simultaneously, a TOC detector after the biological treatment unit shows periodic fluctuations in organic carbon content, with the fluctuation amplitude gradually increasing, indicating a change in water composition. In addition, an odor sensor located in the clear water tank detects an abnormal odor similar to humus at specific times each day, and the odor intensity intermittently increases over time, constituting a change in water odor. These different types of abnormal data are classified and summarized according to the occurrence time, detection location, and abnormal characteristics to form a structured set of water source quality anomaly data. Each record includes an anomaly type label, start timestamp, duration, and associated detection device number.

[0015] Step S112: Collect process abnormal data during the development of unconventional water sources. The process abnormal data includes stagnation in the development process, disorder in the connection between processes, and deviation in the operation of processes. The collected process abnormal data during the development of unconventional water sources is summarized as abnormal performance of the development process.

[0016] The development process of this reclaimed water reuse system collects relevant data through PLC control systems and MES production execution systems deployed at each process stage. For example, in the membrane filtration stage, the system recorded instances where the filter membrane module automatically stopped operating during normal operation and could not be restarted using conventional start commands, indicating a development stage stagnation. Simultaneously, during material transfer between the pretreatment and reverse osmosis units, the raw water transfer pump started, but the valves did not open according to the preset logic, causing material transfer interruptions and constituting a stage connection disorder. Furthermore, in the chemical cleaning stage, the operator did not adjust the cleaning agent ratio according to standard procedures, resulting in unsatisfactory cleaning effects and exceeding the specified cleaning time, representing a stage operation deviation. These abnormal process data are standardized, and abnormal events are arranged according to the process stage sequence. Each abnormal event is assigned a corresponding stage number, abnormal code, and description of its impact scope, thus summarizing it as an abnormal manifestation of the development process.

[0017] Step S113: Collect abnormal equipment operation data for unconventional water source development. The abnormal equipment operation data includes abnormal equipment start-up and shutdown behavior, abnormal equipment parameter fluctuation behavior, and abnormal equipment output behavior. The collected abnormal equipment operation data for unconventional water source development is converted into abnormal equipment operation behavior.

[0018] The main equipment involved in this reclaimed water reuse system includes bar screens, aeration blowers, dosing pumps, and reverse osmosis membrane modules. Operational data is collected through built-in sensors and an IoT gateway. For example, a bar screen may automatically shut down and fail to restart without a fault alarm signal, even though the control cabinet indicator lights show normal power; this indicates an abnormal start-up / shutdown behavior. Similarly, the aeration blower's pressure sensor may show irregular and drastic pressure fluctuations within a short period, exceeding the upper and lower limits of the normal operating range; and the dosing pump may consistently output chemicals below the theoretical value at the set dosage, even after calibration, indicating an abnormal output. This abnormal data is categorized and statistically analyzed according to equipment type, installation location, and frequency of occurrence. Detailed records are created for each type of abnormality, including equipment model, abnormal parameter name, duration of abnormality, and associated process flow, transforming it into structured data on abnormal equipment operation.

[0019] Step S114: Monitor abnormal environmental data around unconventional water source development areas. The abnormal environmental data includes abnormal soil conditions, abnormal vegetation growth, and abnormal surrounding water bodies. The monitored abnormal environmental data around unconventional water source development areas are integrated into environmentally related abnormal data.

[0020] Environmental monitoring stations have been set up around the industrial park where the reclaimed water reuse system is located, including soil moisture monitors, vegetation growth cameras, and water quality sampling points for surrounding water bodies. For example, soil sensors detected an abnormal decrease in soil pH and a simultaneous increase in soil conductivity in a specific area around the treatment plant, indicating abnormal soil conditions. Large areas of vegetation in the plant's green belt showed yellowing leaves, and their growth height was significantly lower than in previous years, constituting abnormal vegetation growth. Abnormal algal growth, reduced water transparency, and a decreasing dissolved oxygen content were observed in the rainwater collection pond downstream of the system's discharge outlet, indicating abnormal water quality in the surrounding area. Spatiotemporal matching of these environmental anomaly data was performed, and correlation analysis was conducted on anomaly data of different environmental elements within the same time period and area, integrating them into a dataset of environmentally correlated anomalies that includes environmental element types, anomaly severity levels, and spatial distribution characteristics.

[0021] Step S115: Combine the abnormal water quality, abnormal development process, abnormal equipment operation, and environmental correlation to form problem characterization information.

[0022] After collecting and organizing the aforementioned anomaly data, the four categories of data—anomalies in water source quality, anomalies in the development process, anomalies in equipment operation, and anomalies related to the environment—are combined. For example, changes in water turbidity, stagnation in membrane filtration, abnormal output of dosing pumps, and decreases in pH of surrounding soil occurring within the same time window (e.g., a continuous 48-hour period) are correlated to form a complete problem event record. Each problem characterization information includes a unique event number, the time range of the anomaly, the involved process units, a list of associated equipment, and the environmental impact area, ensuring that effective spatiotemporal correlations can be established between various anomalies. Step S116: Invoke the pre-built unconventional water source knowledge graph, wherein the water source feature units in the pre-built unconventional water source knowledge graph correspond to water source quality related attributes, the process feature units in the pre-built unconventional water source knowledge graph correspond to development process related attributes, the equipment feature units in the pre-built unconventional water source knowledge graph correspond to equipment operation related attributes, and the environmental feature units in the pre-built unconventional water source knowledge graph correspond to environmental association related attributes.

[0023] A pre-constructed unconventional water source knowledge graph is invoked through a knowledge graph management platform. This unconventional water source knowledge graph is stored in the Neo4j graph database. The knowledge graph includes water source characteristic units such as entity nodes like "turbidity after sand filtration," "TOC content," and "olfactory threshold," each associated with attributes such as testing standards, normal ranges, and influencing factors. Process characteristic units cover entities such as "membrane filtration," "chemical cleaning process," and "chemical dosing steps," associated with attributes such as standard operating procedures, time requirements, and sequential connections. Equipment characteristic units include equipment entities such as "reverse osmosis membrane modules," "dosing pumps," and "aeration fans," with attributes such as model specifications, operating parameter ranges, and maintenance cycles. Environmental characteristic units include environmental entities such as "surrounding soil," "green belt vegetation," and "rainwater collection ponds," associated with attributes such as environmental quality standards, monitoring frequency, and background values. Connections between characteristic units are established through semantic association edges; for example, there is an "execution" type association edge between the "dosing pump" equipment characteristic unit and the "chemical dosing steps" process characteristic unit.

[0024] Step S117: Extract key descriptive content from the problem representation information, and perform semantic matching between the extracted key descriptive content and the attribute descriptions of each feature unit in the pre-constructed unconventional water source knowledge graph to determine the feature units corresponding to the problem representation information.

[0025] Natural language processing is applied to the generated problem representation information. Named entity recognition technology is used to extract key descriptive content. For example, keywords such as "sand filter outlet," "suspended solids concentration," and "continuous increase" are extracted from abnormal water quality phenomena; key phrases such as "membrane filtration stage," "automatic shutdown," and "unrecoverable" are extracted from abnormal development process phenomena. A BERT-based semantic matching model is used to calculate the similarity between these key descriptive contents and the attribute descriptions of each feature unit in the knowledge graph. For example, the key description "dosing pump output is lower than theoretical value" is matched with the attribute description "output flow deviation" of the "dosing pump" feature unit. When the similarity exceeds a set threshold, the problem description corresponds to the "dosing pump" feature unit. For keywords with ambiguity, disambiguation processing is performed based on the context and entity relationships in the knowledge graph to ensure the accuracy of the matching results.

[0026] Step S118: Mark the matching feature units in the pre-constructed unconventional water source knowledge graph to complete the mapping of problem representation information to the pre-constructed unconventional water source knowledge graph. The semantic association edges between each feature unit are used to represent the attribute association relationship between different feature units.

[0027] In the knowledge graph visualization interface, feature units identified through semantic matching are highlighted. For example, the water source feature unit node corresponding to "abnormal turbidity at the sand filter outlet" is marked in red, and its associated abnormal time range and detection value change trend are displayed. Simultaneously, semantic association edges directly connected to this feature unit are automatically retrieved, such as the "affected by" type association edge between "turbidity at the sand filter outlet" and "sand filter pump operating status," and the "association" type association edge between "turbidity at the sand filter outlet" and "backwashing frequency." The feature units corresponding to each abnormal manifestation in the problem representation information and their associations are collectively constructed into a subgraph structure and stored in a temporary workspace as the basis for subsequently constructing the problem association transmission chain. Through the above mapping method, the transformation of unstructured problem representation information into structured knowledge graph nodes is realized.

[0028] Step S120: Based on the mapping results of the problem representation information and the unconventional water source knowledge graph, construct a problem association transmission chain. The problem association transmission chain starts from the feature unit corresponding to the mapping and extends along the semantic association edge to other feature units in the knowledge graph. Record the association type of the feature units and semantic association edges passed through during the extension process.

[0029] After mapping the problem representation information to the knowledge graph, starting from the marked feature units, the path is extended according to the semantic association edges in the knowledge graph to construct the problem association transmission chain. This process aims to clarify the propagation path of the impact of anomalies.

[0030] Step S121: Select the core feature unit corresponding to each abnormal behavior from the feature units corresponding to the problem representation information mapping, and take the core feature unit corresponding to each abnormal behavior as the starting point of the problem association transmission chain.

[0031] For multiple feature units obtained from mapping, core feature units are selected based on the type and degree of impact of the anomalies. For example, in the case of abnormal water quality, the feature unit corresponding to "abnormal turbidity at the sand filter outlet" is selected as the core feature unit because it directly reflects the key indicators of the treated water; in the case of abnormal equipment operation, the feature unit corresponding to "abnormal output of the dosing pump" is selected as the core feature unit because it may affect the dosing effect of chemicals in subsequent treatment stages. Each type of anomaly corresponds to one core feature unit to ensure that the beginning of the transmission chain is representative. During the selection process, factors such as the frequency of anomaly occurrence, the scope of impact on the overall system operation, and the degree of criticality in the process flow are comprehensively considered. Automatic screening is performed through an expert rule base. For feature units with disputes, entity importance scoring in a knowledge graph is used to assist in decision-making.

[0032] Step S122: Query all semantic association edges connected to the starting feature unit of the problem association transmission chain in the unconventional water source knowledge graph, and determine the next-level feature unit pointed to by each semantic association edge found.

[0033] Taking the initial feature unit "abnormal output of the dosing pump" as an example, a query operation is performed using a knowledge graph query language (such as Cypher) to retrieve all outgoing and incoming edges of this node. For example, the query results show that this feature unit has incoming edges of type "controlled by," connecting to the "dosing pump control cabinet" equipment feature unit; outgoing edges of type "affect," connecting to the "reverse osmosis membrane fouling rate" water source feature unit; and edges of type "related," connecting to the "reagent ratio" process feature unit. For each semantically related edge, its related type and the target feature unit it points to are recorded, forming an edge-node correspondence table. During the query process, the directional attribute of the edges needs to be considered to distinguish the influence of directed and undirected edges on the transmission path, ensuring that the determination of the next-level feature unit conforms to the topological structure of the knowledge graph.

[0034] Step S123: Identify the association type corresponding to each semantic association edge found in the query. The association type includes attribute dependency association type, causal trigger association type, condition constraint association type, and state linkage association type.

[0035] The semantic association edges retrieved are categorized. For example, the association edge between "abnormal dosing pump output" and "reagent ratio" is labeled as an "attribute-dependent association type," because the reagent ratio directly determines the output concentration attribute of the dosing pump; the association edge between "abnormal dosing pump output" and "reverse osmosis membrane fouling rate" is a "causal triggering association type," because insufficient dosing directly leads to an accelerated membrane fouling rate; the association edge between "dosing pump control cabinet" and "abnormal dosing pump output" is a "conditional constraint association type," because the voltage stability of the control cabinet constrains the operating state of the dosing pump; if there is a simultaneous start-stop relationship between "dosing pump" and "aeration fan" due to changes in system load, then their association edge is a "state-linked association type." Association type identification is based on a predefined edge type system in the knowledge graph, achieved by comparing edge type labels. For edges without explicitly labeled types, an association type prediction model based on graph neural networks is used for inference.

[0036] Step S124: Starting from the initial feature unit of the problem association transmission chain, extend along the next level feature unit pointed to by each semantic association edge found, and record the sequence of feature units passed through during the extension process and the association type of semantic association edges between each feature unit.

[0037] Starting from the initial feature unit "Abnormal Dosing Pump Output," the path extends along the "Influence" type association edge to the "Reverse Osmosis Membrane Fouling Rate" feature unit, recording this path as [Abnormal Dosing Pump Output] - Influence -> [Reverse Osmosis Membrane Fouling Rate]. Simultaneously, it extends along the "Attribute Dependence" type association edge to the "Reagent Ratio" feature unit, recording the path as [Abnormal Dosing Pump Output] - Attribute Dependence -> [Reagent Ratio]. During the extension process, a timestamp is added to each feature unit traversed to distinguish the order of different extension paths. For nodes containing multiple next-level feature units, a breadth-first search strategy is used to extend sequentially, ensuring that every possible path is covered. The recorded feature unit sequence is arranged according to the extension order, and the association type is stored one-to-one with the corresponding edge.

[0038] Step S125: For the next level feature unit reached by the extension, repeat the operation of querying the semantic association edge connected to the next level feature unit and the feature unit pointed to by the semantic association edge, and continue to extend along the semantic association edge until the terminal feature unit in the unconventional water source knowledge graph has no subsequent associated feature units.

[0039] Taking the extended "reverse osmosis membrane fouling rate" feature unit as an example, a semantic association edge query operation is performed again. It is found that there are outgoing edges of type "cause" pointing to the "shortened membrane cleaning cycle" process feature unit, and edges of type "association" pointing to the "feed water pressure" equipment feature unit. Continuing to extend along the "cause" type edges to the "shortened membrane cleaning cycle" feature unit, the associated edges of this unit are queried, revealing that they point to the "increased chemical cleaning agent consumption" process feature unit. Since this unit has no further associated downstream feature units, it is identified as the terminal feature unit. During this process, an extension depth threshold needs to be set to avoid infinite path extension due to an excessively large knowledge graph. When the extension depth reaches the threshold without encountering a terminal feature unit, the extension is forcibly terminated, and the current feature unit is marked as a temporary terminal node.

[0040] Step S126: Arrange the feature unit sequence and corresponding association type formed by extending the starting feature unit of each problem association transmission chain in the order of extension to form multiple initial transmission paths with the starting feature unit of the problem association transmission chain as the starting point and the ending feature unit as the ending point.

[0041] For example, one initial transmission path extending from the starting point of "abnormal dosing pump output" is: [abnormal dosing pump output] - attribute dependency -> [reagent ratio] - condition constraint -> [operator training record] - association -> [standard operating procedure version], where "standard operating procedure version" is the terminal characteristic unit; another path is: [abnormal dosing pump output] - impact -> [reverse osmosis membrane fouling rate] - result -> [shortened membrane cleaning cycle] - association -> [increased chemical cleaning agent consumption], where "increased chemical cleaning agent consumption" is the terminal characteristic unit. These paths are grouped according to their starting point characteristic units. Each group contains all possible paths extending from the same starting point. Each path is stored in the form of a characteristic unit sequence and an association type sequence, for example, a path represented as (characteristic unit list: [A, B, C, D], association type list: [T1, T2, T3]).

[0042] Step S127: Perform deduplication on the multiple initial conduction paths formed, delete initial conduction paths with identical path structures, and retain initial conduction paths with unique path structures.

[0043] Deduplication is performed by comparing the complete consistency of the feature unit sequence and the association type sequence. For example, two paths are considered duplicate paths if they contain feature units in the exact same order and have the same association type. A hash algorithm is used to generate a unique identifier for each path. The hash value is calculated by concatenating the feature unit ID sequence and the association type ID sequence of the path. Paths with the same hash value are considered duplicate paths. During deduplication, the first path instance found is retained, and subsequent duplicate instances are deleted. Paths with the same feature unit order but different association types, or paths with the same association type but different feature units, are considered different paths and retained. The deduplication operation is implemented in the path database through an indexing mechanism to improve processing efficiency.

[0044] Step S128: Combine the deduplicated initial transmission paths to form a problem association transmission chain that covers all associated feature units, so that the formed problem association transmission chain contains the complete extension path from the starting feature unit to each terminal feature unit and the corresponding association type information.

[0045] All deduplicated initial transmission paths are categorized and summarized according to their starting feature units, forming a forest structure containing multiple directed trees, i.e., the problem-related transmission chain. For example, the tree rooted at "abnormal output of the dosing pump" contains multiple branch paths, each branch corresponding to an initial transmission path; the tree rooted at "abnormal turbidity at the sand filter outlet" contains another set of branch paths. A graph merging algorithm is used to integrate these tree structures into a unified directed acyclic graph, ensuring that all associated feature units are included. During the merging process, the intersection nodes and shared sub-paths between paths are recorded to avoid redundant processing during subsequent path strength calculations. The final problem-related transmission chain is stored in the form of an adjacency list, containing three parts: a set of nodes, a set of edges, and a set of paths.

[0046] Step S130: Calculate the influence attenuation coefficient of each semantic association edge in the problem association transmission chain, and generate an attenuation coefficient matrix containing the influence attenuation coefficient of each semantic association edge. The influence attenuation coefficient is determined based on the transmission strength corresponding to the association type and the semantic distance between feature units.

[0047] To quantify the influence of different associated edges during the propagation process, it is necessary to calculate the influence attenuation coefficient of each semantic associated edge and organize these coefficients in matrix form.

[0048] Step S131: Establish a correspondence table between association types and transmission strengths, where attribute dependency association types correspond to the first transmission strength, causal trigger association types correspond to the second transmission strength, conditional constraint association types correspond to the third transmission strength, and state linkage association types correspond to the fourth transmission strength. The transmission strengths corresponding to different association types have different fixed value ranges.

[0049] Through domain expert surveys and historical data statistical analysis, a mapping relationship between association types and transmission strength is established. For example, the numerical range of causal trigger association types (second transmission strength) is set to a higher range because it represents a direct causal relationship with a strong transmission effect; the range of attribute dependency association types (first transmission strength) is next, representing the dependency relationship between attributes; the range of conditional constraint association types (third transmission strength) is lower, representing the indirect conditional constraint effect; and the range of state linkage association types (fourth transmission strength) is the lowest, representing a weak association with synchronous state changes. Each numerical range is defined by a fuzzy membership function, for example, the first transmission strength corresponds to [0.6, 0.8], the second transmission strength corresponds to [0.8, 1.0], the third transmission strength corresponds to [0.4, 0.6], and the fourth transmission strength corresponds to [0.2, 0.4]. This mapping table is stored in the metadata layer of the knowledge graph and can be dynamically adjusted according to the actual application effect.

[0050] Step S132: Extract the transmission strength value corresponding to the association type of each semantic association edge in the problem association transmission chain from the established correspondence table between association type and transmission strength.

[0051] For each semantic association edge in the problem association transmission chain, the corresponding relationship table is queried according to its association type to obtain the transmission strength value range. For example, if the association type of an edge is causal triggering, then the range of the second transmission strength [0.8, 1.0] is extracted. When selecting a specific value within the range, it is weighted by combining the historical transmission effect data of the edge in the knowledge graph. For example, if the average transmission strength of this type of edge in past source tracing cases is 0.85, then values ​​close to this value are selected first in this extraction. For newly emerging association types or edges lacking historical data, the median value of the range is taken as the initial transmission strength value. During the extraction process, a list of transmission strength values ​​is generated, corresponding one-to-one with the unique identifier of the edge.

[0052] Step S133: Calculate the semantic distance between two feature units connected by each semantic association edge in the problem association transmission chain. The semantic distance is determined based on the difference in node level of the two feature units in the unconventional water source knowledge graph and the similarity of the attributes of the two feature units. The smaller the difference in node level of the two feature units in the unconventional water source knowledge graph and the higher the similarity of the attributes of the two feature units, the smaller the semantic distance between the two feature units.

[0053] First, determine the hierarchical position of the feature units in the knowledge graph. A topological sorting algorithm is used to divide all nodes in the knowledge graph into different levels; for example, basic equipment nodes are divided into level 1, process node nodes into level 2, and system performance index nodes into level 3, etc. The hierarchical difference between two feature units is calculated, which is the absolute value of the difference in hierarchical values. Then, the attribute similarity between the two feature units is calculated. After vectorizing the attribute values ​​of each feature unit, the similarity value is calculated using the cosine similarity formula. Attributes include type, functional description, and associated entities. After normalizing the hierarchical difference and attribute similarity, they are substituted into the semantic distance calculation formula, for example, semantic distance = α × normalized hierarchical difference + (1-α) × (1- normalized attribute similarity), where α is a weight coefficient, and the optimal value is determined using a grid search method. The smaller the final semantic distance value, the closer the two feature units are semantically.

[0054] Step S134: Set the semantic distance decay function. Input the calculated semantic distance into the set semantic distance decay function to obtain the semantic distance decay value. The larger the calculated semantic distance, the larger the semantic distance decay value.

[0055] An exponential decay function is used as the semantic distance decay function. For example, the decay function is set as semantic distance decay value = 1 - exp(-β × semantic distance), where β is the decay coefficient, adjusted according to the average path length of the knowledge graph to ensure that the decay value is close to 1 when the semantic distance is equal to the average diameter of the knowledge graph. The calculated semantic distance is then substituted into this function to obtain the corresponding decay value. For example, when the semantic distance between two feature units is small (e.g., 0.2), a smaller decay value (e.g., 0.18) is obtained after substituting into the function; when the semantic distance is large (e.g., 1.5), a larger decay value (e.g., 0.78) is obtained. The semantic distance decay value reflects the degree of attenuation of the influence transmission caused by semantic differences; a larger value indicates a more severe attenuation.

[0056] Step S135: Calculate the influence attenuation coefficient based on the extracted transmission intensity value and the obtained semantic distance attenuation value. The calculation method is to multiply the extracted transmission intensity value by (1 minus the obtained semantic distance attenuation value) to obtain the influence attenuation coefficient of each semantic association edge in the problem association transmission chain.

[0057] For example, if the propagation strength of a semantic association edge is 0.85 and the semantic distance attenuation value is 0.3, then the influence attenuation coefficient = 0.85 × (1 - 0.3) = 0.595. In this calculation, the propagation strength value represents the inherent propagation capability of the association type, and (1 - semantic distance attenuation value) represents the proportion of propagation capability retained by the semantic distance. The product of the two reflects the actual effective strength of the edge in the influence propagation process. After the calculation is completed, the influence attenuation coefficient is truncated to ensure that its value is within the interval [0, 1]. Calculation results less than 0 are taken as 0, and results greater than 1 are taken as 1.

[0058] Step S136: Standardize the calculated impact attenuation coefficient, statistically analyze the identification information of all semantic related edges in the problem-related transmission chain, and match the identification information of each semantic related edge with the corresponding standardized impact attenuation coefficient.

[0059] The min-max normalization method is used to process the influence attenuation coefficients, mapping all coefficients to the interval [0, 1]. The normalization formula is: Normalized coefficient = (Original coefficient - Minimum coefficient) / (Maximum coefficient - Minimum coefficient). For example, if the minimum original coefficient is 0.2 and the maximum is 0.9, and the original coefficient of a certain edge is 0.595, then the normalized coefficient = (0.595 - 0.2) / (0.9 - 0.2) = 0.564. A unique identifier, such as an edge ID, is used for all semantically related edges in the statistical problem's correlation propagation chain. A mapping table between edge IDs and normalized influence attenuation coefficients is established to ensure that each edge ID corresponds to a unique normalized coefficient value.

[0060] Step S137: Using the identification information of the semantically related edges as the row and column indices, fill the corresponding influence attenuation coefficients into the corresponding positions of the matrix to generate an attenuation coefficient matrix containing the influence attenuation coefficients of each semantically related edge. Fill the positions of the feature units that are not directly connected in the generated attenuation coefficient matrix with zero values.

[0061] Construct an N×N square matrix, where N is the total number of feature units in the problem-related propagation chain, and the row and column indices are unique identifiers of the feature units. For two feature units (i, j) connected by a direct semantic edge, fill the matrix with the normalized influence attenuation coefficient corresponding to that edge at the i-th row and j-th column position; for feature unit pairs without a direct connection, fill with zero. For example, if feature unit A has an ID of 101 and feature unit B has an ID of 102, and there is a semantic edge between them with a normalized influence attenuation coefficient of 0.564, then the value at position (101, 102) in the matrix is ​​0.564. The value at position (102, 101) is determined based on the directionality of the edge; if it is a directed edge from A to B, then position (102, 101) is filled with 0. The generated attenuation coefficient matrix is ​​stored in sparse matrix form to save storage space.

[0062] Step S140: Based on the attenuation coefficient matrix, perform path filtering on the problem-related transmission chain, retain the transmission paths whose influence attenuation coefficient meets the transmission requirements, and remove the transmission paths whose influence attenuation coefficient does not meet the transmission requirements, to obtain the effective related transmission paths.

[0063] The initial conduction path is quantitatively evaluated using the attenuation coefficient matrix, and effective paths with practical significance are selected, reducing the computational load of subsequent source tracing analysis.

[0064] Step S141: Analyze the number of semantically related edges contained in each initial transmission path in the problem-related transmission chain, and determine the set of semantically related edges corresponding to each initial transmission path after analysis.

[0065] Traverse each initial propagation path and count the number of semantically related edges contained in the path. For example, if a path contains 3 feature units and 2 semantically related edges, then the number of edges is 2. For each path, uniquely identify its exact semantically related edges using edge IDs to form a set of semantically related edges. For example, the edge set of path P1 is {E101, E105, E109}, where E is the edge ID prefix. When determining the edge set, it is necessary to record them according to the order of the edges in the path to form an ordered set, in order to preserve the directional information of the path.

[0066] Step S1411: Traverse each initial propagation path in the problem-related propagation chain, and record the starting end feature unit identifier and the ending feature unit identifier of each initial propagation path.

[0067] Each initial conduction path is visited sequentially using a path traversal algorithm (such as depth-first search), starting from the beginning of the path and tracking the order of the feature units until the end. For example, for the path [abnormal dosing pump output] -> [reverse osmosis membrane fouling rate] -> [shortened membrane cleaning cycle], the starting feature unit is recorded as "PUMP001" and the ending feature unit is recorded as "CYCLE002". This identification information is stored in the path metadata table, with one record for each path, containing fields such as path ID, starting identifier, and ending identifier.

[0068] Step S1412: Along each initial propagation path traversed, traverse sequentially from the starting feature unit of the initial propagation path to the ending feature unit of the initial propagation path, and record the identifier of each feature unit passed through the initial propagation path and the identifier of the semantic association edge connecting adjacent feature units.

[0069] Taking the initial propagation path with path ID P001 as an example, starting from the initial feature unit "PUMP001", it sequentially passes through feature units such as "MEM003" and "CYCLE002", recording the identifier of each feature unit. Simultaneously, the semantic association edge identifiers between adjacent feature units are recorded; for example, the edge identifier between "PUMP001" and "MEM003" is "E102", and the edge identifier between "MEM003" and "CYCLE002" is "E205". The feature unit identifier sequence and the edge identifier sequence are stored separately to form a detailed topological structure record of the path.

[0070] Step S1413: Count the number of semantically related edge identifiers in each initial transmission path. The counted number is the number of semantically related edges contained in the initial transmission path.

[0071] Calculate the length of the edge identifier sequence for each path to obtain the number of semantically related edges. For example, if the length of the edge identifier sequence [E102, E205] is 2, then the path contains 2 semantically related edges. Update the "Number of Edges" field in the path metadata table with the statistical results.

[0072] Step S1414: Arrange all semantic association edge identifiers recorded in each initial transmission path according to the extension order of the initial transmission path to form the semantic association edge sequence corresponding to the initial transmission path.

[0073] The edge identifiers are organized into an ordered sequence according to the order in which the edges appear during the path extension process. For example, if the path extension order is A->B->C->D, and the corresponding edge identifiers are E1, E2, and E3 respectively, then the semantically associated edge sequence is [E1, E2, E3]. This semantically associated edge sequence preserves the structural information of the path.

[0074] Step S1415: Uniquely identify the semantic association edge sequence corresponding to each initial transmission path, and associate and store the semantic association edge sequence corresponding to each initial transmission path with the starting end feature unit identifier of the corresponding initial transmission path, the ending end feature unit identifier of the corresponding initial transmission path, and the number of semantic association edges contained in the corresponding initial transmission path.

[0075] A unique sequence ID is generated for each semantically related edge sequence. The edge identifier sequence is hashed using a hash function (such as SHA-1) to obtain a fixed-length hash value as the sequence ID. The sequence ID, start node identifier, end node identifier, and edge count are stored in an associated data table to enable multi-dimensional information lookup. For example, the sequence ID can be used to quickly retrieve the corresponding path start and end nodes and the number of edges.

[0076] Step S1416: Check whether there are duplicate semantic association edge identifiers in the semantic association edge sequence corresponding to each initial transmission path. If there are duplicate semantic association edge identifiers in the semantic association edge sequence corresponding to an initial transmission path, confirm whether the initial transmission path is a path loop. If the initial transmission path is a path loop, delete the semantic association edge identifiers of the path loop part of the initial transmission path and retain the semantic association edge identifiers of the non-loop part of the initial transmission path.

[0077] Repeated element detection is performed on semantically related edge sequences. For example, in the sequence [E101, E102, E101, E103], there is a repeated E101. By examining the sequence of feature units between repeated edge identifiers, it is determined whether a cyclic path is formed, such as A->B->A->C where A->B->A constitutes a cycle. If a cycle is confirmed, the edge identifiers of the cyclic part are deleted from the sequence, retaining the non-cyclic part, such as modifying the above sequence to [E101, E102, E103]. When deleting the cyclic part, it is necessary to ensure that the starting and ending feature units of the path are preserved to avoid destroying the integrity of the path.

[0078] Step S1417: Standardize the format of the semantic association edge sequence corresponding to each processed initial propagation path so that all processed semantic association edge sequences are presented with the same identifier format and arrangement order.

[0079] The format of semantically related edge sequences is standardized. For example, edge identifiers are specified as strings of "E + number", separated by commas, and arranged strictly in the direction of path extension from the start to the end. Sequences of length 0 (paths containing only a single feature unit) are uniformly represented as empty sequences. All sequences are standardized using a format validation tool to ensure format consistency and facilitate subsequent batch processing.

[0080] Step S1418: Take the processed semantic association edge sequence of each initial transmission path as the semantic association edge set corresponding to that initial transmission path, and establish a one-to-one correspondence table between the initial transmission path and the corresponding semantic association edge set.

[0081] Using the path ID as the primary key and the processed semantically related edge sequences as field values, a path-edge set mapping table is constructed. For example, the edge set corresponding to path IDP001 is [E101, E105, E109], and the edge set corresponding to path IDP002 is [E203, E207], etc. This mapping table is stored in the database, supporting quick querying of the corresponding edge set by path ID, or reverse querying of the path containing that edge by edge ID.

[0082] Step S142: Extract the influence attenuation coefficient of each semantically associated edge from the set of semantically associated edges corresponding to each initial propagation path from the attenuation coefficient matrix.

[0083] Based on the semantically associated edge set corresponding to the initial propagation path, the influence attenuation coefficient of each edge is retrieved from the attenuation coefficient matrix. For example, for the edge set [E101, E105, E109], the corresponding position is located in the row and column indices of the attenuation coefficient matrix by the edge ID, and the coefficient value in the matrix is ​​extracted. If an edge in the edge set does not exist in the attenuation coefficient matrix (possibly due to deduplication or removal during loop processing), the edge is skipped or marked with a coefficient value of 0. The extracted coefficient values ​​are arranged in the order of the edges in the sequence, forming a coefficient sequence.

[0084] Step S143: Calculate the product of the influence attenuation coefficients of all semantically related edges in each initial conduction path, and use the calculated product as the overall conduction strength value of the initial conduction path.

[0085] Multiply all the values ​​in the coefficient sequence of each path. For example, the product of the coefficient sequence [0.6, 0.7, 0.8] is 0.6 × 0.7 × 0.8 = 0.336. This product is the overall conduction strength value of the path. For paths containing 0 edges (only the starting feature unit), the overall conduction strength value is set to 1 (indicating that the conduction strength to itself is the maximum). Floating-point multiplication is used in the calculation process, retaining sufficient decimal places to avoid precision loss.

[0086] Step S144: Set an overall conduction intensity threshold, which is determined based on the accuracy requirements of tracing unconventional water source problems and the effective path intensity range of historical tracing data.

[0087] Taking into account the accuracy requirements for tracing unconventional water source issues (such as the minimum allowable error rate) and the overall transmission intensity distribution of effective paths in historical tracing cases, a threshold is set using statistical analysis methods. For example, cluster analysis is performed on the overall transmission intensity values ​​of effective paths in historical data, and the cluster center value with the highest proportion is taken as the initial threshold. Then, it is dynamically adjusted based on the actual tracing results. If high-precision tracing is required, the threshold is set higher; if more potential paths need to be covered, the threshold is appropriately lowered. The threshold is stored in the system configuration file after being set, and manual modification is supported.

[0088] Step S145: Compare the overall conduction strength value of each initial conduction path with the set overall conduction strength threshold. If the overall conduction strength value of an initial conduction path is greater than or equal to the set overall conduction strength threshold, then mark the initial conduction path as a candidate valid path; if the overall conduction strength value of an initial conduction path is less than the set overall conduction strength threshold, then mark the initial conduction path as an invalid path.

[0089] The process iterates through all initial conduction paths, comparing their overall conduction strength value with a threshold. For example, with a threshold of 0.3, if a path has an overall conduction strength value of 0.336, it is marked as a candidate valid path; if another path has an overall conduction strength value of 0.25, it is marked as an invalid path. During the comparison, paths with an overall conduction strength value equal to the threshold are considered to meet the conduction requirements and are marked as candidate valid paths. The marking results are stored in the "Status" field of the path metadata table.

[0090] Step S146: Perform a path integrity check on the initial propagation path marked as a candidate valid path to confirm whether there are missing feature units or missing semantic association edges in the initial propagation path marked as a candidate valid path. If there are missing feature units or missing semantic association edges, complete them and recalculate the overall propagation strength value of the initial propagation path.

[0091] The candidate valid paths undergo a completeness check, encompassing both feature unit completeness and semantic association edge completeness. Feature unit completeness is checked by comparing the feature unit sequence in the path with the actual feature units in the knowledge graph to identify any missing or obsolete feature unit identifiers. Semantic association edge completeness is checked by comparing the path's edge set with the edge set in the attenuation coefficient matrix to identify any unrecorded edges or edges with a coefficient value of 0 (non-directly connected edges). If missing edges are found, they are supplemented using knowledge graph completion algorithms. For example, this involves inferring potentially missing intermediate edges based on indirect relationships between feature units, or migrating edge information from similar paths. After completion, the path's feature unit sequence and edge set are regenerated, and the overall transmission strength value is recalculated.

[0092] Step S1461: Extract the feature unit sequence of the initial propagation path labeled as a candidate valid path and the semantic association edge sequence of the initial propagation path labeled as a candidate valid path.

[0093] Extract the feature unit sequence and semantic association edge sequence from the metadata of the candidate valid paths. For example, the feature unit sequence of path P001 is [A, B, C, D], and the semantic association edge sequence is [E1, E2, E3]. Convert the above sequences into list form for easy element-level inspection.

[0094] Step S1462: Compare the extracted feature unit sequence with the feature unit library in the unconventional water source knowledge graph, and check whether each feature unit identifier in the extracted feature unit sequence exists in the feature unit library in the unconventional water source knowledge graph. If there is a feature unit identifier in the extracted feature unit sequence that does not exist in the feature unit library in the unconventional water source knowledge graph, then mark the feature unit identifier as a missing feature unit.

[0095] Access the feature unit library of the knowledge graph to obtain a list of identifiers for all valid feature units. Compare each identifier in the feature unit sequence of the path with this list one by one. If an identifier is found not in the list, mark it as missing. For example, if the identifier "X123" appears in the feature unit sequence but is not found in the feature unit library, then "X123" is marked as a missing feature unit.

[0096] Step S1463: Compare the extracted semantic association edge sequence with the semantic association edge library in the unconventional water source knowledge graph. Check whether each semantic association edge identifier in the extracted semantic association edge sequence exists in the semantic association edge library in the unconventional water source knowledge graph. If there is a semantic association edge identifier in the extracted semantic association edge sequence that does not exist in the semantic association edge library in the unconventional water source knowledge graph, then mark the semantic association edge identifier as missing semantic association edge.

[0097] Similarly, access the semantic association edge library of the knowledge graph to obtain a list of identifiers for all valid edges. Compare each identifier in the semantic association edge sequence of the path, and mark edges that do not exist in the list as missing. For example, if "E500" appears in the edge sequence but is not found in the edge library, then "E500" is marked as a missing semantic association edge.

[0098] Step S1464: For candidate valid paths marked with missing feature units, query the feature unit attribute information corresponding to the feature unit identifier marked as missing, select the feature unit with the most matching feature unit attribute information from the feature unit library in the unconventional water source knowledge graph as a supplementary feature unit, and insert the selected supplementary feature unit into the corresponding missing position of the feature unit sequence of the candidate valid path.

[0099] For missing feature unit identifiers, if their attribute information (such as name and description) is identifiable, a semantic search is used to find the feature unit with the most similar attributes from the feature unit library. For example, if the attribute description of the missing identifier "X123" is "pump used to deliver medicine", then the library is searched for feature units of type "pump" with a function description containing "medicine delivery", and "dosing pump B" with the highest similarity is selected as the supplementary feature unit and inserted into the missing position of the feature unit sequence.

[0100] Step S1465: For candidate valid paths marked with missing semantic association edges, based on the attribute information of the feature units before and after the marked missing semantic association edges, select the semantic association edge with the most matching association type from the semantic association edge library in the unconventional water source knowledge graph as a supplementary semantic association edge, and insert the selected supplementary semantic association edge into the corresponding missing position of the semantic association edge sequence of the candidate valid path.

[0101] Based on the attribute properties of the feature units before and after the missing edge, the possible association type can be inferred. For example, if the missing edge connects "dosing pump" and "pharmaceutical storage tank", and the attributes of the feature units before and after indicate that the former is equipment and the latter is a storage container, then an association edge of the "connection" type is selected from the edge library as a supplement and inserted into the missing position of the edge sequence.

[0102] Step S1466: After the supplementation is completed, a new sequence of feature units and a new sequence of semantic association edges are regenerated to form a complete candidate valid path.

[0103] After inserting the supplemented feature units and semantically related edges into the original sequence, a new sequence is formed. For example, the original sequence [A, missing, B] becomes [A, C, B] after supplementation, and the original edge sequence [E1, missing, E2] becomes [E1, E3, E2] after supplementation. The new sequence constitutes a complete candidate valid path.

[0104] Step S1467: Extract the influence attenuation coefficient of each semantically associated edge in the new semantically associated edge sequence of the complete candidate effective path from the attenuation coefficient matrix.

[0105] Based on the new semantically related edge sequence, the influence attenuation coefficient of each edge is extracted again from the attenuation coefficient matrix, using the same method as step S142.

[0106] Step S1468: Recalculate the product of the influence attenuation coefficients of all semantically related edges in the completed candidate effective path to obtain a new overall transmission strength value for the completed candidate effective path, and replace the original overall transmission strength value of the candidate effective path with the new overall transmission strength value.

[0107] The newly extracted coefficient sequence is multiplied to obtain a new overall conduction strength value, and the path metadata is updated to replace the original value.

[0108] Step S147: Verify the overall conduction strength value of the completed candidate valid path again to ensure that the overall conduction strength value of the completed candidate valid path still meets the requirement of being greater than or equal to the set overall conduction strength threshold.

[0109] The supplemented overall conduction strength value is compared again with the set threshold. If it still meets the requirements, it is retained as a candidate valid path; otherwise, it is marked as an invalid path. For example, if the supplemented overall conduction strength value is 0.28, which is less than the threshold of 0.3, it is marked as an invalid path.

[0110] Step S148: Remove all initial propagation paths marked as invalid paths, and organize the candidate valid paths that have been verified and meet the requirements into a set of valid associated propagation paths.

[0111] All validated candidate valid paths are collected, and their path IDs, feature unit sequences, edge sequences, and overall propagation strength values ​​are stored in a valid path database to form a set of valid associated propagation paths. This set of valid associated propagation paths contains all propagation paths that may have a significant impact on the problem.

[0112] Step S150: Based on the effective correlation transmission path, trace back to the starting feature unit, integrate the information of the effective correlation transmission path and the root cause feature unit, and generate the source tracing result of the unconventional water source problem. The starting feature unit is the root cause feature unit of the problem.

[0113] By tracing the effective correlation transmission path in reverse, the root cause characteristic unit of the problem is identified, and relevant information is integrated to form the final source tracing result report.

[0114] Step S151: Perform a reverse traversal of each valid association transmission path in the set of valid association transmission paths, backtracking from the terminal feature unit of each valid association transmission path to the starting feature unit of each valid association transmission path, and recording the feature units passed during the backtracking process and the semantic association edges passed during the backtracking process.

[0115] A reverse traversal is performed on each path in the set of valid associative transmission paths. For example, the reverse traversal order of the path [A->B->C->D] is D->C->B->A. During the backtracking process, the identifiers of the feature units traversed and their corresponding semantic association edge identifiers are recorded to form a reverse path record. For example, the sequence of feature units traversed during the backtracking is [D, C, B, A], and the sequence of semantic association edges is [E3, E2, E1] (the original path edge sequence is [E1, E2, E3]).

[0116] Step S1511: Obtain the forward feature unit sequence and the forward semantic association edge sequence of each effective association transmission path in the effective association transmission path set. The forward feature unit sequence and the forward semantic association edge sequence are arranged in order from the starting feature unit of the effective association transmission path to the ending feature unit of the effective association transmission path.

[0117] Read the forward sequence of each path from the set of effective associated transmission paths. For example, the forward feature unit sequence of path P001 is [A, B, C, D], and the forward semantic association edge sequence is [E1, E2, E3].

[0118] Step S1512: Reverse the forward feature unit sequence of each valid associated transmission path to obtain the reverse feature unit sequence from the terminal feature unit of the valid associated transmission path to the starting feature unit of the valid associated transmission path.

[0119] Perform a reversal operation on the forward feature unit sequence, for example, reverse [A, B, C, D] to [D, C, B, A] to obtain the reverse feature unit sequence.

[0120] Step S1513: Reverse the forward semantic association edge sequence of each effective association transmission path to obtain the reverse semantic association edge sequence from the terminal feature unit side of the effective association transmission path to the starting feature unit side of the effective association transmission path.

[0121] Perform a reversal operation on the forward semantic association edge sequence, for example, reverse [E1, E2, E3] to [E3, E2, E1] to obtain the reverse semantic association edge sequence.

[0122] Step S1514: Starting from the first feature unit of the reverse feature unit sequence of each valid associated propagation path, begin the reverse traversal of that valid associated propagation path.

[0123] The traversal begins from the first element of the reverse feature unit sequence (the terminal feature unit of the original path), for example, the reverse traversal of path P001 starts from D.

[0124] Step S1515: Read each feature unit identifier in the reverse feature unit sequence of each valid associated propagation path in sequence, and record the name of the feature unit corresponding to each read feature unit identifier and the attribute summary information of the feature unit.

[0125] Traverse the reverse feature unit sequence, read the feature unit identifier one by one, and query the knowledge graph to obtain its name and attribute summary. For example, the identifier "MEM003" corresponds to the name "reverse osmosis membrane module", and the attribute summary is "material: polyamide, molecular weight cutoff: 1000Da". Record the above information in the reverse traversal log.

[0126] Step S1516: Read the semantic association edge identifier of each semantic association edge in the reverse semantic association edge sequence of each valid association propagation path, and record the association type of the semantic association edge corresponding to each semantic association edge identifier and the influence attenuation coefficient of the semantic association edge.

[0127] Synchronously traverse the reverse semantic association edge sequence, read the edge identifier, and query the knowledge graph to obtain the association type and influence decay coefficient. For example, the association type corresponding to the edge identifier "E105" is "causal trigger association type", and the influence decay coefficient is 0.75. Record this information to the reverse traversal log and associate it with the corresponding feature unit information.

[0128] Step S1517: During the reverse traversal process, a reverse traversal record table is established for each valid association transmission path. The feature unit information in the reverse feature unit sequence of each valid association transmission path and the semantic association edge information in the reverse semantic association edge sequence of that valid association transmission path are filled into the reverse traversal record table of that valid association transmission path in a one-to-one correspondence according to the reverse traversal order.

[0129] Create a reverse traversal record table for each path. The table contains columns: sequence number, feature unit identifier, feature unit name, attribute summary, semantic association edge identifier, association type, and influence attenuation coefficient. Fill in the information sequentially according to the traversal order. For example, sequence number 1 corresponds to feature unit D and edge E3, sequence number 2 corresponds to feature unit C and edge E2, and so on.

[0130] Step S1518: When the reverse traversal reaches the last feature unit of the reverse feature unit sequence of each valid associated propagation path, stop the reverse traversal of that valid associated propagation path.

[0131] When the traversal reaches the last element of the reverse feature unit sequence (the starting feature unit of the original path), the reverse traversal of the path is stopped.

[0132] Step S1519: Repeat the above reverse traversal operation for all valid association transmission paths in the set of valid association transmission paths, and generate a corresponding reverse traversal record table for each valid association transmission path, so that the generated reverse traversal record table completely records the feature units passed through during the backtracking process of the valid association transmission path and the semantic association edge information passed through during the backtracking process.

[0133] Perform steps S1511 to S1518 on all paths in the set to generate their respective reverse traversal record tables, which are then stored in the database.

[0134] Step S152: During the reverse traversal, identify the earliest feature unit in each valid association propagation path that is not included in other valid association propagation paths, and tentatively designate the identified feature unit as the potential root feature unit.

[0135] Analyze the reverse traversal record table for each path, searching from the starting point (original path end) to the ending point (original path start) for the first feature unit that does not appear in the reverse traversal record table of any other path. For example, if the reverse traversal sequence of path P001 is [D, C, B, A] and the sequence of path P002 is [D, C, E, A], then feature units B and E are the earliest feature units in the two paths that are not included in the other, and are tentatively designated as potential root feature units. During the identification process, set operations are used to compare the feature unit sets of different paths to find the unique feature unit.

[0136] Step S153: Collect all potential root cause feature units corresponding to all effective correlation transmission paths, and count the frequency of occurrence of each collected potential root cause feature unit in different effective correlation transmission paths.

[0137] Summarize the potential root cause feature units of all paths into a list. For example, the list of potential root cause feature units is [B, E, B, F], where B appears twice, E appears once, and F appears once. Count the frequency of each feature unit.

[0138] For example, step S1531: Create a statistical table of potential root cause feature units, which includes three columns: feature unit identifier, feature unit type, and frequency of occurrence.

[0139] Create a three-column table to record the statistical results.

[0140] Step S1532: Traverse the reverse traversal record table of each valid association propagation path in the set of valid association propagation paths, and extract the potential root cause feature unit identifier and the feature unit type corresponding to the potential root cause feature unit identifier from the reverse traversal record table of each valid association propagation path.

[0141] Extract the identifier and type of the potential root cause feature unit from the reverse traversal record table of each path. For example, the type of the identifier "PUMP002" is "device feature unit".

[0142] Step S1533: Check whether the extracted potential root cause feature unit identifiers already exist in the feature unit identifier column of the potential root cause feature unit statistics table.

[0143] Query the statistics table to determine whether the current feature unit identifier has been recorded.

[0144] Step S1534: If the extracted potential root cause feature unit identifier already exists in the feature unit identifier column of the potential root cause feature unit statistics table, then add 1 to the value of the occurrence frequency column corresponding to the potential root cause feature unit identifier.

[0145] For existing identifiers, update the frequency count. For example, if the original frequency is 2, add 1 to make it 3.

[0146] Step S1535: If the extracted potential root cause feature unit identifier does not exist in the feature unit identifier column of the potential root cause feature unit statistics table, then add a new row to the potential root cause feature unit statistics table, fill in the potential root cause feature unit identifier in the feature unit identifier column of the new row, fill in the feature unit type corresponding to the potential root cause feature unit identifier in the feature unit type column of the new row, and set the occurrence frequency column value of the new row to 1.

[0147] For the new identifier, add a new row to the table and initialize the frequency to 1.

[0148] Step S1536: After extracting and counting all potential root cause feature units corresponding to all valid associated transmission paths, check whether there are duplicate feature unit identifiers in the potential root cause feature unit statistics table. If there are duplicate feature unit identifiers in the potential root cause feature unit statistics table, merge the duplicate rows and add up the occurrence frequencies corresponding to the duplicate rows.

[0149] Merge rows in the table that have duplicate identifiers, ensuring that each identifier occupies only one row and the frequency is the sum of the frequencies.

[0150] Step S1537: Sort the statistical table of potential root cause feature units after merging in descending order of the frequency column value, so that the row corresponding to the potential root cause feature unit with the highest frequency is placed at the top of the statistical table of potential root cause feature units.

[0151] The table is arranged in descending order of frequency to facilitate the rapid identification of high-frequency feature units.

[0152] Step S1538: In the sorted statistical table of potential root cause feature units, label the row corresponding to each potential root cause feature unit with the identifier of the effective associated transmission path corresponding to that potential root cause feature unit.

[0153] Add a "Path ID" column to the table to record all valid associated transmission path IDs containing the potential root cause feature unit, making it easier to trace the source.

[0154] Step S1539: Generate a statistical results report, which includes a sorted statistical table of potential root cause feature units and a description of the source path of each potential root cause feature unit.

[0155] The statistical tables and path descriptions were compiled into a report as an intermediate result of the source tracing analysis.

[0156] Step S154: Select the potential root cause feature unit with the highest frequency of occurrence as the final starting feature unit, that is, determine the selected starting feature unit as the root cause feature unit of the problem.

[0157] The most frequent potential root cause feature unit is selected from the sorted statistical table. For example, if the most frequent one is "abnormal output of the dosing pump", then it is identified as the root cause feature unit. If there are multiple feature units with the same frequency, a secondary screening is performed by calculating the centrality index (such as betweenness centrality and compact centrality) of these feature units in the knowledge graph, and the one with the highest centrality is selected as the root cause feature unit.

[0158] Step S155: Extract the attribute information of the root source feature unit. The attribute information of the root source feature unit includes the type of the root source feature unit, the specific parameters corresponding to the root source feature unit, the hierarchical position of the root source feature unit in the unconventional water source knowledge graph, and other feature units associated with the root source feature unit.

[0159] Query the detailed attributes of the root feature unit from the knowledge graph. For example, the type is "equipment feature unit," and specific parameters include model, rated flow rate, operating pressure, etc. The hierarchical position is level 2, and other related feature units include "dosing pump control cabinet," "chemical storage tank," etc. Organize the above information into structured data, including fields such as attribute name, attribute value, and data type.

[0160] Step S156: Organize the complete information of each effective association transmission path. The complete information of each effective association transmission path includes the feature unit sequence in each effective association transmission path, the semantic association edge sequence in each effective association transmission path, the association type of each semantic association edge in each effective association transmission path, the influence attenuation coefficient of each semantic association edge in each effective association transmission path, and the overall transmission strength value of each effective association transmission path.

[0161] For each path in the set of effective correlation propagation paths, its feature unit sequence, semantic correlation edge sequence, correlation type list, influence attenuation coefficient list, and overall propagation strength value are summarized to form a path detailed information package. For example, the detailed information of path P001 includes: feature unit sequence [A, B, C, D], semantic correlation edge sequence [E1, E2, E3], correlation type list [T1, T2, T3], influence attenuation coefficient list [0.8, 0.7, 0.6], and overall propagation strength value of 0.336.

[0162] Step S157: Associate and integrate the attribute information of the extracted root feature units with the complete information of all valid associated transmission paths, and organize the associated and integrated content according to the transmission order from the root feature units to the problem representation information mapping feature units.

[0163] Starting with the root cause feature unit, its attribute information is integrated with all valid associated transmission path information containing that root cause to form a radial structure centered on the root cause. For example, the path associated with the root cause feature unit "abnormal output of the dosing pump" includes P001, P003, etc. The information of the above paths is arranged in the transmission order (from root cause to terminal) to show the propagation path of the impact.

[0164] Step S158: Add a source tracing process description, which includes the basis for calculating the impact attenuation coefficient, the path screening criteria, and the method for determining the root cause characteristic unit, forming a complete document containing root cause characteristic unit information, effective associated transmission path information, and a source tracing process description, as the source tracing result for unconventional water source problems.

[0165] Write a description of the source tracing process, explaining the calculation methods for the attenuation coefficient (e.g., based on association type and semantic distance), the basis for setting the threshold for path selection (e.g., historical data statistics), and the rules for determining the root cause feature units (e.g., highest frequency and highest centrality). Combine the root cause feature unit information, effective association transmission path information, and process description to generate a PDF source tracing results report. The report includes sections such as summary, root cause analysis, transmission path visualization, and recommended measures to facilitate understanding and application by technical personnel.

[0166] Based on the same inventive concept, please refer to Figure 2 This paper shows a schematic block diagram of an unconventional water source problem tracing system 100 using a knowledge graph, provided in an embodiment of this application, for executing the above-described unconventional water source problem tracing method using a knowledge graph. The unconventional water source problem tracing system 100 using a knowledge graph may include a communication unit 110, a machine-readable storage medium 120, and a processor 130.

[0167] In this embodiment, both the machine-readable storage medium 120 and the processor 130 are located within the unconventional water source problem tracing system 100 utilizing a knowledge graph and are separately configured. However, it should be understood that the machine-readable storage medium 120 may also be independent of the unconventional water source problem tracing system 100 utilizing a knowledge graph and may be accessed by the processor 130 via a bus interface. Alternatively, the machine-readable storage medium 120 may also be integrated into the processor 130 and may communicate with external systems via the communication unit 110.

[0168] The processor 130 is the control center of the knowledge graph-based unconventional water source problem tracing system 100. It connects to various parts of the system via various interfaces and lines. By running or executing software programs and / or modules stored in the machine-readable storage medium 120, and by calling data stored in the machine-readable storage medium 120, it performs various functions and processes data of the knowledge graph-based unconventional water source problem tracing system 100, thereby providing overall monitoring of the system. Optionally, the processor 130 may include one or more processing cores; for example, the processor 130 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor. The machine-readable storage medium 120 is used to store machine-executable instructions for executing the scheme of this application, and the processor 130 is used to execute the machine-executable instructions stored in the machine-readable storage medium 120 to implement the unconventional water source problem tracing method using knowledge graphs provided in the aforementioned method embodiments.

[0169] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.

Claims

1. A method for tracing the origins of unconventional water source problems using knowledge graphs, characterized in that, The method includes: Obtain problem representation information of unconventional water sources, and map the problem representation information to a pre-constructed unconventional water source knowledge graph. The problem representation information includes abnormal water source quality, abnormal development process, abnormal equipment operation, and abnormal environmental correlation. The unconventional water source knowledge graph includes water source feature units, process feature units, equipment feature units, and environmental feature units, and each feature unit is connected by semantic association edges. Based on the mapping results between the problem representation information and the unconventional water source knowledge graph, a problem association transmission chain is constructed. The problem association transmission chain starts from the feature unit corresponding to the mapping and extends along the semantic association edge to other feature units in the knowledge graph. The association type of the feature units and semantic association edges passed through during the extension process is recorded. Calculate the influence attenuation coefficient of each semantic association edge in the problem association transmission chain, and generate an attenuation coefficient matrix containing the influence attenuation coefficient of each semantic association edge. The influence attenuation coefficient is determined based on the transmission strength corresponding to the association type and the semantic distance between feature units. Based on the attenuation coefficient matrix, the problem-related transmission chain is filtered to retain the transmission paths whose attenuation coefficients meet the transmission requirements, and to remove the transmission paths whose attenuation coefficients do not meet the transmission requirements, thus obtaining the effective related transmission paths. Based on the effective correlation transmission path, the starting feature unit is traced back to the effective correlation transmission path and the information of the root cause feature unit to be integrated to generate the source tracing result of the unconventional water source problem. The starting feature unit is the root cause feature unit of the problem.

2. The method for tracing the source of unconventional water source problems using knowledge graphs according to claim 1, characterized in that, The step of obtaining problem representation information of unconventional water sources and mapping the problem representation information to a pre-constructed knowledge graph of unconventional water sources includes: Anomalies related to water quality are collected through unconventional water source monitoring systems. These anomalies include changes in water source composition, turbidity, and odor. The collected anomalies related to water quality are then compiled into a list of water quality anomalies. Collect process anomaly data during the development of unconventional water sources. The process anomaly data includes stagnation in the development process, disordered connection between processes, and deviation in operation. The collected process anomaly data during the development of unconventional water sources is summarized as development process anomaly performance. Collect abnormal equipment operation data for unconventional water source development. The abnormal equipment operation data includes abnormal equipment start-up and shutdown behavior, equipment parameter fluctuation behavior, and abnormal equipment output behavior. The collected abnormal equipment operation data for unconventional water source development is transformed into abnormal equipment operation behavior. Monitor environmental anomalies around unconventional water source development areas. These anomalies include abnormal soil conditions, abnormal vegetation growth, and abnormal surrounding water bodies. The monitored environmental anomalies around unconventional water source development areas are integrated into environmentally correlated anomalies. The abnormal manifestations of water source quality, development process, equipment operation, and environmental correlation are combined to form problem characterization information; The system invokes a pre-built unconventional water source knowledge graph, in which water source feature units correspond to water source quality-related attributes, process feature units correspond to development process-related attributes, equipment feature units correspond to equipment operation-related attributes, and environmental feature units correspond to environmental association-related attributes. Extract key descriptive content from the problem representation information, and semantically match the extracted key descriptive content with the attribute descriptions of each feature unit in the pre-constructed unconventional water source knowledge graph to determine the feature units corresponding to the problem representation information; In the pre-constructed unconventional water source knowledge graph, the matching corresponding feature units are marked to complete the mapping of problem representation information to the pre-constructed unconventional water source knowledge graph. The semantic association edges between each feature unit are used to represent the attribute association relationship between different feature units.

3. The method for tracing the source of unconventional water source problems using knowledge graphs according to claim 1, characterized in that, Based on the mapping results between the problem representation information and the unconventional water source knowledge graph, a problem association transmission chain is constructed. This chain starts with the feature unit corresponding to the mapping and extends along semantic association edges to other feature units in the knowledge graph. The association types of the feature units traversed and the semantic association edges during the extension process are recorded, including: Select the core feature unit corresponding to each abnormal behavior from the feature units corresponding to the problem representation information mapping, and take the core feature unit corresponding to each abnormal behavior as the starting point of the problem association transmission chain; The query identifies all semantic association edges connecting the starting feature unit of the query problem's association transmission chain to the unconventional water source knowledge graph, and determines the next-level feature unit pointed to by each semantic association edge found. Identify the association type corresponding to each semantic association edge found in the query. The association types include attribute dependency association type, causal trigger association type, conditional constraint association type, and state linkage association type. Starting from the initial feature unit of the problem-related propagation chain, the extension is carried out along the next level feature unit pointed to by each semantic association edge found in the query, and the sequence of feature units passed through during the extension process and the association type of semantic association edges between each feature unit are recorded. For the next level feature unit reached by the extension, repeat the operation of querying the semantic association edge connected to the next level feature unit and the feature unit pointed to by the semantic association edge, and continue to extend along the semantic association edge until the terminal feature unit in the unconventional water source knowledge graph has no subsequent associated feature units. The feature unit sequence and corresponding association type formed by extending the starting feature unit of each problem association transmission chain are arranged in the extension order to form multiple initial transmission paths with the starting feature unit of the problem association transmission chain as the starting point and the ending feature unit as the ending point. The multiple initial conduction paths formed are deduplicated, and initial conduction paths with identical path structures are deleted, while initial conduction paths with unique path structures are retained. The initial transmission paths after deduplication are combined to form a problem-related transmission chain that covers all associated feature units. This ensures that the formed problem-related transmission chain contains the complete extension path from the starting feature unit to each terminal feature unit and the corresponding association type information.

4. The method for tracing the source of unconventional water source problems using knowledge graphs according to claim 1, characterized in that, The calculation of the influence attenuation coefficient of each semantically related edge in the problem-related propagation chain, and the generation of an attenuation coefficient matrix containing the influence attenuation coefficients of each semantically related edge, includes: Establish a correspondence table between association types and transmission strengths, where attribute dependency association types correspond to the first transmission strength, causal trigger association types correspond to the second transmission strength, conditional constraint association types correspond to the third transmission strength, and state linkage association types correspond to the fourth transmission strength. The transmission strengths corresponding to different association types have different fixed value ranges. Extract the transmission strength value corresponding to the association type of each semantic association edge in the problem association transmission chain from the established correspondence table between association type and transmission strength; The semantic distance between two feature units connected by each semantic association edge in the problem association transmission chain is calculated. The semantic distance is determined based on the difference in node level of the two feature units in the unconventional water source knowledge graph and the similarity of the attributes of the two feature units. The smaller the difference in node level of the two feature units in the unconventional water source knowledge graph and the higher the similarity of the attributes of the two feature units, the smaller the semantic distance between the two feature units. Define a semantic distance decay function, input the calculated semantic distance into the defined semantic distance decay function, and obtain the semantic distance decay value. The larger the calculated semantic distance, the larger the semantic distance decay value. The influence attenuation coefficient is calculated based on the extracted transmission intensity value and the obtained semantic distance attenuation value. The calculation method is to multiply the extracted transmission intensity value by (1 minus the obtained semantic distance attenuation value) to obtain the influence attenuation coefficient of each semantic association edge in the problem association transmission chain. The calculated impact attenuation coefficient is standardized, and the identification information of all semantic related edges in the problem-related transmission chain is statistically analyzed. The identification information of each semantic related edge is matched one-to-one with the corresponding standardized impact attenuation coefficient. Using the identifier information of the semantically related edges as row and column indices, the corresponding influence attenuation coefficients are filled into the corresponding positions in the matrix to generate an attenuation coefficient matrix containing the influence attenuation coefficients of each semantically related edge. The positions of the feature units that are not directly connected in the generated attenuation coefficient matrix are filled with zero values.

5. The method for tracing the source of unconventional water source problems using knowledge graphs according to claim 1, characterized in that, The step of filtering the problem-related transmission chains based on the attenuation coefficient matrix, retaining transmission paths whose influencing attenuation coefficients meet the transmission requirements, and removing transmission paths whose influencing attenuation coefficients do not meet the transmission requirements, yields effective related transmission paths, including: Analyze the number of semantically related edges contained in each initial transmission path in the problem-related transmission chain, and determine the set of semantically related edges corresponding to each initial transmission path after analysis; Extract the influence attenuation coefficient of each semantically associated edge from the set of semantically associated edges corresponding to each initial propagation path from the attenuation coefficient matrix; Calculate the product of the influence attenuation coefficients of all semantically related edges in each initial transmission path, and use the calculated product as the overall transmission strength value of that initial transmission path; Set an overall conduction intensity threshold, which is determined based on the accuracy requirements of tracing the source of unconventional water source problems and the effective path intensity range of historical tracing data; The overall conduction strength value of each initial conduction path is compared with the set overall conduction strength threshold. If the overall conduction strength value of an initial conduction path is greater than or equal to the set overall conduction strength threshold, the initial conduction path is marked as a candidate valid path; if the overall conduction strength value of an initial conduction path is less than the set overall conduction strength threshold, the initial conduction path is marked as an invalid path. Perform a path integrity check on the initial propagation path marked as a candidate valid path to confirm whether there are missing feature units or missing semantic association edges in the initial propagation path marked as a candidate valid path. If there are missing feature units or missing semantic association edges, fill them in and recalculate the overall propagation strength value of the initial propagation path. The overall conduction strength value of the completed candidate valid path is verified again to ensure that the overall conduction strength value of the completed candidate valid path still meets the requirement of being greater than or equal to the set overall conduction strength threshold. All initial propagation paths marked as invalid are removed, and the candidate valid paths that have been verified and meet the requirements are organized into a set of valid associated propagation paths.

6. The method for tracing the source of unconventional water source problems using knowledge graphs according to claim 5, characterized in that, The analysis determines the number of semantically related edges contained in each initial transmission path in the problem's correlation transmission chain, and establishes the set of semantically related edges corresponding to each initial transmission path, including: Traverse each initial propagation path in the problem-related propagation chain, and record the starting end feature unit identifier and the ending feature unit identifier of each initial propagation path. Along each initial propagation path traversed, the feature unit at the beginning of the initial propagation path is traversed sequentially to the feature unit at the end of the initial propagation path, and the identifier of each feature unit passed through in the initial propagation path and the identifier of the semantic association edge connecting adjacent feature units are recorded. The number of semantically related edge identifiers in each initial propagation path is counted, and the counted number is the number of semantically related edges contained in that initial propagation path; Arrange all semantically related edge identifiers recorded in each initial transmission path according to the extension order of that initial transmission path to form the semantically related edge sequence corresponding to that initial transmission path; Each initial propagation path is uniquely identified by its corresponding semantic association edge sequence. The semantic association edge sequence of each initial propagation path is associated with and stored in conjunction with the starting feature unit identifier of the initial propagation path, the ending feature unit identifier of the initial propagation path, and the number of semantic association edges contained in the initial propagation path. Check if there are duplicate semantic association edge identifiers in the semantic association edge sequence corresponding to each initial transmission path. If there are duplicate semantic association edge identifiers in the semantic association edge sequence corresponding to an initial transmission path, confirm whether the initial transmission path is a path loop. If the initial transmission path is a path loop, delete the semantic association edge identifiers of the path loop part of the initial transmission path and retain the semantic association edge identifiers of the non-loop part of the initial transmission path. The semantic association edge sequence corresponding to each initial propagation path after processing is formatted uniformly so that all processed semantic association edge sequences are presented with the same identifier format and arrangement order; The semantically associated edge sequence processed by each initial transmission path is used as the set of semantically associated edges corresponding to that initial transmission path, and a one-to-one correspondence table between the initial transmission path and the corresponding set of semantically associated edges is established.

7. The method for tracing the source of unconventional water source problems using knowledge graphs according to claim 5, characterized in that, The initial propagation path marked as a candidate valid path undergoes a path integrity check to confirm whether there are missing feature units or missing semantic association edges in the initial propagation path marked as a candidate valid path. If there are missing feature units or missing semantic association edges, the path is supplemented and the overall propagation strength value of the initial propagation path is recalculated, including: Extract the feature unit sequence of the initial propagation path labeled as a candidate valid path and the semantic association edge sequence of the initial propagation path labeled as a candidate valid path; The extracted feature unit sequence is compared with the feature unit library in the unconventional water source knowledge graph. It is checked whether each feature unit identifier in the extracted feature unit sequence exists in the feature unit library in the unconventional water source knowledge graph. If there is a feature unit identifier in the extracted feature unit sequence that does not exist in the feature unit library in the unconventional water source knowledge graph, then the feature unit identifier is marked as a missing feature unit. The extracted semantic association edge sequence is compared with the semantic association edge library in the unconventional water source knowledge graph. It is checked whether each semantic association edge identifier in the extracted semantic association edge sequence exists in the semantic association edge library in the unconventional water source knowledge graph. If there is a semantic association edge identifier in the extracted semantic association edge sequence that does not exist in the semantic association edge library in the unconventional water source knowledge graph, then the semantic association edge identifier is marked as missing semantic association edge. For candidate valid paths marked with missing feature units, query the feature unit attribute information corresponding to the feature unit identifier marked as missing, select the feature unit with the most matching feature unit attribute information from the feature unit library in the unconventional water source knowledge graph as a supplementary feature unit, and insert the selected supplementary feature unit into the corresponding missing position of the feature unit sequence of the candidate valid path. For candidate valid paths marked with missing semantic association edges, based on the attribute information of the feature units connected by the missing semantic association edges, the semantic association edge with the most matching association type is selected from the semantic association edge library in the unconventional water source knowledge graph as a supplementary semantic association edge, and the selected supplementary semantic association edge is inserted into the corresponding missing position of the semantic association edge sequence of the candidate valid path. After the supplementation is completed, a new sequence of feature units and a new sequence of semantically related edges are generated to form a complete candidate valid path; Extract the influence attenuation coefficient of each semantically associated edge from the new semantically associated edge sequence of the complete candidate effective path from the attenuation coefficient matrix; The influence attenuation coefficients of all semantically related edges in the completed candidate effective path are recalculated to obtain a new overall transmission strength value for the completed candidate effective path. The original overall transmission strength value of the candidate effective path is replaced with the new overall transmission strength value.

8. The method for tracing the source of unconventional water source problems using knowledge graphs according to claim 1, characterized in that, The process of tracing back to the starting feature unit based on the effective correlation transmission path, integrating the information of the effective correlation transmission path and the root cause feature unit, and generating the source tracing result for unconventional water source problems includes: For each valid association transmission path in the set of valid association transmission paths, perform a reverse traversal, backtracking from the terminal feature unit of each valid association transmission path to the starting feature unit of each valid association transmission path, and record the feature units passed during the backtracking process and the semantic association edges passed during the backtracking process; During the reverse traversal, the earliest feature unit that appears in each valid association propagation path and is not included in other valid association propagation paths is identified, and the identified feature unit is temporarily designated as the potential root feature unit. Collect all potential root cause feature units corresponding to all effective correlation transmission paths, and count the frequency of occurrence of each collected potential root cause feature unit in different effective correlation transmission paths; The most frequently occurring potential root cause feature unit is selected as the final starting feature unit, that is, the selected starting feature unit is determined to be the root cause feature unit of the problem. Extract the attribute information of the root source feature unit. The attribute information of the root source feature unit includes the type of the root source feature unit, the specific parameters corresponding to the root source feature unit, the hierarchical position of the root source feature unit in the unconventional water source knowledge graph, and other feature units associated with the root source feature unit. The complete information of each effective association transmission path is compiled, including the feature unit sequence in each effective association transmission path, the semantic association edge sequence in each effective association transmission path, the association type of each semantic association edge in each effective association transmission path, the influence attenuation coefficient of each semantic association edge in each effective association transmission path, and the overall transmission strength value of each effective association transmission path. The attribute information of the extracted root cause feature units is associated and integrated with the complete information of all effective associated transmission paths, and the associated and integrated content is organized according to the transmission order from the root cause feature units to the problem representation information mapping feature units. Add a source tracing process description, which includes the basis for calculating the impact attenuation coefficient, the path selection criteria, and the method for determining the root cause characteristic units. This forms a complete document containing root cause characteristic unit information, effective associated transmission path information, and a source tracing process description, serving as the source tracing result for unconventional water source problems.

9. The method for tracing the source of unconventional water source problems using knowledge graphs according to claim 8, characterized in that, The reverse traversal of each valid association propagation path in the set of valid association propagation paths, traversing from the terminal feature unit of each valid association propagation path back to the starting feature unit of each valid association propagation path, and recording the feature units traversed and the semantic association edges traversed during the backtracking process, includes: Obtain the forward feature unit sequence and the forward semantic association edge sequence of each effective association transmission path in the set of effective association transmission paths. The forward feature unit sequence and the forward semantic association edge sequence are arranged in order from the starting feature unit of the effective association transmission path to the ending feature unit of the effective association transmission path. Reverse the forward feature unit sequence of each valid associated propagation path to obtain the reverse feature unit sequence from the terminal feature unit of the valid associated propagation path to the starting feature unit of the valid associated propagation path. Reverse the forward semantic association edge sequence of each effective association transmission path to obtain the reverse semantic association edge sequence from the terminal feature unit side of the effective association transmission path to the starting feature unit side of the effective association transmission path. Starting with the first feature unit of the reverse feature unit sequence of each valid association propagation path, begin the reverse traversal of that valid association propagation path; Read each feature unit identifier in the reverse feature unit sequence of each valid associated propagation path in sequence, and record the name of the feature unit corresponding to each read feature unit identifier and the attribute summary information of the feature unit; For each valid association propagation path, read the semantic association edge identifier of each semantic association edge in the reverse semantic association edge sequence, and record the association type of the semantic association edge corresponding to each semantic association edge identifier and the influence attenuation coefficient of the semantic association edge. During the reverse traversal process, a reverse traversal record table is established for each valid association transmission path. The feature unit information in the reverse feature unit sequence of each valid association transmission path and the semantic association edge information in the reverse semantic association edge sequence of that valid association transmission path are filled into the reverse traversal record table of that valid association transmission path in a one-to-one correspondence according to the reverse traversal order. When the reverse traversal reaches the last feature unit of the reverse feature unit sequence of each valid associated propagation path, the reverse traversal of that valid associated propagation path is stopped. Repeat the above reverse traversal operation for all valid association transmission paths in the set of valid association transmission paths, and generate a corresponding reverse traversal record table for each valid association transmission path, so that the generated reverse traversal record table completely records the feature units and semantic association edge information passed during the backtracking process of the valid association transmission path.

10. A source tracing system for unconventional water source problems utilizing knowledge graphs, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; The processor is configured to execute the unconventional water source tracing method using a knowledge graph as described in any one of claims 1 to 9 by executing the machine-executable instructions.