Traffic event analysis method, device and equipment based on multi-hop causal path exploration

By constructing a knowledge graph and combining heuristic scoring functions and Monte Carlo tree search algorithm, the problem of inefficiency of existing causal reasoning methods in smart transportation systems is solved, and efficient and accurate exploration of multi-hop causal paths is achieved, and it is suitable for scenarios such as accident tracing, emergency decision-making and dynamic planning.

CN120372304APending Publication Date: 2025-07-25CHINESE SCI CLOUD COMPUTING ACAD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510285159.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing causal reasoning methods are difficult to adapt to the dynamic traffic environment in smart transportation systems and cannot effectively capture complex multi-hop causal chains, resulting in lack of flexibility and accuracy in reasoning results, especially inefficient in dealing with emergencies and real-time data flows.

Method used

Traffic event analysis method based on multi-hop causal path exploration, by constructing a knowledge graph, calculating the semantic correlation between nodes and problem semantic vectors, combining heuristic scoring functions and Monte Carlo tree search algorithm, chain reasoning is performed to generate multi-hop causal paths, dynamically optimized path expansion and decision support.

Benefits of technology

It significantly improves the efficiency and accuracy of causal path exploration, can effectively respond to the needs of multi-dimensional event analysis, is suitable for causal analysis and dynamic decision-making support for complex events in smart transportation systems, and is widely used in scenarios such as accident tracing, emergency decision-making and dynamic planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372304A_ABST
    Figure CN120372304A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of causal reasoning, and discloses a traffic event analysis method, device and equipment based on multi-hop causal path exploration, and the method comprises the steps: constructing a knowledge graph based on multi-source heterogeneous data in the traffic field; calculating the semantic correlation between each node in the knowledge graph and the question semantic vector, and taking the node with the highest semantic correlation as an initial reasoning node; and performing reasoning by starting from the initial reasoning node and combining a heuristic scoring function and a Monte Carlo tree search algorithm to obtain a multi-hop causal path. According to the method, a complex input problem can be effectively analyzed, rapid matching with the most relevant nodes of the problem is realized by utilizing the knowledge graph, a complex causal path is effectively identified and explored by combining chain reasoning and path optimization technologies, and the method is more efficient by combining a heuristic scoring function and a Monte Carlo tree search algorithm. And carrying out multi-dimensional event analysis and decision support on the basis of dynamic optimization. And the efficiency and accuracy of causal path exploration are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of causal reasoning, and in particular, to a traffic event analysis method, device, and equipment based on multi-hop causal path exploration. Background Art

[0002] With the continuous development of intelligent transportation systems, a large amount of multi-source heterogeneous data has been generated in the transportation field. This data not only involves the physical attributes of entities, but also contains complex event causal relationships and spatio-temporal characteristics, which are important resources for analyzing and optimizing traffic flow and improving event response capabilities. However, existing causal reasoning methods often rely on limited causal relationship assumptions and cannot adapt to the complexity of dynamic traffic environments. This limitation makes it difficult for existing methods to effectively capture complex multi-hop causal chains, especially when dealing with emergencies and real-time data streams, and the reasoning results lack flexibility and accuracy. Summary of the Invention

[0003] The main purpose of the present application is to provide a traffic event analysis method, device, and equipment based on multi-hop causal path exploration, aiming to solve the technical problem of inaccurate traffic event reasoning and analysis in the prior art.

[0004] The first aspect of the present application provides a traffic event analysis method based on multi-hop causal path exploration. The traffic event analysis method based on multi-hop causal path exploration includes:

[0005] Constructing a knowledge graph based on multi-source heterogeneous data in the transportation field;

[0006] Calculating the semantic relevance between each node in the knowledge graph and the problem semantic vector, and taking the node with the highest semantic relevance as the initial reasoning node, where the problem semantic vector is obtained by converting the user input problem;

[0007] Starting from the initial reasoning node, combining a heuristic scoring function and a Monte Carlo tree search algorithm to perform reasoning to obtain a multi-hop causal path, where the multi-hop causal path is a response to the user input problem.

[0008] The present application also provides a traffic event analysis device based on multi-hop causal path exploration. The traffic event analysis device based on multi-hop causal path exploration includes:

[0009] A knowledge graph construction module for constructing a knowledge graph based on multi-source heterogeneous data in the transportation field;

[0010] An initial node determination module for calculating the semantic relevance between each node in the knowledge graph and the problem semantic vector, and taking the node with the highest semantic relevance as the initial reasoning node, where the problem semantic vector is obtained by converting the user input problem;

[0011] A path expansion module, configured to start from an initial inference node, combine a heuristic scoring function and a Monte Carlo tree search algorithm, and perform inference to obtain a multi-hop causal path, where the multi-hop causal path is a response to a user input question.

[0012] A third aspect of the present application provides a computer device, including: a memory and at least one processor, where instructions are stored in the memory; the at least one processor invokes the instructions in the memory to cause the computer device to execute the above-mentioned traffic event analysis method based on multi-hop causal path exploration.

[0013] A fourth aspect of the present application provides a computer-readable storage medium, in which instructions are stored, and when the instructions are run on a computer, the computer is caused to execute the above-mentioned traffic event analysis method based on multi-hop causal path exploration.

[0014] The present application can effectively parse complex input questions, use a knowledge graph to achieve fast matching of the nodes most relevant to the questions, effectively identify and explore complex causal paths by combining chain reasoning, knowledge graph construction, and path optimization techniques, and perform multi-dimensional event analysis and decision support on the basis of dynamic optimization by combining a heuristic scoring function and a Monte Carlo tree search algorithm. The efficiency and accuracy of causal path exploration are significantly improved. The present application is particularly suitable for complex event causal analysis and dynamic decision support in intelligent transportation systems, can effectively meet the requirements of multi-dimensional event analysis, and is widely applied to scenarios such as accident cause tracing, emergency decision-making, and dynamic planning. It has broad application prospects in fields such as intelligent transportation, medical causal analysis, and risk assessment. Description of the Drawings

[0015] Figure 1 It is a schematic flowchart of the first embodiment of the traffic event analysis method based on multi-hop causal path exploration in the embodiment of the present application;

[0016] Figure 2 It is a schematic diagram of the functional modules of an embodiment of the traffic event analysis device based on multi-hop causal path exploration in the embodiment of the present application;

[0017] Figure 3 It is a schematic diagram of an embodiment of the computer device in the embodiment of the present application. Detailed Embodiments

[0018] In the description, claims and the above-mentioned drawings of this application, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the term "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0019] With the continuous development of intelligent transportation systems, a large amount of multi-source heterogeneous data has emerged in the transportation field. These data not only involve the physical attributes of entities, but also contain complex event causal relationships and spatio-temporal characteristics, and are important resources for analyzing and optimizing traffic flow and enhancing event response capabilities. However, existing causal reasoning methods often rely on static data models and limited causal relationship assumptions and cannot adapt to the complexity of dynamic traffic environments. This limitation makes it difficult for existing methods to effectively capture complex multi-hop causal chains, especially when dealing with emergencies and real-time data streams, and the reasoning results lack flexibility and accuracy.

[0020] Traditional methods also face the problem of low path reasoning efficiency. In complex traffic systems, the interaction effects between multi-dimensional data are complex, and path reasoning requires identifying potential causal chains in massive data. In addition, the generated causal paths also lack logical consistency and do not fully combine semantic relevance, data support, and dynamic optimization requirements. Therefore, the existing technologies cannot meet the high-efficiency requirements for complex event root cause analysis and dynamic decision-making support in intelligent transportation systems.

[0021] Based on this, this application provides a traffic event analysis solution based on multi-hop causal path exploration.

[0022] Refer to Figure 1 , in an embodiment of this application, a traffic event analysis method based on multi-hop causal path exploration is provided. The traffic event analysis method based on multi-hop causal path exploration includes:

[0023] S100: Construct a knowledge graph based on multi-source heterogeneous data in the transportation field.

[0024] Specifically, collect multi-source heterogeneous data in the transportation field, including real-time traffic flow data, event logs, weather data or meteorological data, road topology information, geographical information, etc. After standardizing these multi-source heterogeneous data, a structured knowledge graph (Knowledge Graph, KG) can be generated, denoted as KG = (V, E), where V is the set of nodes, including events (such as accidents) and entities (such as road segments), and E is the set of relationships, i.e., the set of edges, including causal relationships, spatio-temporal proximity relationships, etc.

[0025] Obtain real-time information such as road congestion, vehicle speed, and traffic accidents through sensors, cameras, or GPS devices. Meteorological data: Obtain meteorological information such as temperature, humidity, wind speed, and precipitation through weather stations or satellite remote sensing data. Geographical information: Extract geographical locations, terrain features, land use types, etc. from a geographic information system (GIS) or remote sensing data.

[0026] Semantically embed the nodes and edges in the knowledge graph to generate an initial semantic representation.

[0027] S200: Calculate the semantic relevance between each node in the knowledge graph and the problem semantic vector, and use the node with the highest semantic relevance as the initial inference node, where the problem semantic vector is obtained by converting the user input problem.

[0028] Specifically, semantically embed the nodes or edges in the knowledge graph to generate the corresponding initial semantic representation.

[0029] Convert the user input problem to generate a problem semantic vector.

[0030] Based on the semantic representation of the node and the problem semantic vector of the user input problem, the semantic relevance between the node and the user input problem can be calculated.

[0031] Select the node with the highest semantic relevance as the initial inference node v start 。

[0032] By screening the initial inference node, it is ensured that the inference process starts from the most relevant information point, improving the accuracy of causal inference.

[0033] S300: Starting from the initial inference node, combine the heuristic scoring function and the Monte Carlo tree search algorithm to perform inference to obtain a multi-hop causal path, where the multi-hop causal path is the answer to the user input problem.

[0034] Specifically, starting from the initial inference node, combining the heuristic scoring function and the Monte Carlo Tree Search (MCTS) algorithm, considering multiple factors comprehensively, the path quality is improved through stages such as selection, expansion, simulation, and backtracking, and chain reasoning, path expansion, and dynamic optimization are carried out to obtain the multi-hop causal path after the path expansion is completed.

[0035] Among them, the Monte Carlo Tree Search (MCTS) algorithm is mainly used to dynamically optimize the inference path. By using the Monte Carlo Tree Search (MCTS) algorithm, MCTS selects the optimal inference path by simulating and evaluating multiple paths.

[0036] After meeting the path expansion termination condition, the generated multi-hop causal path is:

[0037] Path = {v start , v1, v2, …, v n}

[0038] Among them, v start , v1, v2, …, v n are the initial inference node and the expansion nodes generated during the path expansion process respectively, including intermediate nodes and termination nodes.

[0039] According to the multi-hop causal path, the inference result can be output or converted into the final answer output to help users understand the causal relationship of complex traffic events.

[0040] More specifically, when the path expansion terminates, the final multi-hop causal chain is generated. Assume the multi-hop causal path or inference path is:

[0041] P = {v1, v2, …, v n}

[0042] Then the multi-hop causal chain can be expressed as: v1 → v2 → … → v n , where v1 is the initial node and v n is the termination node, and the arrow represents the causal relationship.

[0043] The final result output is the inference path dynamically optimized by the MCTS algorithm after the multi-hop causal chain is generated, and the multi-hop causal chain is gradually expanded. Starting from the initial node "traffic congestion", through multi-hop reasoning, the following causal chain is generated:

[0044] Traffic congestion → Traffic accident → Road slippery → Rainfall

[0045] Another is that according to the generated causal chain, the system can generate the final answer. For example: "The reason for the traffic congestion in a certain area may be due to a traffic accident, and the reason for the traffic accident may be that the road is slippery, and the reason for the road being slippery may be rainfall."

[0046] The key to multi-hop causal reasoning lies in constructing an efficient and accurate path expansion strategy. In practical applications, the selection of causal paths usually faces problems such as low node correlation and unreliable reasoning results. The chain reasoning method proposed in this application combines a heuristic scoring function with Monte Carlo Tree Search (MCTS), which can not only dynamically evaluate the importance of path nodes, but also optimize the path exploration process to achieve efficient expansion and accurate reasoning of causal paths.

[0047] This embodiment implements a multi-hop causal path exploration scheme based on chain reasoning. This scheme is particularly suitable for causal analysis and dynamic decision support of complex events in intelligent transportation systems, can effectively meet the needs of multi-dimensional event analysis, and is widely used in scenarios such as accident cause tracing, emergency decision-making, and dynamic planning.

[0048] Most traditional causal reasoning methods stay at single-hop analysis and are difficult to handle the multi-level causal relationship chains in complex events. This limitation is particularly obvious in traffic analysis with multi-source data, and the causal reasoning results lack context information and comprehensive relevance. This application generates semantic vectors through problem parsing, combines knowledge graph semantic embedding technology, and accurately matches the initial reasoning nodes most relevant to the problem, significantly improving the accuracy and interpretability of causal path construction.

[0049] This embodiment can effectively parse complex input problems, use the knowledge graph to quickly match the nodes most relevant to the problem, and through combining chain reasoning, knowledge graph construction, and path optimization technologies, effectively identify and explore complex causal paths. By combining the heuristic scoring function and the Monte Carlo tree search algorithm, multi-dimensional event analysis and decision support are carried out on the basis of dynamic optimization. The efficiency and accuracy of causal path exploration are significantly improved. This embodiment is particularly suitable for causal analysis and dynamic decision support of complex events in intelligent transportation systems, can effectively meet the needs of multi-dimensional event analysis, and is widely used in scenarios such as accident cause tracing, emergency decision-making, and dynamic planning. It has broad application prospects in fields such as intelligent transportation, medical causal analysis, and risk assessment.

[0050] In one embodiment, in step S300, starting from the initial reasoning node, combining the heuristic scoring function and the Monte Carlo tree search algorithm, performing reasoning to obtain a multi-hop causal path, including:

[0051] Based on the heuristic scoring function, calculate the heuristic scores of each candidate expansion node of the current node, where the current node is the initial reasoning node or an expansion node selected during the path reasoning process;

[0052] Based on the Monte Carlo tree search algorithm and the heuristic scores of the candidate expansion nodes, evaluate the comprehensive value scores corresponding to each candidate expansion node;

[0053] Select the candidate expansion node with the highest comprehensive value score as the next expansion node;

[0054] In response to the multi-hop causal chain reaching a preset path termination condition, generate a multi-hop causal path based on the multi-hop causal chain, where the multi-hop causal chain includes an initial inference node and each subsequent expansion node obtained through inference.

[0055] Specifically, there are association relationships between nodes in the knowledge graph. As a parent node, it can include one or more adjacent child nodes, and the child nodes also belong to neighbor nodes.

[0056] Based on the knowledge graph, candidate expansion nodes of the current node can be obtained. Among them, the candidate expansion nodes are all or part of the child nodes of the current node, and the current node is the parent node of its candidate expansion nodes. The current node is an initial inference node or an expansion node obtained during path expansion. The expansion nodes also belong to inference nodes.

[0057] The heuristic score can be calculated through a heuristic scoring function, and the heuristic scoring function can include one or more scoring metrics. For example, it includes one or more of semantic relevance, node support, logical consistency score, and adjacency, etc.

[0058] For example, substituting the semantic relevance between the candidate expansion node and the user input question, the node support of the candidate expansion node, the logical consistency score, and the corresponding node scoring weights into the heuristic scoring function, the heuristic score of the candidate expansion node can be obtained. Among them, the logical consistency score is used to indicate the semantic consistency between the candidate expansion node and the current node (parent node).

[0059] The heuristic score is the node importance or the comprehensive score.

[0060] Based on the comprehensive scoring function in the Monte Carlo Tree Search (MCTS) algorithm and the heuristic scores of the candidate expansion nodes, the path transition probabilities and comprehensive value scores of different candidate sub-paths from the current node to the candidate expansion nodes can be evaluated. Among them, the candidate sub-path is the path from the current node to the candidate expansion node; select the candidate expansion node corresponding to the highest comprehensive value score as the expansion node.

[0061] After path expansion is completed, according to the initial inference node and the expansion nodes obtained during the path expansion process, the multi-hop causal path after path expansion termination is obtained.

[0062] In this embodiment, by integrating chain reasoning, heuristic scoring function, and Monte Carlo Tree Search (MCTS) algorithm, this method realizes the efficient construction and dynamic optimization of complex causal paths. First, the heuristic scoring function is used to comprehensively evaluate the semantic relevance, logical consistency, and data support degree of path nodes, and on this basis, the reasoning path is dynamically extended. Secondly, combined with the MCTS algorithm to optimize path selection, based on path transition probability and value function, a causal chain with logical consistency and high interpretability is generated. This method can adapt to the dynamic changes of uncertainty and real-time data in the traffic system, and has broad application prospects in scenarios such as accident cause tracing, emergency decision-making, and dynamic planning, which can significantly improve the event analysis efficiency and intelligent level of the intelligent transportation system.

[0063] In one embodiment, before evaluating the comprehensive value scores corresponding to each candidate expansion node based on the Monte Carlo Tree Search algorithm and the heuristic scoring of candidate expansion nodes, it further includes:

[0064] Take multiple child nodes with the top heuristic scores among the child nodes adjacent to the current node as candidate expansion nodes, where the number of candidate expansion nodes is less than the number of child nodes of the current node.

[0065] Specifically, the candidate expansion nodes in this embodiment are the preset number or preset proportion of child nodes with the highest heuristic scores selected from all the child nodes of the current node.

[0066] For example, if the current node has 10 child nodes, the top 5 child nodes with the highest heuristic scores can be selected as candidate expansion nodes. Or, select the top 50% of the child nodes with the highest heuristic scores as candidate expansion nodes. Subsequently, calculate the comprehensive value scores of these 5 candidate expansion nodes.

[0067] When facing a graph with high connectivity, first perform pre-screening (such as Top-K, etc.), and reduce the candidate expansion nodes of the current node to a few more promising candidate expansion nodes (such as 3-5). Among them, the Top-K screening calculates the heuristic score h(v) for each neighbor node (child node) of the current node and scores them. Sort them in descending order of score, and only keep the top K neighbors with the highest scores as candidate expansion nodes. For example: if K = 3, only select the top 3 nodes with the highest scores to enter the subsequent search, and the other 5 nodes will not be expanded.

[0068] This embodiment introduces a node screening mechanism as a supplementary condition for path expansion. It can screen out candidate expansion nodes through heuristic scoring, limit the number of candidate nodes, and increase the efficiency of subsequent expansion.

[0069] Method for Exploring Multi-Hop Causal Paths Based on Chain Reasoning. By integrating chain reasoning, heuristic scoring functions, and the Monte Carlo Tree Search (MCTS) algorithm, the method of this application achieves the efficient construction and dynamic optimization of complex causal paths. First, a heuristic scoring function is used to comprehensively evaluate the semantic relevance, logical consistency, and data support of path nodes, and based on this, the reasoning path is dynamically extended. Secondly, the MCTS algorithm is combined to optimize path selection, and based on path transition probabilities and value functions, a causal chain with logical consistency and high interpretability is generated. This method can adapt to the uncertainties and dynamic changes of real-time data in the transportation system and has broad application prospects in scenarios such as accident cause tracing, emergency decision-making, and dynamic planning, which can significantly improve the event analysis efficiency and intelligent level of the intelligent transportation system.

[0070] In one embodiment, determining that the multi-hop causal chain reaches the preset path termination condition is performed through the following steps:

[0071] If the path length of the current path reaches the path length limit, stop the path extension of the current path, where the current path includes the initial reasoning node and the extended nodes selected during path extension;

[0072] Or,

[0073] If the heuristic score of the current node is lower than the first set threshold, stop the path extension of the current path where the current node is located;

[0074] Or,

[0075] If the semantic relevance of the current node exceeds the second set threshold, stop the path extension of the current path where the current node is located.

[0076] Specifically, the reasonable design of the path termination condition is the key to ensuring the reasoning efficiency and result accuracy. This embodiment combines the dual criteria of path length and the heuristic score of the node to dynamically adjust the termination condition to adapt to the complexity and accuracy requirements of different reasoning tasks. At the same time, by generating a causal chain and logically verifying the path results, the reliability of the reasoning results is further improved.

[0077] During the path extension process, this embodiment sets the following 3 termination conditions to ensure that the reasoning process can be efficiently completed:

[0078] Maximum path length limit: When the path length reaches the predefined path length limit or upper bound L max stop the extension to avoid a sharp increase in computational complexity or the generation of meaningless reasoning results due to an overly long path.

[0079] Or, when the current path reaches the maximum number of hops (hop count), stop the extension.

[0080] Length(P) = Number of hops in path P Length(P) ≥ L max , then the path extension terminates.

[0081] Heuristic scoring threshold: The heuristic score H(v) of the current node is lower than the first set threshold Q min , or, H(v) < H th , then stop further expanding this path to avoid interference of invalid nodes on the inference result. The first set threshold is the comprehensive scoring threshold H th , which is used to determine whether the comprehensive score of the inference path is high enough. The comprehensive score is the heuristic score and can be calculated through the heuristic scoring function H(v). If the comprehensive score of the current node is lower than the comprehensive scoring threshold, the path extension terminates.

[0082] Or, if the target node is found in the current path, stop expanding. The target node is a node whose semantic relevance exceeds the second set threshold.

[0083] If a node in the inference path matches the target node (find a node directly relevant to the problem), the path extension terminates. The target node can be judged by the semantic relevance score S(q, v):

[0084] S(q, v) ≥ S th

[0085] where S th is the second set threshold or the semantic relevance scoring threshold.

[0086] In this embodiment, by setting conditions such as path length limit, comprehensive scoring threshold, and target node matching, the termination timing of the inference path extension is determined.

[0087] This embodiment can not only dynamically control the termination condition of the inference process, but also output a causal chain that conforms to logical rules, providing reliable support for the causal inference of complex problems.

[0088] In one embodiment, the comprehensive value score is calculated through the following formula 1:

[0089]

[0090] Or,

[0091] The comprehensive value score is calculated through the following formula 2:

[0092]

[0093] Where:

[0094] Q(s,a) represents the expected return of transferring to the candidate expansion node s' by performing action a at the current node s, and can be used as a comprehensive value score;

[0095] R(s,a) is the immediate reward, which is equal to the heuristic score H(v) of the candidate expansion node s';

[0096] P(s′|s,a) is the path transition probability, representing the probability of transferring from the current node s to the candidate expansion node s' through action a;

[0097] V(s′) is the value function of the subsequent state, obtained by aggregating the heuristic scores of the children nodes of the candidate expansion node s' by taking the mean;

[0098] λ is the discount factor, which controls the impact of future rewards on the current value;

[0099] UCT(s,a) is the upper confidence bound formula and can be used as a comprehensive value score;

[0100] N(s,a) is the number of times action a has been visited;

[0101] c is the exploration constant;

[0102] b represents all candidate actions at the current node s;

[0103] ∑ b N(s,b) represents the total number of visits of all executable actions b at the current node s.

[0104] Specifically, path expansion optimizes path selection through the Monte Carlo Tree Search (MCTS) algorithm, and the comprehensive value scoring function used for the comprehensive value score is shown in Formula 2 or Formula 3.

[0105] s and s' represent the relationship between the current node (current state) and the next node (child node or next state) in the search tree / knowledge graph. a is the action taken to decide which next node (or which edge) to jump to starting from node s. When the state is simply understood as a node in the knowledge graph, s is the current node, a is the choice starting from the current node, and s' is the next node or child node of s. s″ is the next-next node or grandchild node of s.

[0106] When calculating V(s′), it is obtained by aggregating the heuristic scores of all child nodes s″ (belonging to the grandchild nodes of s) of the candidate expansion node s' by taking the mean.

[0107] The path transition probability P(s′|s,a) is usually allocated according to the heuristic score. By calculating the proportion of the heuristic score of each child node s′ relative to the sum of the heuristic scores of all child nodes, the path transition probability can be obtained. For example, the current node s has 3 child nodes s′ 1. s ′ 2. s ′ 3. If the heuristic scores of these three child nodes are h(v)1, h(v)2, and h(v)3 respectively, then the corresponding path transition probabilities are as follows:

[0108]

[0109]

[0110] This method allows child nodes with higher scores to have higher transition probabilities, so that these child nodes with higher scores and their corresponding sub-paths are more likely to be selected during the search.

[0111] Q(s,a) is used to measure the expected return of performing action a at the current node s, providing a basis for MCTS in the selection or final decision-making phase. If exploration is not considered, the combined expectation or combined value score of the current value + future value is used for selection. When finally implemented, the path corresponding to the action with the highest Q(s,a) value under the root node is usually selected.

[0112] For candidate expansion nodes, the Monte Carlo Tree Search (MCTS) algorithm can calculate and compare Q(s,a) + the exploration term using the selection strategy.

[0113] The action selection (Selection) combines the UCB (Upper Confidence Bound) balancing strategy and uses UCT(s,a) as the combined value score, which can avoid only greedily selecting the branch with the largest Q(s,a), leaving room for exploration of branches with few visits but high potential.

[0114] The child node or candidate expansion node with the highest combined value score is used as the expansion node to continue the downward search or expansion in the Monte Carlo tree search.

[0115] As can be seen from the above, the path reasoning in this embodiment includes: Selection phase: Starting from the current node, select the action with the highest combined value and gradually expand the reasoning path. Expansion phase: When an unexplored node is encountered, expand new child nodes and calculate their heuristic scores. Simulation phase: Starting from the expansion node, randomly simulate the reasoning path until the termination condition is met. Backtracking phase: Update the value function and transition probability of each node in the path according to the simulation results.

[0116] In this embodiment, the Monte Carlo tree search algorithm is used to dynamically expand the path using the above combined value scoring function, and the exploration and exploitation strategies can be combined to balance the comprehensiveness and efficiency of the path search.

[0117] In one embodiment, the heuristic score is calculated by the following formula 3:

[0118] H(v) = α·S rel (v) + β·S sup (v) + γ·S con (v)

[0119] Formula 3

[0120] Wherein, H(v) is the heuristic score of the node, and S rel (v), S sup (v), and S con (v) respectively represent the three scoring metrics of the semantic relevance, support score, and logical consistency score of the node, and α, β, and γ are the node scoring weights for different scoring metrics.

[0121] Specifically, to achieve effective expansion of the path, the present application designs a heuristic scoring function and a path exploration algorithm based on MCTS.

[0122] The heuristic scoring function evaluates the importance of the node by integrating multiple scoring metrics.

[0123] In this embodiment, other auxiliary scoring metrics (such as the support score and logical consistency score of the node) are introduced on the basis of the semantic relevance score to form a heuristic scoring function (comprehensive scoring function).

[0124] S rel (v): Semantic relevance or semantic relevance score, used to quantify the semantic similarity between node v and the user input problem to be inferred.

[0125] S sup (v): Support score or node support, indicating the support of the node in the knowledge graph (such as the importance or confidence of the node), which can be calculated by the degree centrality or prior knowledge of the node. Reflects the importance of the node in the knowledge graph, and the calculation formula is:

[0126]

[0127] Wherein, in-degree(v) represents the in-degree of node v, is the in-degree of the node v with the maximum in-degree in the graph ′ of the in-degree.

[0128] S con (v): Logical consistency score, measuring the semantic consistency between the current node and the parent node or representing the logical consistency between the node and the current inference path, which can be calculated by rule matching or graph neural network:

[0129]

[0130] Wherein, Sim() is the cosine similarity function, and Node semantic vectors for the parent node and the current node respectively.

[0131] The parameters α, β, and γ represent the weights of three different scoring metrics, namely semantic relevance, support score, and logical consistency score, which belong to edge weights or marginal weights and can be determined through cross-validation.

[0132] In one embodiment, the traffic event analysis method based on multi-hop causal path exploration further includes:

[0133] If there are path node pairs with conflicting causal relationships and / or path node pairs where the causal relationship is inconsistent with the predefined logical rules in the knowledge graph in the multi-hop causal path, then update the node scoring weights of the target node in the knowledge graph, where the node scoring weights include at least one of the node scoring weights corresponding to semantic relevance, the node scoring weights corresponding to support score, and the node scoring weights corresponding to logical consistency score;

[0134] Re-optimize the path selection according to the updated node scoring weights of the target node.

[0135] Specifically, result verification and optimization: Perform logical consistency verification on the generated multi-hop causal path or multi-hop causal path chain to check for contradictions or unreasonableness. If there are contradictions or unreasonableness, the node and edge weights can be updated through a feedback mechanism to optimize the inference result.

[0136] In multi-hop causal path reasoning, the logical consistency and accuracy of the generated path are crucial. Path verification aims to ensure that the inference result conforms to the causal relationship rules in the knowledge graph, while optimization improves the inference efficiency by adjusting the scoring weights and path selection strategies. This application realizes dynamic weight adjustment and result optimization through the combination of path logical verification and feedback mechanism.

[0137] Verifying the logical consistency of the path is the first step in optimizing the inference result. The specific methods include:

[0138] Contradictory path detection: Check whether there are conflicting causal relationships in the path. If the path contains v i →v j and v j →v i , then it is determined as a contradictory path and the rationality of this path needs to be re-evaluated. The mathematical representation is as follows:

[0139] v i →v j ,v j →v i represents a contradictory path.

[0140] Causal consistency verification: Combine the predefined causal rules in the knowledge graph to check whether the relationship between each pair of nodes in the path conforms to the logical relationship in the graph. If the causal relationship in a certain path is inconsistent with the rules in the knowledge graph, it is marked as a potential error path.

[0141] Perform logical consistency verification on the generated causal path chain or causal path to check for contradictions or unreasonableness. Specific verification methods include: Rule matching, which uses predefined logical rules (such as "Traffic accidents do not directly cause rainfall") to check whether the causal relationship in the path is reasonable. Graph Neural Network (GNN) reasoning, which uses a graph neural network to reason about the nodes and edges in the path to judge its logical consistency. Manual review: Submit the generated causal path to domain experts for manual review to ensure its reasonableness and accuracy. Of course, this application is not limited to the above verification methods.

[0142] Path optimization and feedback mechanism: For the contradictions or inconsistent paths found in the verification, this application optimizes the reasoning process through a feedback mechanism, including the following steps:

[0143] Weight adjustment: Update the weight value of the edge in the knowledge graph, and the adjustment formula is as follows:

[0144] w ij =w ij +Δw

[0145] where: w ij represents the edge weight between nodes v i and v j ; Δw is the weight adjustment amount, and its value can be positive or negative: a positive value indicates increasing the edge weight and strengthening the causal relationship between nodes, and a negative value indicates decreasing the edge weight and weakening the causal relationship between nodes.

[0146] For example: w 交通事故,道路湿滑 =w 交通事故,道路湿滑 -Δw

[0147] Path re-selection: According to the updated node scoring weights, the updated heuristic scores can be obtained, and the nodes in the path expansion process are re-ordered and selected to ensure that the finally generated path conforms to logical consistency and has a higher comprehensive score. By dynamically adjusting the scoring parameters and weight values in the path reasoning process, the generation of invalid or incorrect paths is avoided. At the same time, combined with the feedback mechanism, the adaptability of the knowledge graph and the reasoning algorithm is continuously improved, providing efficient and reliable technical support for the causal analysis of complex problems.

[0148] In this embodiment, logical consistency verification is performed on the generated path to check whether the causal relationship between path nodes conforms to the established rules in the knowledge graph. If there are contradictions in the path, the node scoring weights α, β, and γ are adjusted through the feedback mechanism, and the path selection is re-optimized to ensure that the output causal chain has logical consistency and high interpretability.

[0149] After the path is generated in this embodiment, the generated causal path is verified using a logical consistency verification module to ensure that there are no conflicts in the causal relationship in the path. When a conflicting path is found, the path weight is adjusted through the feedback mechanism, and the inference result and path selection are re-optimized. In this embodiment, through logical consistency verification and the feedback mechanism, multi-way verification is performed on the causal path, and when unreasonable points are found, the weights are updated through the feedback mechanism. By focusing on result verification and optimization, it is ensured that the generated results are accurate, reliable, and easy to understand.

[0150] This embodiment introduces edge weights as supplementary conditions for path expansion. By dynamically adjusting the node scoring weights, this solution can adapt to the inference requirements in different problem scenarios. In practical applications, this method can quickly expand multi-hop causal paths with clear logic and strong relevance, providing accurate inference support for complex fields such as traffic analysis.

[0151] In one embodiment, the semantic relevance between each node in the knowledge graph and the problem semantic vector is calculated, including:

[0152] Convert the user input problem into a problem semantic vector;

[0153] Calculate the cosine similarity between the node semantic vector of each node in the knowledge graph and the problem semantic vector, where the cosine similarity is used to quantify the semantic relevance between the node and the user input problem.

[0154] Specifically, problem parsing and initial node selection: For the user input problem (such as "What is the reason for the congestion in a certain area?"), use an embedding model to generate the problem semantic vector q = g(query) of the user input problem, or q = Embed(q).

[0155] Among them, g(·) represents the embedding function, and the embedding model embedding is used to capture the semantic features of the problem and embed them into a high-dimensional space to represent the vector q. Embed(·) can be a pre-trained language model (such as BERT, Sentence-BERT, etc.) used to capture the semantic information of the problem. This step provides a basis for subsequent node relevance matching.

[0156] In addition, the semantic vector generation process of problem parsing can adjust the embedding model according to specific application scenarios. For example, in the field of traffic analysis, a pre-trained traffic domain-specific model can be introduced.

[0157] Immediately followed by the calculation of node semantic relevance. In the knowledge graph, each node v ∈ V has a corresponding semantic embedding vector v. By calculating the semantic relevance between the problem semantic vector q and each node vector v, the most relevant node is selected as the initial reasoning node.

[0158] Through the following formula 4, the cosine similarity can be used to calculate the semantic relevance between each node in the knowledge graph and the problem semantic vector:

[0159]

[0160] Or,

[0161]

[0162] Where, is the problem semantic vector, is the semantic embedding vector of node v. The cosine similarity quantifies their semantic relevance by measuring the angle between the two vectors.

[0163] q·v represents the dot product of the problem semantic vector and the node semantic vector. ∥q∥ and ∥v∥ represent the norms of the problem semantic vector and the node semantic vector respectively. The formula calculates the cosine similarity between the problem semantic vector and the node semantic vector, with a value range of [-1, 1]. The larger the value, the higher the semantic relevance.

[0164] Initial reasoning node selection: In the knowledge graph node set V, the node with the highest semantic relevance score is selected as the initial reasoning node v start :

[0165]

[0166] Or,

[0167] The initial reasoning node v * can be determined by the following formula:

[0168]

[0169] Where, v * is the node most relevant to the problem semantics and serves as the starting point for causal reasoning.

[0170] This process ensures that the initial node closest to the semantics of the user input question can be matched, ensuring that the reasoning path starts from the most relevant information starting point and minimizing the reasoning deviation.

[0171] For example, in the causal reasoning process, it starts from the initial reasoning node v * and proceeds along the causal relationship edges in the knowledge graph. In the above formula, v *It is a "traffic congestion" node, and its possible cause nodes (such as "traffic accident", "road construction", etc.) can be found through causal relationship edges. For each cause node, its semantic relevance to the problem vector can be further calculated to determine the most likely cause.

[0172] In this embodiment, by screening the initial inference nodes, it is ensured that the inference process starts from the most relevant information points, improving the accuracy of causal inference.

[0173] In one embodiment, based on multi-source heterogeneous data in the traffic field, a knowledge graph is constructed, including:

[0174] Perform standardization processing on the multi-source heterogeneous data;

[0175] Perform word segmentation and part-of-speech tagging on the obtained standardized data in sequence, and then perform named entity recognition and relationship extraction;

[0176] By eliminating redundancy and conflicts, the obtained node entities and the relationships between the node entities are fused and linked to the corresponding nodes in the knowledge graph, where the node entities include entities and events, and the nodes include entity nodes and event nodes.

[0177] Specifically, in the intelligent transportation system, the causal analysis of traffic events requires the integration of multi-source heterogeneous data, such as real-time traffic flow data, historical event logs, weather data or meteorological data, geographic information, etc. These data often contain structured and unstructured information, and it is difficult to effectively construct a causal model by directly applying traditional data analysis methods. In this embodiment, through unified modeling technology, multi-source data is embedded in the knowledge graph to provide semantic support for subsequent causal inference.

[0178] First, data collection is carried out, and the collected multi-source heterogeneous data includes, but is not limited to: real-time traffic flow, road construction records, meteorological data, and historical traffic event logs.

[0179] Due to the heterogeneity of multi-source data, it is necessary to perform standardization processing on the data to ensure the consistency and availability of the data. Data preprocessing or standardization processing:

[0180] Missing value filling, for example, using the Gaussian interpolation method;

[0181] Outlier removal, for example, using the three-standard-deviation rule to clean abnormal data;

[0182] Text standardization: Combine a natural language processing (NLP) model to generate semantic vectors of event descriptions, and generate semantic vectors through natural language processing of text descriptions.

[0183] Knowledge graph KG=(V,E) construction:

[0184] KG = {V, E}, V = {v1, v2, …, v n}, E = {(v i , v j )}

[0185] Where:

[0186] V: represents the set of nodes, including event nodes and entity nodes;

[0187] E: represents the set of relationships, including causal relationships and proximity relationships;

[0188] Node representation (node semantic vector):

[0189] v i = f(x i ; θ), x i is the node feature, and f(x i ; θ) is a deep learning embedding function, such as GNN.

[0190] Each node (event or entity) may have several numerical attributes (such as traffic flow, weather type, construction time, etc.), and may also have text descriptions (such as event brief, construction notice, etc.).

[0191] The above standardization process ensures the consistency and accuracy of the data, providing a reliable basis for the construction of the knowledge graph.

[0192] x i (Node feature): is a single vector formed by splicing after cleaning, standardizing, and vectorizing from multi-source heterogeneous data (numerical, text, etc.).

[0193] More specifically, in the intelligent transportation knowledge graph, each node corresponds to an "event" or "entity" (such as a certain traffic accident, a certain road, a certain weather condition, etc.).

[0194] Integrate all available information (numerical attributes + text description) of the node into a feature vector x i , and the specific steps are as follows:

[0195] 1. Text feature transformation:

[0196] For the text description associated with the node (such as "The accident occurred on the highway, causing two vehicles to rear-end", etc.), use the same pre-trained NLP model (such as BERT) to extract its semantic vector. Suppose a text vector with a dimension of dtd_t is obtained.

[0197] 2. Numerical feature standardization:

[0198] For the numerical attributes associated with nodes (such as traffic flow, temperature, rainfall, construction duration, etc.), perform normalization or standardization (for example, transform them to be between 0 and 1 or follow a standard normal distribution) to obtain a numerical vector of dimension \(d_n\).

[0199] 3. Concatenate and obtain \(x\). i :

[0200] Concatenate the text vector and the numerical vector dimensionally (concatenate) to form the comprehensive feature representation of the node:

[0201] \(x_i = [\text{text vector} d_t, \text{numerical vector} d_n] \in \mathbb{R}^{(d_t + d_n)}\). \(x_i=\left[\underbrace{\text{text vector}}{d_t},\;\underbrace{\text{numerical vector}}{d_n}\right]\quad\in\mathbb{R}^{(d_t + d_n)}\).

[0202] Through the above steps, \(x\) i is the single vector representation formed after cleaning, feature extraction, and integration of the node at the raw data level.

[0203] \(\theta\) (parameters of the embedding function): Learnable parameters obtained by training the known causal or associative relationships between nodes in the graph neural network.

[0204] \(f(x\) i ; \(\theta\)) (embedding function): Used to map the node features to a low-dimensional vector space to capture the semantic structure information of the nodes in the knowledge graph.

[0205] For missing data, methods such as Gaussian Process Regression or Mean Imputation are used for filling. The formula for Gaussian Process Regression is:

[0206]

[0207] where \(m(x)\) is the mean function and \(k(x, x')\) is the covariance function.

[0208] Outlier removal: Identify and remove outliers based on the three-sigma rule (3-Sigma Rule).

[0209] Assume the data follows a normal distribution

[0210]

[0211] Then the decision condition for outliers is:

[0212] \(\vert x\)i -μ|>3σ

[0213] where μ is the mean and σ is the standard deviation.

[0214] Convert data from different sources into a unified format (such as JSON, CSV, or RDF) for subsequent processing and analysis. Tokenize, perform part-of-speech tagging, named entity recognition (NER), etc. on text description information (such as event logs, social media text), and generate semantic vectors. Use the Word2Vec model to generate word vectors:

[0215] v w = Word2Vec(w)

[0216] where v w is the vector representation of word w.

[0217] Knowledge graph construction and optimization:

[0218] Use the standardized data to construct a structured knowledge graph (Knowledge Graph, KG). The formal representation of the knowledge graph is

[0219] KG = (V, E)

[0220] where:

[0221] Node set V: includes event nodes and entity nodes.

[0222] Relationship set E: describes the association relationships between nodes, including: causal relationships, spatio-temporal proximity relationships, and attribute relationships.

[0223] First is entity recognition and linking, that is, extract node entities from the data through entity recognition technology and link them to existing nodes in the knowledge graph. Next is relationship extraction, extract the relationships between entities from the data through rule matching, TransE. The relationship representation formula of the TransE model is:

[0224] h + r ≈ t

[0225] where h is the head entity vector, r is the relationship vector, and t is the tail entity vector. Finally, it is knowledge fusion, fuse the knowledge from different data sources, eliminate redundancy and conflicts, and ensure the consistency and accuracy of the knowledge graph.

[0226] To learn node representations on the knowledge graph, the graph neural network (Graph Neural Network, GNN) method can be adopted:

[0227] First, utilize the node's own feature x iInitialize node representations. In each layer of the GNN, nodes interact and aggregate information with their neighboring nodes. After stacking multiple layers of GNNs, richer node embeddings can be obtained.

[0228] The model parameters $\theta$ include:

[0229] Parameters of the input mapping layer: Map $x$ i to the initial hidden layer of the model (such as the weights and biases of a linear transformation).

[0230] Parameters of the GNN aggregation and update layer: The convolution or message passing weights for each layer in the graph neural network.

[0231] Parameters of the output layer: Used to obtain the final node embedding vector or node semantic vector $v_i$.

[0232] Training data:

[0233] Utilize the associations (causal relationships, spatio-temporal proximity relationships, etc.) between known traffic events as training signals. Partially manually annotated causal or correlation labels (such as the causal label "rainy days cause congestion") can also be introduced to enable the model to learn to capture causal clues.

[0234] Loss function: In knowledge graph tasks, a common approach is link prediction or causal relationship classification:

[0235] Link prediction: Given two nodes, predict whether there is a specific relationship (such as "causal" or "proximity") between them.

[0236] Causal relationship classification: Given a pair of nodes, predict the causal direction or relationship type between them.

[0237] During training, continuously update $\theta$ through backpropagation so that the node representations learned by the model can achieve better performance on these tasks.

[0238] Embedding result: When the model training is completed, the embedding function

[0239] $f(x_i;\theta):\mathbb{R}^{(d_t + d_n)}\to\mathbb{R}^k$. For each node $v_i$, we can calculate its final low-dimensional vector representation $v_i = f(x_i;\theta)$, where $k$ is the dimension of the GNN output layer (such as 128, 256, etc.).

[0240] The meaning of $x_i$: It is obtained by cleaning and fusing single or multiple pieces of raw data (numerical values, text, etc.). It is the integrated vector of all information describing the node in the input layer.

[0241] Meaning of \(\theta\): The total sum of learnable parameters such as weights and biases in a graph neural network (or other deep learning models).

[0242] During the training process, it is obtained by iteratively updating through minimizing the loss function and using labeled data or structural information.

[0243] Final goal: To enable the node representation \(v_i\) to better identify and express causal relationships in downstream tasks such as causal path exploration and chain reasoning. For example, determining whether a road construction causes congestion in the nearby area, or whether a certain extreme weather leads to more accidents during peak hours, etc.

[0244] Summary:

[0245] x i (Node feature): A single vector formed by splicing after cleaning, standardizing, and vectorizing from multi-source heterogeneous data (numerical, text, etc.).

[0246] \(\theta\) (Embedding function parameter): Learnable parameters obtained by training the known causal or associative relationships between nodes in a graph neural network.

[0247] \(f(x_i;\theta)\) (Embedding function): Maps the node feature to a low-dimensional vector space to capture the semantic structure information of the node in the knowledge graph.

[0248] In one embodiment, based on multi-source heterogeneous data in the traffic field, constructing a knowledge graph further includes:

[0249] Verifying the nodes and relationships in the constructed knowledge graph based on automated rules;

[0250] And / or,

[0251] Updating the knowledge graph according to incremental data.

[0252] Specifically, knowledge graph optimization:

[0253] To ensure the reliability of the knowledge graph, the following measures are taken: Conduct data verification, verify the nodes and relationships in the knowledge graph through manual review or automated rules to ensure its correctness. Next is knowledge update: Regularly update the knowledge graph to reflect the latest data changes (such as new events, entities, or relationships). Finally is knowledge reasoning: Use rule-based reasoning or reasoning algorithms of graph neural networks to complete the knowledge graph and discover potential relationships or events. The node update formula of the graph neural network is:

[0254]

[0255] Where, is the representation of node v at the l-th layer, is the set of neighbors of node v, W (l) and b (l) are learnable parameters, and σ is the activation function.

[0256] Knowledge graph application: The constructed knowledge graph provides a reliable basis for subsequent model training and inference.

[0257] In addition, the causal inference model used to implement this solution can be evaluated, and standard metrics (such as accuracy, recall, F1 score, etc.) are used to measure the model performance. According to the evaluation results, adjust the model parameters and optimize the inference strategy. Regularly update the knowledge graph to incorporate new data and knowledge to maintain the timeliness and accuracy of the model. Through continuous optimization, improve the adaptability and accuracy of the system in different scenarios.

[0258] In one embodiment, the traffic event analysis method based on multi-hop causal path exploration further includes:

[0259] Generate and output a causal inference report, where the causal inference report includes the inference process data and the sources of the data supporting the inference process.

[0260] Specifically, finally, output a detailed multi-hop causal path inference report, including the visualization results of the inference process and the sources of the supporting data for each path.

[0261] In addition, display the generated multi-hop causal paths in a visual form to help users intuitively understand the inference process. Use a directed graph to show the causal relationships between nodes: where the nodes represent events or entities (such as "traffic congestion", "traffic accident"). And the edges represent causal relationships (such as "causing"), and different colors or thicknesses are used to represent the weights of the edges.

[0262] In addition, in the report, list in detail the sources of the supporting data for each path, including: First, the data sources, such as real-time traffic flow data, meteorological data, event logs, etc. Second, the semantic relevance score, the semantic relevance score S(q,v) for each path. And the logical consistency score: the logical consistency score C(v) for each path.

[0263] The finally generated multi-hop causal path inference report includes the following content:

[0264] Inference path: The generated multi-hop causal chain (such as "traffic congestion → traffic accident → road slippery → rainfall").

[0265] Visualization result: The visualization chart of the causal path.

[0266] Supporting data: The sources of the supporting data for each path and the relevant scores.

[0267] The key steps of this application include: data collection and knowledge graph construction, problem parsing and initial node selection, chain reasoning and path expansion, path termination and result generation, as well as result verification and optimization. By constructing a knowledge graph containing multi-dimensional causal relationships, combining chain reasoning technology and heuristic scoring functions, dynamically expanding and optimizing causal paths, and using the Monte Carlo Tree Search algorithm (MCTS) to optimize path selection in each step of the reasoning process.

[0268] Suppose a traffic system needs to analyze the causal chain of a complex traffic accident. First, relevant event logs, traffic flow data, and meteorological data are collected to construct a knowledge graph in the traffic domain. After parsing the input problem, the node related to the traffic accident is selected as the starting node of the reasoning. Then, through chain reasoning and the MCTS algorithm, the reasoning path is dynamically expanded, and finally, a causal chain is generated to help relevant departments determine the main cause of the accident and guide emergency decision-making.

[0269] Reference Figure 2 , a traffic event analysis device based on multi-hop causal path exploration, characterized in that the traffic event analysis device based on multi-hop causal path exploration includes:

[0270] A knowledge graph construction module 100, configured to construct a knowledge graph based on multi-source heterogeneous data in the traffic domain;

[0271] An initial node determination module 200, configured to calculate the semantic relevance between each node in the knowledge graph and the problem semantic vector, and use the node with the highest semantic relevance as the initial reasoning node, where the problem semantic vector is obtained by converting the user input problem;

[0272] A path expansion module 300, configured to start from the initial reasoning node, combine a heuristic scoring function and the Monte Carlo Tree Search algorithm, and perform reasoning to obtain a multi-hop causal path, where the multi-hop causal path is a response to the user input problem.

[0273] In one embodiment, the path expansion module 300 includes:

[0274] A first calculation module, configured to calculate the heuristic scores of the candidate expansion nodes of the current node based on the heuristic scoring function, where the current node is the initial reasoning node or an expansion node selected during the path reasoning process;

[0275] A second calculation module, configured to evaluate the comprehensive value scores corresponding to the candidate expansion nodes based on the Monte Carlo Tree Search algorithm and the heuristic scores of the candidate expansion nodes;

[0276] A selection module, configured to select the candidate expansion node with the highest comprehensive value score as the next expansion node;

[0277] A generation module, configured to generate a multi-hop causal path based on a multi-hop causal chain in response to the multi-hop causal chain reaching a preset path termination condition, where the multi-hop causal chain includes an initial inference node and each extended node inferred therefrom later.

[0278] In one embodiment, the traffic event analysis device based on multi-hop causal path exploration further includes:

[0279] A screening module, configured to select multiple child nodes with the top heuristic scores among the child nodes adjacent to the current node as candidate extended nodes, where the number of candidate extended nodes is less than the number of child nodes of the current node.

[0280] In one embodiment, the traffic event analysis device based on multi-hop causal path exploration further includes:

[0281] A first extension stop determination module, configured to stop the path extension of the current path if the path length of the current path reaches the path length limit, where the current path includes an initial inference node and the extended nodes selected during path extension;

[0282] Or,

[0283] A second extension stop determination module, configured to stop the path extension of the current path where the current node is located if the heuristic score of the current node is lower than a first set threshold;

[0284] Or,

[0285] A third extension stop determination module, configured to stop the path extension of the current path where the current node is located if the semantic relevance of the current node exceeds a second set threshold.

[0286] In one embodiment, the comprehensive value score is calculated by the following formula 1:

[0287]

[0288] Or,

[0289] The comprehensive value score is calculated by the following formula 2:

[0290]

[0291] Where:

[0292] Q(s,a) represents the expected return of transferring from the current node s by performing the action a to the candidate extended node s′, and can be used as the comprehensive value score;

[0293] R(s,a) is the immediate reward, which is equal to the heuristic score H(v) of the candidate extended node s′;

[0294] P(s′|s,a) is the path transition probability, representing the probability of transferring from the current node s to the candidate expansion node s′ through the action a;

[0295] V(s′) is the value function of the subsequent state, obtained by aggregating the heuristic scores of the children nodes of the candidate expansion node s′;

[0296] λ is the discount factor, controlling the impact of future rewards on the current value;

[0297] UCT(s,a) is the upper confidence bound formula, which can be used as a comprehensive value score;

[0298] N(s,a) is the number of times the action a has been visited;

[0299] c is the exploration constant;

[0300] b represents all candidate actions under the current node s;

[0301] ∑ b N(s,b) represents the total number of times all executable actions b have been visited under the current node s.

[0302] In one embodiment, the heuristic score is calculated through the following formula 3:

[0303] H(v) = α·S rel (v)+β·S sup (v)+γ·S con (v)

[0304] Formula 3

[0305] where H(v) is the heuristic score of the node, and S rel (v), S sup (v), and S con (v) respectively represent the three scoring metrics of semantic relevance, support score, and logical consistency score of the node, and α, β, γ are the node scoring weights of different scoring metrics.

[0306] In one embodiment, the traffic event analysis device based on multi-hop causal path exploration further includes:

[0307] A node update module, configured to update the node scoring weights of the target node in the knowledge graph if there are path node pairs with causal relationship contradictions and / or path node pairs with causal relationships inconsistent with the predefined logical rules in the knowledge graph in the multi-hop causal path, where the node scoring weights include at least one of the node scoring weights corresponding to semantic relevance, the node scoring weights corresponding to support score, and the node scoring weights corresponding to logical consistency score;

[0308] A re-optimization module for re-optimizing path selection according to the node scoring weights of updated target nodes.

[0309] In one embodiment, the initial node determination module 200 includes:

[0310] A vector conversion module for converting a user input question into a question semantic vector;

[0311] A third calculation module for calculating the cosine similarity between the node semantic vectors of each node in the knowledge graph and the question semantic vector, where the cosine similarity is used to quantify the semantic relevance between the node and the user input question.

[0312] In one embodiment, the knowledge graph construction module 100 includes:

[0313] A normalization processing module for normalizing multi-source heterogeneous data;

[0314] An identification and extraction module for successively performing word segmentation, part-of-speech tagging, named entity recognition, and relationship extraction on the obtained normalized data;

[0315] A graph construction module for fusing the obtained node entities and the relationships between the node entities by eliminating redundancy and conflicts, and then linking them to the corresponding nodes in the knowledge graph, where the node entities include entities and events, and the nodes include entity nodes and event nodes.

[0316] In one embodiment, the knowledge graph construction module 100 further includes:

[0317] A verification module for verifying the nodes and relationships in the constructed knowledge graph;

[0318] And / or,

[0319] A graph update module for updating the knowledge graph according to incremental data.

[0320] In one embodiment, the traffic event analysis device based on multi-hop causal path exploration further includes:

[0321] A report generation module for generating and outputting a causal inference report, where the causal inference report includes inference process data and the sources of the data supporting the inference process.

[0322] Figure 3It is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 7000 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 710 (for example, one or more processors) and a memory 720, and one or more storage media 730 (for example, one or more mass storage devices) storing application programs 733 or data 732. Among them, the memory 720 and the storage media 730 may be transient storage or persistent storage. The program stored in the storage media 730 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the computer device 7000. Further, the processor 710 may be configured to communicate with the storage media 730 and execute a series of instruction operations in the storage media 730 on the computer device 7000.

[0323] The computer device 7000 may further include one or more power supplies 740, one or more wired or wireless network interfaces 750, one or more input / output interfaces 760, and / or one or more operating systems 731, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 3 The shown computer device structure does not constitute a limitation on the computer device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0324] The present application also provides a computer device. The computer device includes a memory and a processor. When the computer-readable instructions stored in the memory are executed by the processor, the processor is caused to execute the steps of the traffic event analysis method based on multi-hop causal path exploration in the above embodiments. The present application also provides a computer-readable storage medium. The computer-readable storage medium may be a non-volatile computer-readable storage medium, or may also be a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the traffic event analysis method based on multi-hop causal path exploration.

[0325] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0326] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0327] The above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. A traffic event analysis method based on multi-hop causal path exploration, characterized in that, The traffic event analysis method based on multi-hop causal path exploration includes: Construct a knowledge graph based on multi-source heterogeneous data in the traffic field; Calculate the semantic relevance between each node in the knowledge graph and the problem semantic vector, and use the node with the highest semantic relevance as the initial reasoning node, where the problem semantic vector is obtained by converting the user input problem; Starting from the initial reasoning node, combine the heuristic scoring function and the Monte Carlo tree search algorithm to perform reasoning to obtain a multi-hop causal path, where the multi-hop causal path is the answer to the user input problem.

2. The traffic event analysis method based on multi-hop causal path exploration according to claim 1, wherein The step of starting from the initial reasoning node, combining the heuristic scoring function and the Monte Carlo tree search algorithm to perform reasoning to obtain a multi-hop causal path includes: Based on the heuristic scoring function, calculate the heuristic scores of the candidate expansion nodes of the current node, where the current node is the initial reasoning node or an expansion node selected during the path reasoning process; Based on the Monte Carlo tree search algorithm and the heuristic scores of the candidate expansion nodes, evaluate the comprehensive value scores corresponding to the candidate expansion nodes; Select the candidate expansion node with the highest comprehensive value score as the next expansion node; In response to the multi-hop causal chain reaching the preset path termination condition, generate a multi-hop causal path based on the multi-hop causal chain, where the multi-hop causal chain includes the initial reasoning node and each expansion node inferred thereafter.

3. The traffic event analysis method based on multi-hop causal path exploration according to claim 2, wherein Before evaluating the comprehensive value scores corresponding to the candidate expansion nodes based on the Monte Carlo tree search algorithm and the heuristic scores of the candidate expansion nodes, it further includes: Take multiple child nodes with the top heuristic scores among the child nodes adjacent to the current node as candidate expansion nodes, where the number of candidate expansion nodes is less than the number of child nodes of the current node.

4. The traffic event analysis method based on multi-hop causal path exploration according to claim 2 or 3, characterized in that The multi-hop causal chain reaching the preset path termination condition is judged through the following steps: If the path length of the current path reaches the path length limit, stop the path expansion of the current path, where the current path includes the initial reasoning node and the expansion nodes selected during the path expansion; Or, If the heuristic score of the current node is lower than the first set threshold, stop the path expansion of the current path where the current node is located; Or, If the semantic relevance of the current node exceeds the second set threshold, stop the path expansion of the current path where the current node is located.

5. The traffic event analysis method based on multi-hop causal path exploration according to claim 2, wherein The comprehensive value score is calculated through the following formula 1: Or, The comprehensive value score is calculated through the following formula 2: Where: Q(s,a) represents the expected reward for transferring from the current node s to the candidate expansion node s′ by performing the action a, and can be used as the comprehensive value score; R(s,a) is the immediate reward, equal to the heuristic score H(v) of the candidate expansion node s′; P(s′|s,a) is the path transition probability, indicating the probability of transferring from the current node s to the candidate expansion node s′ through the action a; V(s′) is the value function of the subsequent state, obtained by performing mean aggregation on the heuristic scores of the child nodes of the candidate expansion node s′; λ is the discount factor, which controls the influence of future rewards on the current value; UCT(s,a) is the upper confidence limit formula and can be used as the comprehensive value score; N(s,a) is the number of times the action a is visited; c is the exploration constant; b represents all candidate actions under the current node s; ∑ b N(s, b) represents the total number of visits to all executable actions b under the current node s.

6. The traffic event analysis method based on multi-hop causal path exploration according to claim 2 or 3, characterized in that The heuristic score is calculated by the following formula 3: H(v) = α·S rel (v) + β·S sup (v) + γ·S con (v) Formula 3 Among them, H(v) is the heuristic score of the node, and S rel (v), S sup (v), and S con (v) represent the three scoring metrics of the semantic relevance, support score, and logical consistency score of the node respectively. α, β, and γ are the node scoring weights of different scoring metrics.

7. The traffic event analysis method based on multi-hop causal path exploration according to claim 6, wherein The traffic event analysis method based on multi-hop causal path exploration further includes: If there are path node pairs with contradictory causal relationships and / or path node pairs with causal relationships inconsistent with the predefined logical rules in the knowledge graph in the multi-hop causal path, then update the node score weights of the target nodes in the knowledge graph, where the node score weights include at least one of the node score weights corresponding to semantic relevance, the node score weights corresponding to support degree scores, and the node score weights corresponding to logical consistency scores; According to the updated node score weights of the target nodes, re-optimize the path selection.

8. The traffic event analysis method based on multi-hop causal path exploration according to claim 1, characterized in that Calculating the semantic relevance between each node in the knowledge graph and the problem semantic vector includes: Converting the user input problem into a problem semantic vector; Calculating the cosine similarity between the node semantic vector of each node in the knowledge graph and the problem semantic vector, where the cosine similarity is used to quantify the semantic relevance between the node and the user input problem.

9. The traffic event analysis method based on multi-hop causal path exploration according to claim 1, characterized in that Constructing a knowledge graph based on multi-source heterogeneous data in the traffic field includes: Performing standardization processing on the multi-source heterogeneous data; Successively performing word segmentation and part-of-speech tagging on the obtained standardized data, and then performing named entity recognition and relationship extraction; By eliminating redundancy and conflicts, fusing the obtained node entities and the relationships between the node entities and linking them to the corresponding nodes in the knowledge graph, where the node entities include entities and events, and the nodes include entity nodes and event nodes.

10. The traffic event analysis method based on multi-hop causal path exploration according to claim 9, characterized in that Constructing a knowledge graph based on multi-source heterogeneous data in the traffic field further includes: Verifying the nodes and relationships in the constructed knowledge graph; And / or Updating the knowledge graph according to the incremental data.

11. The traffic event analysis method based on multi-hop causal path exploration according to claim 1, wherein The traffic event analysis method based on multi-hop causal path exploration further includes: Generating and outputting a causal inference report, where the causal inference report includes inference process data and the sources of the data supporting the inference process.

12. A traffic event analysis device based on multi-hop causal path exploration, characterized in that, The traffic event analysis device based on multi-hop causal path exploration includes: A knowledge graph construction module for constructing a knowledge graph based on multi-source heterogeneous data in the traffic field; An initial node determination module for calculating the semantic relevance between each node in the knowledge graph and the problem semantic vector, and taking the node with the highest semantic relevance as the initial inference node, where the problem semantic vector is obtained by converting the user input problem; A path expansion module for starting from the initial inference node, combining a heuristic scoring function and a Monte Carlo tree search algorithm, and performing inference to obtain a multi-hop causal path, where the multi-hop causal path is a response to the user input problem.

13. A computer device, characterized in that, The computer device includes: a memory and at least one processor, and instructions are stored in the memory; The at least one processor invokes the instructions in the memory to cause the computer device to execute the traffic event analysis method based on multi-hop causal path exploration according to any one of claims 1-11.

14. A computer-readable storage medium, on which instructions are stored, characterized in that, When the instructions are executed by the processor, the traffic event analysis method based on multi-hop causal path exploration according to any one of claims 1-11 is implemented.

Citation Information

Cited By

  • Consumer right protection multi-source legal knowledge graph construction and intelligent retrieval method

    CN120743931A

  • Multi-dimensional police service data intelligent search method based on NLP semantic analysis

    CN121071203A

  • Intelligent Search Method for Multidimensional Police Data Based on NLP Semantic Analysis

    CN121071203B

  • Urban traffic event semantic recognition method based on knowledge graph

    CN121071561A

  • User portrait generation method and device based on knowledge graph, and storage medium

    CN121563587A