Analysis method and device of historical operation data, equipment and storage medium

By combining knowledge graphs and large language models, identifying analytical anchor points and constructing causal subgraphs, the system solves the problem of automated deep causal tracing of complex operational issues in unmanned vehicle fleets, improving the comprehensiveness and accuracy of the analysis.

CN121235237APending Publication Date: 2025-12-30APOLLO INTELLIGENT DRIVING (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511351601.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing technologies are insufficient to automate and deeply reveal the causal chains behind the complex operational anomalies of unmanned vehicle fleets, and traditional expert systems or rule engines are unable to handle the multi-dimensional and cross-system interaction effects.

Method used

By combining structured knowledge graphs and large language models, and by determining analytical anchor points, we can trace relationships in the spatiotemporal knowledge graph, construct causal subgraphs, and transform them into analytical prompts for the large model, thereby achieving deep causal tracing.

Benefits of technology

It enables automation and in-depth causal tracing of complex operational problems, improves the comprehensiveness and accuracy of analysis, reduces manual intervention, and expands the applicable scenarios of analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121235237A_ABST
    Figure CN121235237A_ABST
Patent Text Reader

Abstract

The invention provides a historical operation data analysis method and device, equipment and a storage medium, and relates to the technical field of data processing, in particular to the technical field of intelligent transportation, automatic driving, large models and data analysis. According to the specific implementation scheme, at least one analysis anchor point is determined in a space-time knowledge graph representing the operation state of the unmanned carrying tool according to a to-be-analyzed tracing problem; starting from at least one analysis anchor point, tracing the incidence relation in the space-time knowledge graph so as to identify and construct a cause subgraph containing a potential cause-effect chain; converting the structured information of the incong sub-graph into an analysis prompt for a large model; and obtaining an analysis result of the tracing problem according to the reasoning of the analysis prompt by the large model. According to the technical scheme of the invention, the deep causal tracing of complex operation problems can be realized by combining a structured knowledge graph and a large language model with cognitive intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to the fields of intelligent transportation, autonomous driving, big data models, and data analysis technology. Background Technology

[0002] With the continuous advancement of autonomous driving technology, driverless transportation vehicles, represented by driverless taxis, are gradually moving from the technology verification stage to large-scale commercial operation. In this process, the industry's focus has expanded from improving the autonomous driving capabilities of individual vehicles to ensuring the safe, efficient, and economical operation of the entire fleet. Large-scale fleet operation generates massive amounts of multi-dimensional operational data, providing a data foundation for achieving refined management and intelligent decision-making. Summary of the Invention

[0003] This disclosure provides a method, apparatus, device, and storage medium for analyzing historical operational data. According to one aspect of this disclosure, a method for analyzing historical operational data is provided, comprising:

[0004] Based on the tracing problem to be analyzed, at least one analytical anchor point is determined in the spatiotemporal knowledge graph that represents the fleet's operational status.

[0005] Starting from at least one analytical anchor point, trace its relationships in the spatiotemporal knowledge graph to identify and construct causal subgraphs containing potential causal chains;

[0006] Transform the structured information of the causal subgraph into analytical hints for the larger model;

[0007] Based on the reasoning of the large model's analytical prompts, the analytical results of the tracing problem are obtained.

[0008] According to another aspect of this disclosure, an apparatus for analyzing historical operational data is provided, comprising:

[0009] The extraction module is used to determine at least one analysis anchor point in the spatiotemporal knowledge graph representing the fleet's operational status, based on the traceability question to be analyzed.

[0010] The building module is used to trace the relationships in the spatiotemporal knowledge graph starting from at least one analysis anchor point in order to identify and build a causal subgraph containing potential causal chains.

[0011] The transformation module is used to convert the structured information of the causal subgraph into analytical hints for the larger model;

[0012] The reasoning module is used to deduce the analysis results of the tracing problem based on the analysis prompts from the large model.

[0013] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0014] At least one processor; and

[0015] The memory is communicatively connected to the at least one processor; wherein,

[0016] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.

[0017] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.

[0018] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.

[0019] According to the technical solution disclosed herein, a deep causal tracing of complex operational problems can be achieved by combining structured knowledge graphs and large language models with cognitive intelligence.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0021] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0022] Figure 1 This is a flowchart illustrating a method for analyzing historical operational data provided in one embodiment of this disclosure;

[0023] Figure 2 This is a flowchart illustrating a method for analyzing historical operational data provided in another embodiment of this disclosure;

[0024] Figure 3 This is a flowchart illustrating a method for analyzing historical operational data provided in another embodiment of this disclosure;

[0025] Figure 4 This is a schematic diagram of the structure of a historical operation data analysis device provided in an embodiment of this disclosure;

[0026] Figure 5 This is a block diagram of an electronic device used to implement embodiments of the present disclosure. Detailed Implementation

[0027] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0028] In related technologies, current operational data analysis largely relies on expert systems or pre-defined rule engines. These systems can effectively identify known, clearly patterned single-point operational problems. However, current technologies still have limitations when conducting in-depth operational analysis of unmanned vehicle fleets, especially when performing complex historical event causal tracing. Many serious operational anomalies are often the result of the intertemporal coupling of factors from multiple different domains and systems. For example, an increase in user complaint rates in a region may be related to multiple dimensions of information, including the vehicle hardware status in that region, recent adjustments to dispatch strategies, and even the scarcity of charging station resources.

[0029] For these complex and unexpected operational problems, traditional expert systems or rule engines struggle to effectively capture and diagnose them using pre-set, simple rules. Even when data is stored in a knowledge graph, the analysis process often relies on manual, exploratory correlation queries and logical reasoning by operations personnel, making it difficult to automatically and deeply reveal the hidden causal chains behind the events.

[0030] To at least partially address one or more of the aforementioned problems and other potential issues, embodiments of this disclosure provide an analysis method for historical operational data based on knowledge graphs and large-scale models. By utilizing the technical solutions of embodiments of this disclosure, a deep causal tracing of complex operational problems can be achieved by combining structured knowledge graphs and large-scale language models with cognitive intelligence.

[0031] This embodiment relates to a method for analyzing historical operational data, which aims to achieve in-depth causal tracing of complex operational problems by combining structured knowledge graphs and large language models with cognitive intelligence. Figure 1 This is a flowchart illustrating a method for analyzing historical operational data provided in one embodiment of this disclosure. Figure 1 As shown, the method includes at least the following steps:

[0032] S110. Based on the traceability problem to be analyzed, determine at least one analysis anchor point in the spatiotemporal knowledge graph that represents the operational status of the fleet.

[0033] In this embodiment, the fleet consists of multiple unmanned vehicles, which can be understood as autonomous vehicles performing specific transportation tasks on a ground road network. Specific forms include, but are not limited to, driverless taxis for passenger services and driverless delivery vehicles for freight services. This embodiment is applicable to various operational scenarios for unmanned vehicles.

[0034] For example, in an autonomous driving demonstration zone, there is a fleet of 150 pure electric driverless taxis. The operational goal of this fleet is to maximize the satisfaction of users' travel needs within the area while ensuring absolute safety. However, large-scale operation of driverless taxis can generate numerous operational challenges. For instance, during rush hours, a large number of orders will flow unidirectionally and centrally from residential areas to work areas (or vice versa), leading to a severe shortage of vehicles at the origin and a large backlog at the destination. Furthermore, untimely vehicle dispatching and unreasonable spatial distribution can result in excessively long waiting times for users after placing an order, impacting user experience and potentially causing order cancellations.

[0035] In this step, we first need to understand the tracing problem to be analyzed and find the starting point for the analysis in a pre-constructed spatiotemporal knowledge graph.

[0036] A traceability question can be understood as an analytical request made by operations personnel or automated systems to explore the reasons behind a historical operational phenomenon. This question can be in text form, such as "Analyze the reasons why order ORD-999 was complained about by the user."

[0037] A spatiotemporal knowledge graph is a data model that uses a graph structure to model operational entities and the complex relationships between them. The "spatiotemporal" characteristic means that entities and relationships in the graph possess attributes with temporal and spatial dimensions. "Knowledge" is reflected in the fact that the graph contains not only raw data but also deep causal or relational relationships extracted through rules and models. Conceptually, a knowledge graph consists of "entities" as nodes and "relationships" as edges. In this solution, the spatiotemporal knowledge graph can predefine entity types such as "vehicle," "order," "work order," and "spatiotemporal slice." An entity node refers to a node created in the graph database for each specific data record or operational event, based on its type. For example, when the system receives data with the vehicle ID "VN-123," it creates an entity node of type "vehicle," using "VN-123" as its unique identifier. Operational data related to this vehicle can be used as attributes of this node or as associated nodes connected to it.

[0038] Analysis anchors can be understood as one or more initial nodes located by the system in the spatiotemporal knowledge graph based on the tracing problem, used to initiate subsequent causal chain tracing. The determination of analysis anchors serves as a bridge connecting abstract problems with specific graph data; the specific method for determining them will be detailed in subsequent embodiments.

[0039] For example, suppose a traceability question to be analyzed is: "Analyze the reason why order ORD-999 was complained about by users." After receiving this question, the system will first parse it and determine the node associated with "order ORD-999" in the spatiotemporal knowledge graph as the anchor point for this analysis.

[0040] S120. Starting from at least one analytical anchor point, trace its relationships in the spatiotemporal knowledge graph to identify and construct a causal subgraph containing potential causal chains.

[0041] In this step, after determining the analysis anchor point, we will start from that anchor point and explore the graph to construct a local graph structure focused on causal analysis.

[0042] Tracing the relationships can be understood as a graph traversal or graph reasoning process, which aims to start from the anchor point and search along the edges of the graph for all other nodes that may have a causal relationship with it.

[0043] A causal subgraph can be understood as a subset extracted from the full knowledge graph. It contains the analysis anchor points and all potential causal nodes discovered by tracing relationships and the connection paths between them, thus forming a visualized network of potential causal chains.

[0044] For example, starting with the analysis anchor point "Order ORD-999", the system traces back in the knowledge graph and may find that the order was executed by "Vehicle VN-456", and that "Vehicle VN-456" experienced an "emergency braking" event 5 minutes before the order was executed. The system will then construct a causal subgraph by combining the "Order ORD-999" node, the "Vehicle VN-456" node, the "emergency braking" event node, and the relationships between them.

[0045] S130. Transform the structured information of the causal subgraph into analytical prompts for the large model.

[0046] In this step, the machine-readable graph structure data needs to be translated into a natural language format that can be understood by a large language model (also known in the field as a large model or LLM).

[0047] Analysis hints can be a well-designed text paragraph containing background information, context, and explicit instructions, used to guide large models in high-quality inference.

[0048] For example, the cause-effect subgraph information constructed in the previous step can be transformed into text like this: "Background: Analyze the reason why order ORD-999 was complained about. Related event: The associated vehicle VN-456 experienced a sudden braking incident 5 minutes before the service started." This text will become part of the analysis prompt.

[0049] S140. Based on the reasoning of the analysis prompts from the large model, the analysis results of the tracing problem are obtained.

[0050] In this step, the analysis hints generated in the previous step are sent to a pre-configured large language model via an API interface. The large model will perform logical reasoning based on the information in the hints and return an analysis result in text format. The analysis result can be a comprehensive report that includes inferences about the root cause of the problem, logical explanations, confidence assessments, and suggestions for future improvements.

[0051] For example, upon receiving the analysis prompt, the large model might infer and return: "The root cause of the complaint against order ORD-999 is most likely due to the sudden braking incident that occurred on vehicle VN-456 during the pick-up process, resulting in a poor passenger experience."

[0052] According to the scheme of this disclosure, by determining the analysis anchor point, tracing the relationship to construct the causal subgraph, and then converting the subgraph information into analysis prompts for the large model to reason, a complete technical link from the original problem to the in-depth analysis result is established. This can combine the objectivity of structured graph data with the cognitive intelligence of the large language model, and solve the problem of operation tracing that is difficult to handle with complex and implicit causal relationships by traditional methods.

[0053] In one possible implementation, step S110, based on the tracing problem to be analyzed, determines at least one analysis anchor point in the spatiotemporal knowledge graph representing the fleet's operational status, and may further include the following steps:

[0054] S111. Perform natural language understanding on textual traceability issues to identify traceability intent and extract key entities.

[0055] This step is a prerequisite for determining the analysis anchor point, aiming to understand the user's natural language questions. Natural language understanding can be implemented based on existing algorithms or models for parsing text input. This process mainly includes two sub-tasks:

[0056] Tracing Intent Recognition: Determine the type of user's question, such as whether they want to trace a specific event or attribute a macro-level indicator.

[0057] Key entity extraction: Extract useful information fragments from the text, such as vehicle ID, order ID, time adverbs, and location nouns, and link them with standardized identifiers in the knowledge graph.

[0058] For example, when asked "Why did vehicle VN-123 go offline on Orchard Road yesterday?", the NLU module would identify:

[0059] Tracing intent: "Event-driven tracing".

[0060] Key Entity: {Vehicle ID:"VN-123", Time:"Yesterday", Location:"Orchard Road", Event:"Offline"}.

[0061] S112. Based on the tracing intent and key entities, locate at least one analysis anchor point in the spatiotemporal knowledge graph.

[0062] In this step, the structured information parsed by NLU in the previous step is used to query the spatiotemporal knowledge graph to locate one or more nodes that are most suitable as the starting point for analysis.

[0063] In the example above, based on the key entities {Vehicle ID:"VN-123", Time:"Yesterday", Event:"Offline"}, we can query the knowledge graph for the operational event node of type "Offline" associated with vehicle VN-123 yesterday, and determine that node as the anchor point for this analysis.

[0064] According to the embodiments of this disclosure, by introducing a natural language understanding module, the system is able to process unstructured text queries. Human natural language queries are accurately translated into machine-executable graph localization instructions, greatly improving the usability and intelligence of the analysis system, enabling even operations personnel unfamiliar with database query languages ​​to perform in-depth data analysis.

[0065] In one possible implementation, step S112 locates at least one analysis anchor point in the spatiotemporal knowledge graph based on the tracing intent and key entities, and may further include the following steps:

[0066] When the tracing problem is an event-driven tracing problem, the corresponding target event node can be directly searched in the spatiotemporal knowledge graph based on the event identifier in the key entity.

[0067] At least one entity node associated with the target event node is identified as the analysis anchor point.

[0068] In this embodiment of the disclosure, event-driven tracing problems can be understood as tracing analyses that revolve around a clearly defined operational event that has already occurred. When the NLU module determines that a problem belongs to this type, the system will execute a direct and efficient localization strategy.

[0069] Specific example: For the question "Analyze the reasons for customer complaint C-886", the key entities extracted by NLU include the unique event identifier C-886. The system will use this ID as an index to directly query and locate the target event node representing "customer complaint C-886" in the knowledge graph.

[0070] In some cases, tracing directly from the event node may not provide sufficient information. To obtain richer context, the system can use the core business entity directly related to the event as the anchor point for analysis. Entity nodes can represent physical or business objects, such as vehicle nodes or order nodes.

[0071] In the example above, after locating the event node "Customer Complaint C-886", the system found that it is connected to the "Order ORD-999" node and the "Vehicle VN-456" node through a relationship edge. For comprehensive analysis, the system can jointly determine the two entity nodes "Order ORD-999" and "Vehicle VN-456" as the anchor points for this analysis.

[0072] According to the solutions in this disclosure, by designing a specialized anchor point location strategy for event-driven problems, the starting point of the analysis can be found quickly and accurately. By using the core entities associated with the event as anchor points, subsequent causal chain tracing can be conducted in a broader and more relevant context, thereby improving the comprehensiveness of the analysis.

[0073] In one possible implementation, step S112 locates at least one analysis anchor point in the spatiotemporal knowledge graph based on the tracing intent and key entities, and may further include the following steps:

[0074] In the case of a non-event-driven tracing problem, based on at least one of the entity identifier, time range, or indicator information of the key entity, a set of operational event nodes related to the tracing problem are inferred and generated in the spatiotemporal knowledge graph as at least one anchor node.

[0075] In this embodiment of the disclosure, non-event-driven tracing problems can be understood as tracing analyses of phenomena that do not involve explicit abnormal events but rather target a specific entity or macroscopic indicator. For these types of problems, the system needs to perform a "hypothesis generation" or "problem dimensionality reduction" process, that is, intelligently inferring which underlying events are most likely to have led to the macroscopic phenomena observed by the user.

[0076] Specific example: For the question, "Why was the order volume for vehicle VN-789 so low yesterday?", there are no specific abnormal events. The system will query the knowledge graph for all related events for that vehicle yesterday based on the key entity {Vehicle ID: "VN-789", Time: "Yesterday"}, and infer the most abnormal events, such as a "charging in progress" event lasting up to 4 hours and several "order dispatch failures". The system will then use these inferred event nodes, which are most likely to lead to the "low order volume" phenomenon, to generate the anchor node set for this analysis.

[0077] According to the embodiments of this disclosure, by designing intelligent inference and anchor point generation mechanisms for non-event-driven problems, this method can handle more open and fuzzy analysis requests. This method can automatically reduce the dimensionality of a macroscopic, phenomenon-level problem and focus it on a set of specific, analyzable micro-events, greatly expanding the applicability and analytical capabilities of the method.

[0078] In one possible implementation, such as Figure 2 As shown, step S120 begins with at least one analytical anchor point, tracing its relationships in the spatiotemporal knowledge graph to identify and construct a causal subgraph containing potential causal chains, and may further include the following steps:

[0079] S210. Starting from at least one analysis anchor point, perform a temporal reverse traversal in the spatiotemporal knowledge graph to obtain multiple candidate cause nodes.

[0080] In this embodiment, the temporal reverse traversal can be a graph search algorithm. Starting from the analysis anchor point, it backtracks along the "coming directions" of the relation edges, searching for all nodes that occurred before the anchor point in time and are connected to the anchor point by a path. Candidate cause nodes can be all nodes that meet the conditions found during this traversal, forming the initial set of potential causes.

[0081] For example, starting from the anchor point "order ORD-999 was complained about", the system traverses backwards and may find multiple candidate reason nodes such as "vehicle VN-456 executed this order", "vehicle VN-456 experienced emergency braking at T-5 minutes", and "vehicle VN-456 underwent route replanning at T-10 minutes".

[0082] S220. Score the causal contribution of multiple candidate cause nodes.

[0083] In this embodiment of the disclosure, in order to select the most important cause from numerous candidate causes, it is necessary to quantitatively evaluate each candidate cause node. The causal contribution score can be a comprehensive score used to measure the likelihood of a genuine causal relationship between a candidate cause node and the analysis anchor point.

[0084] In the example above, the "sudden braking" node and the "path replanning" node found in the previous step will be scored separately. It is possible that the causal contribution score of "sudden braking" is 0.92, while the score of "path replanning" is 0.75.

[0085] S230. Select core causal events based on causal contribution scores to construct a causal subgraph.

[0086] In this embodiment of the disclosure, a scoring threshold can be set to filter out all candidate causal nodes whose scores are higher than the threshold. Core causal events are these high-probability causal nodes that have passed the scoring screening. The final causal subgraph is composed of analysis anchor points, these core causal events, and the relationships between them.

[0087] If the scoring threshold is 0.7, then the two nodes "sudden braking" and "path replanning" will be selected as core causal events and included in the final causal subgraph.

[0088] According to the solution of this disclosure, through a three-step process of "traversal-scoring-filtering," the most noteworthy core causes can be intelligently and quantitatively identified from massive amounts of related information. This method replaces manual experience-based judgment with a data-driven scoring model, greatly improving the objectivity and accuracy of cause localization.

[0089] In one possible implementation, step S220 scores the causal contribution of multiple candidate cause nodes, and may further include the following steps:

[0090] The causal contribution score is calculated based on at least one of the following factors: temporal proximity, spatial proximity, pre-defined event correlation strength, and shared core business entities between the candidate cause node and the analysis anchor point.

[0091] In this embodiment of the disclosure, the causal contribution score can be a comprehensive score calculated by weighting multiple dimensions, which can be any one, two, three, four, or even any number of more dimensions. These dimensions can be:

[0092] Temporal proximity: The smaller the time difference between the candidate cause and the anchor point, the higher the score.

[0093] Spatial proximity: The closer the two are geographically, the higher the score.

[0094] Preset event association strength: Based on offline mining of massive historical data, the association probability between different event types is pre-calculated. For example, if historical data shows that the probability of a "complaint" occurring after "sudden braking" is 30%, then this 30% can be used as part of the association strength score.

[0095] Shared core business entities: If the candidate reason is associated with the same vehicle or the same order as the anchor point, a higher score can be obtained.

[0096] Taking the above example, for the candidate reason "sudden braking" event, it is calculated that the time difference between it and the "complaint" anchor point is only 5 minutes (high time score), the location is the same (high spatial score), the historical correlation is strong, and it shares the same vehicle and order (high entity sharing score), and finally it gets a high comprehensive score of 0.92.

[0097] According to the solution of this disclosure, a multi-dimensional scoring model can comprehensively and holistically assess the likelihood of an event being a cause from multiple perspectives, including time, space, historical statistics, and business logic. This comprehensive quantitative assessment method is more reliable and robust than single-dimensional judgment.

[0098] In one possible implementation, such as Figure 3 As shown, step S210 starts from at least one analysis anchor point and performs a temporal reverse traversal in the spatiotemporal knowledge graph to obtain multiple candidate cause nodes. This may further include the following steps:

[0099] S211. Based on the spatiotemporal properties of at least one analysis anchor point, set the spatiotemporal constraint boundary.

[0100] In this embodiment of the disclosure, before the traversal begins, a reasonable search range is defined by spatiotemporal constraint boundaries to improve efficiency. The spatiotemporal constraint boundaries can be a search range defined by both time and space, such as a "30-minute" time window ending at the time the anchor point occurs, and a "5-kilometer" radius range centered on the location where the anchor point occurs.

[0101] S212. Starting from any analysis anchor point, perform a temporal reverse traversal in the spatiotemporal knowledge graph.

[0102] S213. During the reverse temporal traversal, if the current node exceeds the spatiotemporal constraint boundary, prune the path where the current node is located until all paths have been traversed.

[0103] In this embodiment of the disclosure: pruning refers to the process during traversal where, when the algorithm finds that the occurrence time or location of a candidate cause node has exceeded the set boundary, it will immediately stop tracing back along that node.

[0104] For example, starting from the "complaint" anchor point (occurring at 3:00 PM, location A), when traversing backwards, if a "vehicle maintenance" event occurs at 10:00 AM, since its time exceeds the "first 30 minutes" boundary, any events that occurred before "vehicle maintenance" will no longer be explored, thus eliminating this invalid search path.

[0105] S214. Based on the remaining spatiotemporal knowledge graph after pruning, obtain multiple candidate cause nodes related to the tracing problem and containing potential causal chains. The paths include nodes in the spatiotemporal knowledge graph and edges connected to those nodes.

[0106] According to the scheme of this disclosure, by presetting spatiotemporal constraint boundaries and performing dynamic pruning during traversal, computational resources can be concentrated on the most relevant temporal and spatial ranges. This method greatly reduces unnecessary graph search overhead and significantly improves the analysis efficiency of causal tracing under large-scale knowledge graphs.

[0107] In one possible implementation, step S130 transforms the structured information of the subgraph into analytical cues for the larger model, and may further include the following steps:

[0108] S131. Transform the structured information of the causal subgraph into a narrative text containing weighted evidence.

[0109] In this embodiment of the disclosure, the narrative text can be a logical and readable natural language text that simulates a human analyst writing a report. This text organizes the nodes and edges in the causal subgraph according to a certain logic (such as chronological order or importance order). Weighted evidence refers to text that not only states facts but also explicitly indicates the strength of the correlation between those facts (i.e., the core causal event) and the issue; this strength can be derived from causal contribution scores.

[0110] S132. Embed the narrative text and tracing questions into the preset analysis prompt template to obtain analysis prompts for the large model.

[0111] In this embodiment of the disclosure, the analysis prompt template can be a pre-designed text framework that includes role-playing, task instructions, and output format requirements. Filling the "background information" section of the template with the narrative text generated in the previous step can create a complete and high-quality analysis prompt.

[0112] For example, an analysis prompt template could be: "#Role: You are a senior operations analyst. #Background Information: {Narrative text generated by S131 here}. #Task: Based on the background information, please provide the root cause and improvement suggestions."

[0113] According to the scheme of this disclosure, a high-quality way to interact with large models is created by transforming structured causal subgraphs into weighted narrative text and combining them with analysis prompt templates. This method provides highly focused, pre-processed objective evidence in a form best suited to the understanding and reasoning of large models, which is a key step in ensuring the accuracy and depth of subsequent analysis results.

[0114] In one possible implementation, step S131 transforms the structured information of the causal subgraph into narrative text containing weighted evidence, and may further include the following steps:

[0115] S310. Based on the causal contribution score of the core causal events in the causal subgraph, assign different strengths of evidence descriptions to the core causal events in the narrative text.

[0116] In this embodiment of the disclosure, this step is a specific implementation of "weighted evidence". A mapping rule from rating to description can be set.

[0117] A score > 0.9 indicates strong evidence, which can be described as "highly relevant core evidence".

[0118] 0.7 < rating <= 0.9 -> description could be: "Possible causes that are significantly relevant".

[0119] 0.5 < rating <= 0.7 -> description could be: "A supporting factor worth noting".

[0120] Specific example: In Example 5, the "emergency braking" event has a score of 0.92, and the "path replanning" event has a score of 0.75. In the generated narrative text, they would be described as:

[0121] "[Highly relevant core evidence]: The vehicle involved experienced a 'sudden braking' incident approximately 1 minute before the complaint was filed (overall contribution 0.92)."

[0122] "[Possible cause with significant relevance]: The vehicle performed a 'route replanning' approximately 5 minutes before the complaint occurred (overall contribution 0.75)."

[0123] According to the scheme of this disclosure, by assigning quantitatively based descriptions of the strength of evidence to different core causal events, the resulting narrative text not only contains facts but also conveys the hierarchy of importance of these facts. This approach significantly reduces the reasoning burden of large models, enabling them to quickly grasp the key to the problem and conduct logical analysis around the core evidence, thereby improving the accuracy and logicality of the final analysis results.

[0124] In one possible implementation, the spatiotemporal knowledge graph is built based on multi-source operational data reflecting the operational status of autonomous vehicles, obtained from at least multiple data sources, including scheduling systems, vehicle-mounted equipment, and work order systems.

[0125] In this embodiment, data is first collected from various key systems within the operational ecosystem. Multi-source operational data refers to a collection of data generated during the daily operation of autonomous vehicles, encompassing diverse sources and formats. These data sources include at least:

[0126] The dispatch system provides order information, trip information (including the start and end points of the trip), vehicle-order matching records, dispatch strategy logs, pricing information, etc., which are usually structured database records.

[0127] Vehicle-side equipment: mainly refers to the vehicle's telematics processor and autonomous driving computing unit (MPU), which will report the vehicle's location (GPS), speed, acceleration, battery charging status, vehicle hardware status, fault diagnostic codes, etc. in real time. It can also include the vehicle trajectory formed by the location information of the vehicle at multiple consecutive moments, which is usually a high-frequency real-time data stream.

[0128] Work order system: Records all offline operation and maintenance activities, such as charging work orders, battery swapping work orders, repair and maintenance work orders, cleaning work orders, etc., including information such as activity type, creation time, execution location, and completion status.

[0129] The operational status of autonomous vehicles encompasses two levels: micro-level individual status data and macro-level overall situational data.

[0130] 1. Micro-level individual state data mainly refers to the specific and independent operational state of each autonomous vehicle. This data primarily comes from vehicle-side equipment.

[0131] 2. Macro-level situational data mainly refers to overall data reflecting the current supply and demand, resource distribution, and efficiency from the perspective of the entire fleet or a specific operating area. This data primarily comes from real-time statistics and aggregation from the dispatching system and work order system.

[0132] The "spatiotemporal" characteristic in a spatiotemporal knowledge graph means that entities and relationships in the graph model possess attributes with temporal and spatial dimensions, enabling them to describe the evolution of operational states in space and time. "Knowledge" is reflected in the fact that the graph contains not only raw data but also deep causal or relational relationships extracted through rules and models. Conceptually, this knowledge graph consists of "entities" as nodes and "relationships" as edges.

[0133] It is understandable that a spatiotemporal knowledge graph is obtained by deeply integrating and associating multi-source data from all corners of the autonomous vehicle operation ecosystem through a series of steps such as data access, entity recognition, and relationship construction.

[0134] According to the solution of this disclosure embodiment, by constructing a spatiotemporal knowledge graph based on multi-source operational data, it is ensured that all analysis and prediction are based on comprehensive, real-time and accurate data, thus guaranteeing the reliability and practical value of the final output results.

[0135] Based on the aforementioned implementation steps, two complete examples corresponding to "event-driven" and "non-event-driven" tracing problems are given to illustrate the actual application process of this solution in detail.

[0136] Example 1:

[0137] First, let's describe the business scenario for this example. Yesterday at 12:30 PM, during the midday rush hour, the system automatically generated and alerted for a "severe supply-demand imbalance" event. This event occurred within a geographic grid within the Central Business District (CBD). Operations analysts need to investigate why such a severe capacity shortage occurred at this critical time and location.

[0138] The first step in the analysis process is to acquire and parse the tracing question. After the operations staff or analyst selects the alarm event in the system, it will be transformed into a clear tracing question: "Analyze the root cause of the 'Severe Supply-Demand Imbalance' event (ID: IMB-554) that occurred yesterday at 12:30 in the CBD grid (w21z7)." The natural language understanding module will parse this input, identify its tracing intent as "event-driven tracing," and extract key entities, including the event ID "IMB-554," the time "yesterday at 12:30," and the location "w21z7."

[0139] The next step is to determine the anchor point for analysis. Since the problem is event-driven, we can directly locate the operational event node related to "severe supply-demand imbalance" in the knowledge graph using the event ID "IMB-554". The system found that this event node is closely related to the "spatiotemporal slice" node "w21z7@yesterday12:30" representing this spatiotemporal range. To analyze the cause of the spatiotemporal slice imbalance, this spatiotemporal slice node is determined as the anchor point for this analysis.

[0140] After the anchor points were determined, a causal subgraph was constructed. Starting from the anchor point "w21z7@yesterday 12:30", a reverse traversal of the time sequence was performed, focusing on finding events that negatively impacted the transportation capacity supply in the area before the imbalance occurred, i.e., between 12:00 and 12:30. The system identified and scored several candidate causes. One candidate cause was a work order event for "mobile fleet concentrated charging," showing that 5 vehicles collectively swapped batteries at the edge of the CBD at 12:15, with a causal contribution score of 0.90. Another candidate cause was a "traffic congestion" event from third-party data, showing severe congestion on a main road leading into the CBD at 12:00, with a causal contribution score of 0.85. Yet another candidate cause was a "dispatch strategy change" event that occurred at 11:00 AM, which slightly reduced the incentive weight for idle vehicles to actively cruise, with a causal contribution score of 0.60.

[0141] After setting the scoring threshold to 0.5, all three reasons mentioned above were selected as core causal events. Based on this, the system constructed a causal subgraph containing nodes such as "supply and demand imbalance", "centralized energy replenishment", "traffic congestion" and "strategy change".

[0142] The system then transforms the constructed causal subgraph into analytical prompts for the larger model. The transformation process involves generating a narrative text from the structured information of the graph, for example: "Background: Analysis of the reasons for the severe supply-demand imbalance that occurred in the CBD grid (w21z7) at 12:30 yesterday. Highly relevant core evidence: 15 minutes before the imbalance occurred, 5 vehicles near the grid collectively entered 'battery swapping' mode, causing a sudden decrease in supply (overall contribution 0.90). Significantly relevant possible cause: 30 minutes before the imbalance occurred, a 'severe congestion' event occurred on a main road entering the grid, hindering the inflow of external transport capacity (overall contribution 0.85). A noteworthy auxiliary factor: Earlier that day, the incentive weight of the 'idle vehicle active cruising' scheduling strategy was reduced (overall contribution 0.60)."

[0143] This narrative text is then embedded into a pre-defined analysis prompt template, forming the final analysis prompt.

[0144] The final step is to obtain the analysis results. After receiving the analysis prompts, the large language model performs inference and outputs root cause inferences and operational improvement suggestions. The root cause inference is: this supply-demand imbalance is a typical operational risk caused by multiple concurrent factors. The direct cause is the superposition of a concentrated fleet refueling activity and a traffic congestion on a key road just before the midday rush hour, which simultaneously reduced both existing and new supply. Earlier adjustments to the dispatching strategy reduced the system's capacity reserves, making the system insufficiently resilient to such concurrent shocks. Operational improvement suggestions include: coordinating with the refueling plans of large customer fleets to guide them to stagger their concentrated refueling during midday and evening peak hours; simultaneously, the dispatching system should more sensitively integrate real-time traffic data, automatically increasing dispatching incentives for vehicles on surrounding uncongested roads when main roads entering hotspot areas are congested.

[0145] Example 2: Non-Event-Driven Tracing – Attribution Analysis of “Regional Service Quality Decline”

[0146] This example aims to answer a more open, macro-level traceability question raised by operational reporting data.

[0147] First, let's describe the business scenario in this example. The operations manager discovered in the weekly report that the "average passenger wait time" in area A (Jurong East) was as high as 12 minutes last Saturday, far exceeding the average of 7 minutes, and the reasons for this need to be investigated.

[0148] The first step in the analysis process is to acquire and analyze the tracing question. The operations manager inputs the question into the system: "Why was the average passenger wait time higher in area A last Saturday?" The system's natural language understanding module analyzes the input, identifies the tracing intent as "non-event-driven tracing," and extracts key entities, including location "A," time "last Saturday," and the higher-than-expected metric "average passenger wait time."

[0149] Next, we proceed to the step of determining the analysis anchor points. Since the problem is not event-driven, the system performs an intelligent inference. The system first queries the knowledge graph ontology to find key driving factors affecting the "average passenger waiting time," such as "regional supply-demand ratio," "average pick-up distance," and "order dispatch failure rate." By comparing the data for these factors in region A for "last Saturday" and "the previous Saturday," the system finds that the "average pick-up distance" was 80% higher last Saturday than the previous Saturday, making it the most significant factor. At this point, the problem is automatically reduced in dimensionality and focused on "analyzing the reasons for the excessively long average pick-up distance in region A last Saturday." The system further queries all order events in the region that day with a "pick-up distance" greater than 3 kilometers, and intelligently infers and generates the set of analysis anchor points for this analysis from these specific "long-distance order dispatch" events.

[0150] Once the anchor points are determined, the system begins constructing a causal subgraph. Starting from these "long-distance dispatch" event anchor points, the system performs a backward traversal of the time sequence and scores the discovered candidate causes. One candidate cause is the "weather" event node, showing that area A recorded "heavy rain" between 2 PM and 5 PM that day, with a causal contribution score of 0.92. Another candidate cause, discovered through analysis of the vehicle trajectories, is that a large number of vehicles originated from the neighboring area B, directly explaining the long pick-up distance, with a causal contribution score of 0.98.

[0151] Based on these core reasons, the system constructed a complete causal subgraph from "heavy rain" to "reduction in local vehicle supply in area A" to "inter-regional dispatch from area B", ultimately leading to "excessively long average pick-up distance" and "high average waiting time".

[0152] The system then transforms the causal subgraph into analytical prompts. The generated narrative text is as follows: "Background: Analysis of the reasons for the higher average waiting time in area A last Saturday. The core phenomenon is the 'excessively long average pick-up distance' in this area. Highly relevant core evidence: Analysis shows that a large number of orders were dispatched from idle vehicles in the neighboring area B over long distances (overall contribution 0.98). Significantly relevant possible causes: This phenomenon is highly correlated with the 'heavy rain' weather event that occurred in area A from 2 PM to 5 PM that day. Historical data shows that heavy rain is the main cause of temporary shortages in local transportation capacity, thus triggering large-scale cross-regional dispatching (overall contribution 0.92)."

[0153] This narrative text was then embedded into the analysis prompt template.

[0154] Finally, the analysis results are obtained. After receiving prompts, the large language model performs inference and outputs root cause inferences and operational improvement suggestions. The root cause inference is: the direct cause of the high waiting time is the excessively long average pick-up distance, while the root cause is that during heavy rain, a localized and temporary shortage of transport capacity occurred in area A, forcing the dispatch system to mobilize vehicles from neighboring areas on a large scale to meet the demand. Operational improvement suggestions are not elaborated further.

[0155] It's important to note that the two examples above are intended to clearly illustrate the basic process of this method. In actual large-scale operations, the situation is far more complex. The causal subgraph of an operational anomaly (such as "regional supply and demand imbalance") often contains not just two or three clearly identifiable candidate causes, but potentially dozens or even hundreds of spatiotemporally related event nodes with varying scores, forming a complex network. Algorithmic contribution scoring essentially measures strong correlation, not true causation. In complex subgraphs, there may be multiple events with very high scores. For example, a "vehicle malfunction" and a "traffic congestion" might both score highly, but whether the congestion caused the vehicle to idle for too long, triggering a malfunction alarm, or whether the vehicle breakdown caused the traffic congestion, the algorithm itself cannot make such a deep logical and common-sense judgment.

[0156] Real-world operational problems are usually not caused by a single reason, but rather by the interaction and amplification of multiple factors. For example, a slight "schedule strategy adjustment" (score 0.6), a "light rain" (score 0.5), and a few vehicles "charging normally" (score 0.4) are not the main causes individually. However, their combination—the strategy adjustment leading to a slight uneven distribution of vehicles, the light rain slowing down the overall vehicle speed, and a few vehicles going to charge at the same time—collectively creates a severe supply-demand imbalance. Traditional scoring models struggle to capture such non-linear relationships and combinations.

[0157] Moreover, algorithmic scoring cannot understand the deeper semantics and business implications of events; it cannot truly grasp the semantics of "detours" in different scenarios. Is it a "reasonable detour" to avoid congestion, or an "abnormal detour" due to a navigation system error? Or is it a "compliant detour" to accommodate passengers' temporary changes of destination? These all require judgment based on a broader context and common sense, which is beyond the scope of scoring models.

[0158] Therefore, in this complex fog, large language models play a crucial cognitive and reasoning role that algorithms cannot replace. They are responsible for reasoning from "relevance" to "causality." The analytical prompts they receive are not merely lists of events, but narrative texts already containing weighted evidence. Like an experienced operations expert, the large model can examine this "list of evidence," combining its built-in world knowledge and logical reasoning capabilities to construct the most logical causal narrative. It can determine cause and effect in specific scenarios, thus completing causal judgments that scoring models cannot.

[0159] Understandably, the knowledge graph's reverse traversal and contribution scoring are responsible for quickly and accurately collecting and organizing all relevant, weighted clues from massive amounts of data. The large language model, on the other hand, is responsible for the final, decisive logical reasoning and judgment on these complex clues.

[0160] Figure 4 This is a schematic diagram of the structure of a historical operation data analysis device 400 provided according to an embodiment of this disclosure. Figure 4 As shown, the device includes at least:

[0161] Extraction module 401 is used to determine at least one analysis anchor point in the spatiotemporal knowledge graph representing the fleet operation status based on the traceability problem to be analyzed.

[0162] Module 402 is used to trace the relationships in the spatiotemporal knowledge graph starting from at least one analysis anchor point in order to identify and construct a causal subgraph containing potential causal chains.

[0163] The transformation module 403 is used to transform the structured information of the causal subgraph into analytical prompts for the large model;

[0164] The reasoning module 404 is used to reason about the analysis prompts based on the large model to obtain the analysis results of the tracing problem.

[0165] In one possible implementation, the extraction module 401 is used for:

[0166] Natural language understanding is used to analyze textual traceability issues in order to identify traceability intent and extract key entities;

[0167] Based on the tracing intent and key entities, locate at least one analytical anchor point in the spatiotemporal knowledge graph.

[0168] In one possible implementation, the extraction module 401 is used for:

[0169] When the tracing problem is an event-driven tracing problem, the corresponding target event node can be directly searched in the spatiotemporal knowledge graph based on the event identifier in the key entity;

[0170] At least one entity node associated with the target event node is identified as the analysis anchor point.

[0171] In one possible implementation, the extraction module 401 is used for:

[0172] In the case of a non-event-driven tracing problem, based on at least one of the entity identifier, time range, or indicator information of the key entity, a set of operational event nodes related to the tracing problem are inferred and generated in the spatiotemporal knowledge graph as at least one anchor node.

[0173] In one possible implementation, building module 402 is used for:

[0174] Starting from at least one analysis anchor point, perform a temporal reverse traversal in the spatiotemporal knowledge graph to obtain multiple candidate cause nodes;

[0175] Scoring the causal contribution of multiple candidate cause nodes;

[0176] Core causal events are selected based on causal contribution scores to construct a subgraph.

[0177] In one possible implementation, building module 402 is used for:

[0178] The causal contribution score is calculated based on at least one of the following factors: temporal proximity, spatial proximity, pre-defined event correlation strength, and shared core business entities between the candidate cause node and the analysis anchor point.

[0179] In one possible implementation, building module 402 is used for:

[0180] Based on the spatiotemporal properties of at least one analysis anchor point, set the spatiotemporal constraint boundary;

[0181] Starting from any analysis anchor point, perform a temporal reverse traversal in the spatiotemporal knowledge graph;

[0182] During the reverse temporal traversal, if the current node exceeds the spatiotemporal constraint boundary, the path containing the current node is pruned until all paths have been traversed.

[0183] Based on the portion of the spatiotemporal knowledge graph retained after pruning, multiple candidate cause nodes containing potential causal chains related to the tracing problem are obtained; where the path includes nodes in the spatiotemporal knowledge graph and edges connected to the nodes.

[0184] In one possible implementation, the conversion module 403 is used for:

[0185] The structured information of the causal subgraph is transformed into narrative text containing weighted evidence;

[0186] Narrative text and tracing questions are embedded into pre-defined analysis prompt templates to obtain analysis prompts for the large model.

[0187] In one possible implementation, the conversion module 403 is used for:

[0188] Based on the causal contribution score of the core causal events in the causal subgraph, different strengths of evidence are assigned to the core causal events in the narrative text.

[0189] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.

[0190] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0191] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0192] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0193] like Figure 5 As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 508 into random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.

[0194] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0195] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as methods for analyzing historical operating data. For example, in some embodiments, the methods for analyzing historical operating data may be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the methods for analyzing historical operating data described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform methods for analyzing historical operating data by any other suitable means (e.g., by means of firmware).

[0196] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0197] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0198] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0199] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0200] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0201] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0202] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0203] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. An analysis method of historical operation data, comprising: determining at least one analysis anchor point in a time-space knowledge graph representing a state of a fleet operation according to a trace-back problem to be analyzed; tracing back associated relationships in the time-space knowledge graph from the at least one analysis anchor point to identify and construct a causal sub-graph containing a potential causal chain; transforming structured information of the causal sub-graph into an analysis prompt for a large model; obtaining an analysis result of the trace-back problem according to reasoning of the large model on the analysis prompt.

2. The method of claim 1, wherein, The determining at least one analysis anchor point in a time-space knowledge graph representing a state of a fleet operation according to a trace-back problem to be analyzed comprises: performing natural language understanding on the trace-back problem in text form to identify a trace-back intention and extract key entities; locating at least one analysis anchor point in a time-space knowledge graph representing a state of a fleet operation according to the trace-back intention and the key entities.

3. The method of claim 2, wherein, The locating at least one analysis anchor point in a time-space knowledge graph representing a state of a fleet operation according to the trace-back intention and the key entities comprises: in a case where the trace-back problem is an event-driven trace-back problem, directly searching for a target event node in the time-space knowledge graph according to an event identifier in the key entities; determining at least one entity node associated with the target event node as an analysis anchor point.

4. The method of claim 2, wherein, The locating at least one analysis anchor point in a time-space knowledge graph representing a state of a fleet operation according to the trace-back intention and the key entities comprises: in a case where the trace-back problem is a non-event-driven trace-back problem, inferring and generating a group of operation event nodes related to the trace-back problem in the time-space knowledge graph according to at least one of an entity identifier, a time range or index information of the key entities as the at least one anchor node.

5. The method of claim 1, wherein, The tracing back associated relationships in the time-space knowledge graph from the at least one analysis anchor point to identify and construct a causal sub-graph containing a potential causal chain comprises: performing time sequence reverse traversal in the time-space knowledge graph from the at least one analysis anchor point to obtain a plurality of candidate cause nodes; performing causal contribution degree scoring on the plurality of candidate cause nodes; filtering out core causal events according to the causal contribution degree scoring to construct the causal sub-graph.

6. The method of claim 5, wherein, The performing causal contribution degree scoring on the plurality of candidate cause nodes comprises: calculating the causal contribution degree scoring according to at least one of time proximity, space proximity, a preset event association strength and a shared core business entity between the candidate cause nodes and the analysis anchor point.

7. The method of claim 5, wherein, The performing time sequence reverse traversal in the time-space knowledge graph from the at least one analysis anchor point to obtain a plurality of candidate cause nodes comprises: setting a time-space constraint boundary according to time-space attributes of the at least one analysis anchor point; performing time sequence reverse traversal in the time-space knowledge graph from any analysis anchor point; in a case where a current node traversed exceeds the time-space constraint boundary in the time sequence reverse traversal process, pruning a path where the current node is located until all paths are traversed; According to the part of the spatio-temporal knowledge graph retained after pruning, a plurality of candidate cause nodes related to the traceability problem and containing potential causal chains are obtained; wherein, the path includes nodes of the spatio-temporal knowledge graph and edges connected to the nodes.

8. The method of claim 1, wherein, The structured information of the causal subgraph is converted into an analysis prompt for the large model, including: The structured information of the causal subgraph is converted into a narrative text containing weighted evidence. The narrative text and the traceability problem are embedded into a preset analysis prompt template to obtain an analysis prompt for the large model.

9. The method of claim 8, wherein, The structured information of the causal subgraph is converted into a narrative text containing weighted evidence, including: According to the causal contribution score of the core causal event in the causal subgraph, the core causal event is given different evidence strength descriptions in the narrative text.

10. The method of claim 1, wherein, The spatio-temporal knowledge graph is constructed based on multi-source operation data reflecting the operation state of unmanned vehicles, which is obtained from at least a plurality of data sources including a scheduling system, a vehicle terminal device, and a work order system.

11. An analysis device for historical operation data, comprising: an extraction module configured to determine at least one analysis anchor point in a spatio-temporal knowledge graph representing the operation state of a vehicle fleet according to a traceability problem to be analyzed; a construction module configured to trace the associated relationship in the spatio-temporal knowledge graph from the at least one analysis anchor point to identify and construct a causal subgraph containing potential causal chains; a conversion module configured to convert the structured information of the causal subgraph into an analysis prompt for a large model; a reasoning module configured to obtain an analysis result of the traceability problem according to the reasoning of the large model on the analysis prompt.

12. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-10.

13. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to make the computer execute the method according to any one of claims 1-10.

14. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-10.