Cloud-based vehicle cloud full-link observation and adaptive optimization method and system

CN122554346APending Publication Date: 2026-08-11DONGFENG MOTOR GRP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-20
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0008]本申请旨在解决现有技术中观测维度割裂、预警僵化、优化滞后、车端资源消耗过高等问题,提供一种基于观测云的车云全链路智能观测与自适应优化系统及方法

Benefits of technology

[0046] This application discloses a vehicle-to-infrastructure (V2I) end-to-end intelligent observation and adaptive optimization method and system based on an observation cloud. The optimization method includes: in the data acquisition phase, acquiring observation data, link quality data, and cloud-based operational status data reported by the vehicle based on a dynamic adjustment strategy for driving scenarios; in the data fusion phase, performing time-series formatting on the multi-source data and constructing a V2I-Road-Cloud end-to-end topology map; in the intelligent analysis phase, performing anomaly detection and trend prediction based on the topology map and a scenario-based dynamic baseline, and automatically locating the root cause node of the fault along the propagation path of the topology map; in the optimization execution phase, automatically issuing optimization instructions to adjust the operational status of the vehicle, link, or cloud based on the root cause location results, forming an adaptive closed loop. This application solves the problems of high false alarm rate and difficult root cause location in traditional solutions in dynamic V2I environments by using scenario-based dynamic baselines and root cause propagation maps, significantly improving the accuracy of fault warning and the efficiency of fault location.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122554346A_ABST
    Figure CN122554346A_ABST
Patent Text Reader

Abstract

This application discloses a vehicle-to-infrastructure (V2I) cloud-based intelligent observation and adaptive optimization method and system. The optimization method includes: in the data acquisition phase, acquiring observation data, link quality data, and cloud-based operational status data reported by the vehicle based on a dynamic adjustment strategy for driving scenarios; in the data fusion phase, performing time-series formatting on the multi-source data and constructing a V2I-road-cloud full-link topology map; in the intelligent analysis phase, performing anomaly detection and trend prediction based on the topology map and scenario-based dynamic baselines, and automatically locating the root cause node of the fault along the propagation path of the topology map; in the optimization execution phase, automatically issuing optimization instructions to adjust the operational status of the vehicle, link, or cloud based on the root cause location results, forming an adaptive closed loop. This application solves the problems of high false alarm rate and difficult root cause location in traditional solutions in dynamic V2I environments through scenario-based dynamic baselines and root cause propagation maps, significantly improving the accuracy of fault warning and the efficiency of fault location.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of vehicle networking, intelligent operation and maintenance and cloud computing technology, specifically to a system and method for achieving collaborative observation, intelligent analysis and adaptive optimization of data across the entire link between the vehicle, communication link and cloud based on an observation cloud platform. Background Technology

[0002] The high availability of vehicle-to-everything (V2X) systems relies on real-time observation and dynamic optimization of end-to-end data. However, existing technologies suffer from four major bottlenecks: fragmented observation dimensions, rigid early warning logic, passive and lagging optimization, and insufficient vehicle-side adaptation. As a mature end-to-end observation tool, the observation cloud has not achieved deep adaptation in the application of V2X scenarios.

[0003] Specifically, traditional vehicle-to-everything (V2X) monitoring adopts a single-point independent monitoring mode of "vehicle, link, and cloud" - the vehicle only focuses on the hardware status, the cloud only monitors the server load, and the communication link relies on the operator's basic monitoring. The data from the three cannot be linked to analyze the causal relationship of "vehicle lag - link latency - cloud service anomaly". In a certain car company, 45% of the V2X failures were caused by cross-link coordination problems, but due to the fragmented monitoring, it was impossible to quickly locate the problem.

[0004] In terms of early warning logic, the existing solution uses fixed threshold alarms (such as alarms triggered when server CPU > 85%), which is not adapted to dynamic scenarios of vehicle networking (such as link latency naturally increases when driving at high speeds, and battery parameters fluctuate at low temperatures), resulting in a false alarm rate of over 38%; and it lacks timing prediction capabilities, making it impossible to predict potential faults in advance, with 80% of faults only being discovered after user feedback.

[0005] Regarding optimization actions, the observed data is only used for fault alarms and is not linked with the vehicle cloud system control logic—the vehicle-side data collection frequency is fixed and the cloud resource configuration is static. Even if problems such as "low computing power ECU overload" and "link bandwidth waste" are observed, manual optimization commands are still required, with a response delay of more than 2 hours.

[0006] Regarding vehicle-side adaptation, existing monitoring tools are not optimized for low-computing-power ECUs on the vehicle side. Full data collection increases the ECU load by more than 15%, and massive data transmission consumes 30% of the vehicle-to-cloud communication bandwidth, which in turn affects the stability of core functions.

[0007] In summary, existing technologies cannot meet the core requirements of the Internet of Vehicles (IoV) for "end-to-end collaborative observation, scenario-based intelligent early warning, and adaptive dynamic optimization." The end-to-end data integration and intelligent analysis capabilities of the observation cloud provide a foundation for solving these problems, but existing technologies have not achieved deep adaptation between the observation cloud and IoV scenarios, necessitating the construction of a dedicated intelligent observation and adaptive optimization system. Summary of the Invention

[0008] This application aims to address the problems of fragmented observation dimensions, rigid early warning systems, delayed optimization, and excessive vehicle-side resource consumption in existing technologies, and provides a vehicle-cloud end-to-end intelligent observation and adaptive optimization system and method based on an observation cloud. To achieve the above objectives, this application adopts a four-layer closed-loop architecture of "lightweight acquisition layer - time-series fusion layer - intelligent analysis layer - adaptive optimization layer" to achieve deep adaptation of the observation cloud to the vehicle network scenario. The specific technical solution is as follows:

[0009] In a first aspect, embodiments of this application provide a vehicle-cloud end-to-end observation and adaptive optimization method based on observation cloud, including:

[0010] During the data acquisition phase, vehicle-side observation data is obtained from the lightweight acquisition agent deployed on the vehicle, which dynamically adjusts the acquisition strategy based on the vehicle driving scenario. Link quality data of the communication link is also obtained, as well as the operational status data of the cloud service.

[0011] During the data fusion phase, the vehicle-side observation data, the link quality data, and the operational status data are processed by time-series formatting, and a full-link topology map representing the dependencies between the vehicle-side, the link, and the cloud is constructed and updated based on the formatted data.

[0012] During the intelligent analysis phase, based on the full-link topology map and the pre-built dynamic baseline corresponding to the vehicle driving scenario, anomaly detection and trend prediction are performed on the formatted data. When an anomaly is predicted, the root cause node is located by analyzing the propagation path of the fault in the full-link topology map.

[0013] During the optimization execution phase, based on the location results of the root cause nodes and the preset optimization strategy, optimization instructions are generated and issued to adjust the operating status of at least one link in the vehicle, communication link, or cloud, so as to form an adaptive closed loop of observation-analysis-optimization.

[0014] Furthermore, during the data acquisition phase, vehicle-side observation data is obtained, specifically including:

[0015] The lightweight acquisition agent is instructed to dynamically adjust the acquisition frequency of different types of vehicle-side indicators based on the driving scene identification results obtained from the observation cloud platform and / or the abnormal status of indicators obtained through real-time analysis.

[0016] The strategy for dynamically adjusting the data collection frequency includes: automatically increasing the data collection frequency when core functions are running or when indicator values ​​are abnormal, and automatically decreasing the data collection frequency when non-core functions are running or when system resource utilization exceeds a preset threshold.

[0017] Furthermore, during the data fusion phase, a full-link topology map representing the dependencies between the vehicle, the link, and the cloud is constructed and updated, specifically including:

[0018] By utilizing the service topology automatic discovery capability of the observation cloud, and combining the service call relationships and communication endpoint information parsed from the vehicle-side observation data and the operation status data, an initial topology map containing vehicle-side functional domains, communication link nodes, and cloud microservice nodes is automatically generated.

[0019] Based on real-time collected link quality data and node health status data, the connection status and health indicators between nodes in the initial topology diagram are dynamically updated.

[0020] Furthermore, the construction method of the pre-built dynamic baseline corresponding to the vehicle driving scenario includes:

[0021] Acquire time-series indicator data from the vehicle, network, and cloud that are tagged with driving scenarios within a historical time window;

[0022] Time-series index data under the same driving scenario label are aggregated, and statistical models or machine learning clustering algorithms are used to calculate the normal fluctuation range of each index under different scenarios, so as to generate multiple dynamic baselines corresponding to different driving scenarios.

[0023] The threshold range of the dynamic baseline is adaptively switched according to the real-time identified vehicle driving scenario.

[0024] Furthermore, when an anomaly is predicted during the intelligent analysis phase, the root cause node is located by analyzing the propagation path of the fault in the entire link topology diagram. Specifically, this includes:

[0025] When at least one node's metric is detected to deviate from its corresponding dynamic baseline, backtracking analysis is performed starting from that node and along the dependencies in the full-link topology graph.

[0026] Using a graph neural network model, the temporal correlation, spatial correlation, and causal contribution between abnormal indicators of upstream nodes and fault phenomena of downstream nodes are calculated, and a root cause map representing the fault propagation relationship is constructed.

[0027] The node with the highest contribution in the root cause graph and no other fault-introduced edges is identified as the root cause node.

[0028] Furthermore, the intelligent analysis phase also includes:

[0029] The comprehensive availability score of the vehicle cloud system is calculated based on the operating status of core functions, the degree of deviation of indicators from the dynamic baseline, and historical fault data.

[0030] When the overall availability score is lower than a preset threshold, the optimization execution phase is automatically triggered.

[0031] Furthermore, the optimization instructions generated and issued during the optimization execution phase include at least one vehicle-side optimization instruction, which is used to perform any of the following operations:

[0032] Adjust the data collection strategy of the lightweight data collection agent, including reducing the collection frequency of non-core indicators or pausing their collection;

[0033] Trigger the vehicle-side sensors to perform a self-calibration process;

[0034] When the vehicle's battery level falls below a preset threshold, data reporting for non-core functions is suspended to ensure communication bandwidth for core functions.

[0035] Furthermore, the optimization instructions generated and issued during the optimization execution phase include at least one optimization instruction for the communication link, which is used to perform any of the following operations:

[0036] The system instructs the vehicle-side communication module to switch to a backup communication link when the service quality index of the current communication link falls below a preset switching threshold.

[0037] The vehicle-mounted communication module is instructed to allocate different transmission bandwidth ratios to service data streams of different priorities.

[0038] Furthermore, the optimization instructions generated and issued during the optimization execution phase include at least one cloud-based optimization instruction, which is used to perform any of the following operations:

[0039] Trigger cloud resource orchestration services to elastically scale up or down the computing node cluster that provides specific services;

[0040] Update the routing policy of the cloud load balancer to direct new service requests to less loaded compute nodes.

[0041] Secondly, embodiments of this application provide a vehicle-to-cloud end-to-end observation and adaptive optimization system based on an observation cloud, including:

[0042] The lightweight acquisition module is used to acquire vehicle-side observation data reported by the lightweight acquisition agent deployed on the vehicle, which is obtained by dynamically adjusting the acquisition strategy based on the vehicle driving scenario, acquire link quality data of the communication link, and acquire cloud service operation status data.

[0043] The time-series fusion module is used to perform time-series formatting processing on the vehicle-side observation data, the link quality data, and the operating status data, and to construct and update a full-link topology map representing the dependency relationship between the vehicle-side, the link, and the cloud based on the formatted data.

[0044] The intelligent analysis module is used to perform anomaly detection and trend prediction on the formatted data based on the full-link topology map and a pre-built dynamic baseline corresponding to the vehicle driving scenario. When an anomaly is predicted, the root cause node is located by analyzing the propagation path of the fault in the full-link topology map.

[0045] The adaptive optimization module is used to generate and issue optimization instructions to adjust the operating status of at least one link in the vehicle, communication link or cloud based on the location results of the root cause node and the preset optimization strategy, so as to form an adaptive closed loop of observation-analysis-optimization.

[0046] This application discloses a vehicle-to-infrastructure (V2I) end-to-end intelligent observation and adaptive optimization method and system based on an observation cloud. The optimization method includes: in the data acquisition phase, acquiring observation data, link quality data, and cloud-based operational status data reported by the vehicle based on a dynamic adjustment strategy for driving scenarios; in the data fusion phase, performing time-series formatting on the multi-source data and constructing a V2I-Road-Cloud end-to-end topology map; in the intelligent analysis phase, performing anomaly detection and trend prediction based on the topology map and a scenario-based dynamic baseline, and automatically locating the root cause node of the fault along the propagation path of the topology map; in the optimization execution phase, automatically issuing optimization instructions to adjust the operational status of the vehicle, link, or cloud based on the root cause location results, forming an adaptive closed loop. This application solves the problems of high false alarm rate and difficult root cause location in traditional solutions in dynamic V2I environments by using scenario-based dynamic baselines and root cause propagation maps, significantly improving the accuracy of fault warning and the efficiency of fault location. Attached Figure Description

[0047] Figure 1 This application provides a core flowchart of a vehicle-cloud end-to-end observation and adaptive optimization method based on observation cloud as an embodiment of the present application.

[0048] Figure 2 This is a schematic diagram of the module structure of a vehicle-cloud full-link observation and adaptive optimization system based on observation cloud, provided in an embodiment of this application. Detailed Implementation

[0049] To enable those skilled in the art to better understand the technical solutions of this application, exemplary embodiments of this application are described below with reference to the accompanying drawings, including various details of the embodiments of this application to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description. Unless otherwise specified, the various embodiments of this application and the features within those embodiments can be combined with each other.

[0050] As used herein, the term "and / or" includes any and all combinations of one or more of the associated enumerated entries. The terminology used herein is for describing particular embodiments only and is not intended to limit the application. As used herein, the singular forms "a" and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that when the terms "comprising" and / or "made of" are used herein, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Terms such as "connected" or "linked" are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.

[0051] Unless otherwise specified, all terms used in this application (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It should also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this application, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined in this application.

[0052] refer to Figure 1 and Figure 2 One embodiment of this application proposes a vehicle-to-cloud full-link observation and adaptive optimization method and system based on observation cloud. Through a four-layer closed-loop design of "lightweight acquisition layer - time series fusion layer - intelligent analysis layer - adaptive optimization layer", it realizes deep adaptation between observation cloud and vehicle network scenario.

[0053] I. Data Acquisition Phase – Implementation of the Lightweight Acquisition Layer

[0054] During the data acquisition phase, vehicle-side observation data is obtained from the lightweight acquisition agent deployed on the vehicle, which dynamically adjusts the acquisition strategy based on the vehicle driving scenario. Link quality data of the communication link is also obtained, as well as the operational status data of the cloud service.

[0055] 1. Lightweight vehicle-side data acquisition module

[0056] To achieve efficient and low-load data acquisition from the vehicle, this application designs a lightweight observation agent deployed on the vehicle. This agent has a deployment size of less than 5 MB and operates using a "data filtering + local preprocessing" approach. It only reports key data, after preliminary screening and feature extraction, to the observation cloud platform, thereby significantly reducing the computational resource consumption of the vehicle's ECU.

[0057] The core innovation of the lightweight data acquisition agent lies in its dynamic data acquisition frequency adjustment mechanism. By interacting with the observation cloud platform, the agent obtains the current driving scenario identification results of the vehicle (such as "high-speed driving", "low-speed urban driving", "weak network in tunnel", "parking and charging" etc.) and the abnormal status of indicators obtained in real time, and then dynamically adjusts the data acquisition frequency of different types of vehicle indicators.

[0058] Specifically, the data collected includes ECU load, sensor operating status (LiDAR / camera), cockpit chip GPU utilization, core function operation logs (OTA transmission progress, remote control execution results), and vehicle driving status (vehicle speed, ambient temperature, driving mode). The data collection frequency adjustment strategy includes:

[0059] Scenario-adaptive adjustment: When the vehicle is in a "high-speed driving" scenario, core parameters related to driving safety, such as braking system status and ECU load, are collected once per second, while non-core parameters, such as historical fuel consumption and window status, are collected at a lower frequency of once every 10 or 30 seconds to reduce bus load and communication bandwidth usage. When the vehicle enters a "parking and charging" scenario, since driving functions are not affected, the collection frequency of all parameters can be moderately increased from the base value (e.g., from 10 seconds / time to 5 seconds / time) to obtain more detailed battery status data for health analysis.

[0060] Anomaly-triggered encryption: When the intelligent analysis layer of the observation cloud detects a sudden 10% increase in the CPU load of an ECU within one minute, it determines that the node has a potential anomaly. It then sends a command to the lightweight data acquisition agent on the vehicle side to temporarily encrypt the collection frequency of relevant indicators for that ECU to 0.5 seconds per acquisition, in order to capture the complete process of the anomaly event. For non-core indicators, collection is not required by the system default state; collection is only triggered when the observation cloud detects anomalies in related indicators (such as a sudden increase in link latency). This strategy can reduce the amount of data by 70%.

[0061] Through the above mechanism, the computational load of the lightweight acquisition agent can be reduced by 80% compared with the traditional full acquisition tool, and the load impact on low computing power ECUs can be controlled within a very small range, thus effectively solving the pain point of high consumption of vehicle-side observation resources.

[0062] 2. Communication Link Acquisition Module

[0063] The communication link acquisition module embeds a lightweight probe from the observation cloud into the vehicle-mounted T-BOX or communication gateway to achieve real-time acquisition of quality indicators such as 5G / 4G link latency, packet loss rate, and bandwidth utilization. The latency measurement accuracy can reach the 1 ms level. Crucially, each link quality data point is automatically tagged with the current scene on the vehicle (such as "high-speed driving," "weak network area," or "low-temperature environment") upon generation, providing a data foundation for subsequent scene-based analysis.

[0064] 3. Cloud-based data collection module

[0065] The cloud-based data acquisition module utilizes a cloud-native acquisition plugin to synchronize real-time metrics such as server CPU / memory load, database response time, application service error codes (e.g., OTA upgrade service return codes), and cloud computing resource utilization. This module supports elastic scaling to accommodate concurrent access from massive numbers of vehicles, ensuring high availability and scalability of cloud-based monitoring data acquisition.

[0066] II. Data Fusion Phase – Implementation of the Temporal Fusion Layer

[0067] During the data fusion phase, the vehicle-side observation data, the link quality data, and the operational status data are processed by time-series formatting, and a full-link topology map representing the dependencies between the vehicle, the link, and the cloud is constructed and updated based on the formatted data.

[0068] 1. Data normalization processing

[0069] Because the vehicle-side binary data, link JSON data, and cloud log data have different formats, the time series fusion layer first performs format unification processing on all data. All data is converted into the observation cloud standard time series format, which includes standard fields such as timestamp, indicator name, value, and scene label. For example, the ECU load data reported by the vehicle, "{'ecu_id':'VCU', 'load':45, 'ts':1638345600, 'scene':'highway'}", and the CPU load data reported by the cloud, "{'host':'ota-srv-01', 'cpu_usage':75, 'timestamp':1638345600, 'scene':null}", are both standardized into the standard format of "measurement: 'system.load', tags: {'source': 'vehicle', 'node': 'VCU', 'scene': 'highway'}, fields: {'value': 45}, timestamp: 1638345600". This standardization makes cross-vehicle, link, and cloud-based correlation queries possible. For example, users can easily query "the correlation between ECU load and link latency for a certain vehicle model in a highway scenario".

[0070] In addition, this application adopts a hot and cold separation storage strategy: high-frequency core data (such as link latency and ECU load) is stored in the observation cloud hot data node and retained for 7 days for quick query and analysis; low-frequency non-core data (such as historical fault logs) is stored in the cold data node and retained for 90 days, reducing storage costs by about 50%.

[0071] 2. Construction and updating of the end-to-end topology map

[0072] The time-series fusion layer leverages the service topology auto-discovery capability of the observation cloud platform, combining service call relationships and communication endpoint information parsed from vehicle-side observation data and cloud-based operational status data to automatically generate an initial topology diagram containing vehicle-side functional domains, communication link nodes, and cloud-based microservice nodes. For example, by analyzing the target IP address in the vehicle-side T-BOX reporting logs and the service discovery records in the cloud-based microservice registry, the system can automatically construct the dependency relationship of "vehicle-side remote control functional domain" → "uplink command link" → "cloud-based remote control microservice".

[0073] During system operation, this topology is not static. The fusion layer dynamically updates the connection status and health indicators of each node in the topology based on real-time collected link quality data and node health status data. For example, when the latency of the "uplink command link" exceeds 80 ms or the packet loss rate exceeds 1%, the status indicator of that link node will be updated from "healthy" (green) to "degraded" (yellow) or "faulty" (red); when the CPU load of a certain instance of the "remote control microservice" in the cloud exceeds 85%, the node indicator of that instance will be updated to "overloaded" (orange). This dynamically updated full-link topology provides a precise contextual environment for subsequent root cause localization.

[0074] III. Intelligent Analysis Stage – Implementation of the Intelligent Analysis Layer

[0075] During the intelligent analysis phase, based on the full-link topology map and the pre-built dynamic baseline corresponding to the vehicle driving scenario, anomaly detection and trend prediction are performed on the formatted data. When an anomaly is predicted, the root cause node is located by analyzing the propagation path of the fault in the full-link topology map.

[0076] 1. Scenario-based time-series early warning model

[0077] Traditional fixed threshold alarms cannot adapt to the dynamically changing operating environment of the Internet of Vehicles (IoV). Therefore, this application constructs a scenario-based time-series early warning model, the core of which is the generation and application of scenario-based dynamic baselines. The construction process of the dynamic baseline is as follows: Time-series indicator data from the vehicle, network, and cloud, labeled with driving scenario tags, are acquired within a historical time window (e.g., the past 3 months); clustering algorithms such as K-Means are used to aggregate time-series indicator data under the same driving scenario tag, and the normal fluctuation range of each indicator under different scenarios is calculated, thereby generating multiple dynamic baselines corresponding to different scenarios such as "high-speed movement," "low-speed urban traffic," "weak tunnel network," and "low-temperature environment."

[0078] In real-time monitoring, the system adaptively switches the baseline threshold used for anomaly detection based on the scene labels reported by vehicles in real time. For example, the baseline for the "link latency" metric is 20-80 ms in the "high-speed movement" scenario, while it is 60-150 ms in the "weak tunnel network" scenario. When a vehicle enters a tunnel, the scene label switches to "weak tunnel network," and a link latency of 85 ms will be considered normal and will not trigger an alarm. This mechanism significantly reduces the false alarm rate.

[0079] For trend prediction, this application employs a deep learning model combining a Long Short-Term Memory (LSTM) network and an attention mechanism. The model is trained using three months of historical time-series data (such as ECU load and link latency) and can predict the trend of indicators for the next 15 minutes. If the predicted value exceeds 80% of the current scenario-based baseline threshold, the system immediately triggers an alert. The alert triggering logic can be expressed as: when P(Value) exceeds 80% of the current scenario-based baseline threshold, the system immediately triggers an alert. t +15)>Baseline Threshold A warning is issued when (scene) × 80%. Where P(Value) t +15) represents the model's prediction of the indicator value for the next 15 minutes. Baseline Threshold (scene) represents the upper limit of the baseline threshold corresponding to the current scene. Furthermore, the overall early warning score can consider scene weights: Early Warning Score = Σ (scene weights) i ×Indicator Deviation i The weight of weak network scenarios can be set to 1.2, and the weight of normal temperature city scenarios can be set to 0.9, so as to reflect the difference in the severity of anomalies in different scenarios.

[0080] 2. End-to-end root cause automatic localization engine

[0081] Once an anomaly is detected and an alert is triggered, the root cause analysis engine starts working. The analysis process is as follows:

[0082] The starting point for backtracking analysis is determined by starting with the node where an anomaly is observed (such as "cloud OTA service response timeout") and tracing back upstream along the dependencies in the full-link topology graph.

[0083] Correlation Contribution Calculation: Utilizing a Graph Neural Network (GNN) model, this study integrates multi-dimensional data such as vehicle status, link quality, and cloud load to calculate the temporal correlation, spatial correlation, and causal contribution between abnormal upstream node indicators and downstream failure phenomena. For example, in the "remote control failure" scenario, the GNN model analysis yields a contribution of 85% for "vehicle-side subscription topic failure," 10% for "link latency exceeding the threshold," and 5% for "high cloud control service load."

[0084] Root cause graph construction and root cause identification: Based on the contribution calculation results, a root cause graph representing the fault propagation relationship is constructed, highlighting the fault source node and its propagation path. Finally, the node with the highest contribution in the root cause graph and no other fault-incoming edges is identified as the root cause node. In this example, "vehicle-side subscription topic failure" is identified as the root cause. The entire location process takes less than 5 minutes, with an accuracy rate exceeding 98%.

[0085] 3. Comprehensive assessment of vehicle-cloud status

[0086] Beyond anomaly detection for specific metrics, the intelligent analysis layer conducts a higher-level comprehensive availability assessment of the vehicle-cloud system. Based on the operational status of core functions (OTA upgrades, remote control, navigation), the degree to which each observed metric deviates from the dynamic baseline, and historical fault data, the system constructs an availability scoring model to calculate a comprehensive availability score (0-100 points) for the vehicle-cloud system in real time. The scoring model comprehensively considers factors such as the success rate of core functions (60% weight), the quality of critical links (20% weight), and the load of cloud service nodes (20% weight). When the comprehensive availability score falls below a preset threshold (e.g., 90 points), even without triggering an alarm for any single metric exceeding its limit, the system is classified as "sub-healthy" and automatically triggers a subsequent adaptive optimization phase, achieving a shift from passively responding to faults to proactively preventing risks.

[0087] IV. Optimization Execution Phase – Implementation of the Adaptive Optimization Layer

[0088] During the optimization execution phase, based on the root cause node location results and preset optimization strategies, optimization instructions are generated and issued to adjust the operational status of at least one of the vehicle-side, communication link, or cloud components, forming an adaptive closed loop of observation-analysis-optimization. The adaptive optimization layer receives alarm events and root cause location results generated by the intelligent analysis layer and automatically triggers optimization actions for the vehicle-side, link, or cloud based on a preset optimization strategy library. This closed-loop linkage mechanism is one of the core innovations of this application, upgrading the observation cloud from a simple observation tool into an intelligent operation and maintenance brain with automatic repair capabilities.

[0089] 1. Vehicle-side adaptive optimization

[0090] The optimization commands for the vehicle side mainly include:

[0091] Data Acquisition Strategy Optimization: When the load of a low-computing-power ECU is observed to be consistently above 70%, the adaptive optimization controller sends a command to the lightweight acquisition agent on that ECU to automatically reduce the acquisition frequency of non-core indicators (e.g., from 10 seconds / time to 30 seconds / time) to free up ECU computing resources. The original frequency is restored after the load decreases.

[0092] Functional degradation adaptation: When the vehicle battery SOC (State of Charge) is observed to be below 10%, the system automatically sends a degradation policy to the vehicle's infotainment system, suspending the reporting of logs and statistics of all functions unrelated to driving, such as the entertainment system and app store, and reserving the limited power and communication bandwidth for core functions such as remote control, emergency call, and fault alarm.

[0093] Sensor calibration trigger: When system analysis detects abnormal fluctuations (such as periodic jumps in values) in the data reported by a tire pressure sensor over a period of time, deviating from the historical baseline, a self-calibration command is automatically sent to the sensor. After the sensor executes its internal calibration procedure, the reported data returns to normal, with a calibration success rate exceeding 95%.

[0094] 2. Adaptive optimization of communication links

[0095] The optimization instructions for communication links mainly include:

[0096] Dynamic link switching: When the packet loss rate of the currently connected 5G link is observed to exceed a preset switching threshold of 3%, the adaptive optimization controller immediately sends a link switching command to the vehicle-side T-BOX. Within 100 ms, the T-BOX switches the data transmission channel from 5G to a more stable 4G backup link. When the 5G signal quality is subsequently observed to recover (packet loss rate <0.5% and remains stable), a command is sent to switch back to the 5G link to take advantage of its high bandwidth.

[0097] Bandwidth allocation on demand: When the vehicle is simultaneously downloading high-priority OTA upgrade packages and updating low-priority offline maps, the controller optimizes the bandwidth allocation strategy, instructing the vehicle communication module to allocate 80% of the available bandwidth to the OTA download task and only 20% to the map update task, thereby ensuring the transmission efficiency of critical tasks.

[0098] 3. Cloud-based adaptive optimization

[0099] The optimization commands for the cloud mainly include:

[0100] Elastic scaling of computing power: When a sudden 50% increase in vehicle access volume within a geographical area is observed in a short period (e.g., 10 minutes), causing the average CPU load of the corresponding cloud access gateway cluster to exceed 80%, the adaptive optimization controller automatically adds 30% more compute node instances to the cluster by calling the cloud service provider's API. The scaling process is initiated within 200 ms, and the new nodes are online to handle traffic within minutes, effectively avoiding service overload. When access volume decreases, the system automatically scales down idle nodes, reducing resource costs and improving cloud resource utilization by more than 30%.

[0101] Service routing optimization: When a node instance in the cloud OTA service cluster is observed to have a CPU utilization of 90% due to processing large upgrade packages, the optimization controller issues an instruction to the cluster's front-end load balancer to temporarily mark the high-load node as unavailable. New OTA download requests will be automatically routed to other healthy nodes in the cluster with CPU loads below 50%. Once the CPU utilization of the high-load node drops below 70% and remains stable, the system issues another instruction to add it back to the service list, balancing service pressure.

[0102] V. Core Process Example: Taking "OTA Upgrade Anomaly Warning and Optimization" as an Example

[0103] To make the technical solution of this application clearer, the following uses a complete OTA upgrade anomaly handling process as an example to connect the above four stages of the work process.

[0104] 1. Observation and Data Collection: A vehicle model undergoes an OTA upgrade while driving on a highway. The vehicle-side observation agent collects OTA transmission progress and ECU load data once per second; the link probe detects a current 5G link latency of 65 ms and a packet loss rate of 2.8%; the cloud-based data collection plugin detects an OTA service CPU load of 75%.

[0105] 2. Time series fusion: The observation cloud normalizes the above multi-source data into a standard time series format, adds a "high-speed driving" scenario label, and updates the full-link topology diagram of "vehicle-side T-BOX - 5G base station - core network - cloud OTA service".

[0106] 3. Intelligent Analysis: The intelligent analysis layer utilizes the dynamic baseline (normal link packet loss rate range of 0-1%) under the "high-speed driving" scenario. If the current packet loss rate of 2.8% exceeds the baseline threshold, and the LSTM model predicts a further increase in the packet loss rate within the next 15 minutes, an alert is triggered. The root cause localization engine, based on topology graph and GNN model analysis, identifies "5G link packet loss" as the core root cause of slow OTA transmission.

[0107] 4. Adaptive Optimization: Based on the root cause location results of "5G link packet loss", the adaptive optimization controller immediately executes the preset strategy: sends a command to the vehicle to switch the data link from 5G to 4G (the packet loss rate drops to 0.3% after the switch); at the same time, in order to ensure upgrade efficiency, it sends a command to the cloud to temporarily expand the computing power of the OTA service cluster in this area by 30%.

[0108] 5. Closed-loop verification: The observation cloud continuously monitors the 4G link status and OTA transmission progress after the switch, confirms that the transmission rate has returned to normal, and the optimization actions take effect, forming a complete closed loop.

[0109] As can be seen from this embodiment, the optimization method of this application can automatically discover potential problems, accurately locate the root cause, and autonomously execute optimization operations without human intervention throughout the process, which greatly improves the reliability, operation and maintenance efficiency, and user experience of the vehicle network system.

[0110] The aforementioned embodiments of the optimization methods and the embodiments of the optimization systems are identical or related in technical concept, and can be referred to each other in terms of technical details and technical effectiveness, which will not be repeated here.

[0111] Overall, the beneficial effects of this application compared to the prior art mainly include:

[0112] 1. Resource consumption issues in vehicle-side data collection

[0113] Traditional vehicle-to-everything (V2X) monitoring tools are not optimized for low-performance ECUs (Engine Control Units). They employ a full-volume, fixed-frequency data collection mode, leading to an over 15% increase in ECU load and consuming approximately 30% of the vehicle-to-cloud communication bandwidth, negatively impacting the stability of core functions such as remote control and navigation. This application addresses this by deploying a lightweight vehicle-side data collection agent (less than 5 MB in size) and innovatively introducing a "scenario-based dynamic data collection frequency + anomaly-triggered data collection" strategy. This strategy adaptively adjusts the collection frequency based on the vehicle's driving scenario (e.g., high-speed driving, parking and charging) and real-time load status, encrypting data collection during core function operation or abnormal situations, and reducing the collection frequency during non-core function operation or resource constraints. Simultaneously, the agent performs data filtering and feature extraction preprocessing on the vehicle side, reporting only critical data. This reduces the impact of vehicle-side data collection on ECU load by over 80% and vehicle-to-cloud communication bandwidth consumption by approximately 70%, ensuring data integrity while preventing interference with core driving functions from the monitoring system.

[0114] 2. Issues regarding the accuracy and false alarm rate of early warnings.

[0115] Existing technologies use fixed threshold alarms (e.g., alarming when server CPU usage exceeds 85%), failing to consider the variability of the dynamic operating environment of the vehicle network. For example, link latency naturally increases during high-speed driving, and battery parameters fluctuate normally in low-temperature environments. Fixed thresholds cannot adapt to these changing scenarios, resulting in a false alarm rate exceeding 38%, leaving maintenance personnel overwhelmed with dealing with invalid alarms. This application constructs a "scenario-based time-series early warning model," generating multi-scenario dynamic baselines by integrating vehicle driving scenario labels (high-speed, urban, weak network, low temperature, etc.) to replace traditional static thresholds. In real-time monitoring, the system adaptively switches the anomaly detection baseline based on the vehicle's current scenario. Simultaneously, a deep learning model using LSTM + attention mechanisms predicts the trend of indicators for the next 15 minutes, providing early warning when the predicted value exceeds the scenario-based baseline threshold by 80%. As a result, the false alarm rate is significantly reduced by approximately 40%, the fault warning accuracy is improved to over 98%, and proactive warnings more than 15 minutes in advance are successfully achieved, enabling maintenance to shift from passive response to proactive defense.

[0116] 3. The efficiency of root cause localization for cross-domain failures.

[0117] In existing technologies, observation data from the vehicle, the road network, and the cloud are fragmented, lacking a global collaborative perspective. When cross-domain faults such as "vehicle-road-cloud service anomalies" occur, maintenance personnel need to log into different systems one by one and manually analyze logs. For one automaker, 45% of faults originated from cross-link collaboration issues, yet the root cause could not be quickly located, with average location time measured in hours. This application utilizes an automatically constructed and dynamically updated vehicle-road-cloud full-link topology map built on the observation cloud, combined with graph neural networks (GNNs) to perform correlation analysis on multi-dimensional indicators. It automatically calculates the causal contribution of upstream node anomalies to downstream faults, generates a visualized fault propagation graph, and identifies the node with the highest contribution and no fault-incoming edges as the root cause node. In this way, the root cause location time for cross-domain faults is reduced from hours to less than 5 minutes, with a location accuracy exceeding 98%, significantly improving maintenance efficiency and fault handling speed.

[0118] 4. Closed-loop linkage problem between observation and control

[0119] In existing technologies, observation data is only used for fault alarms and post-event analysis, and is independent of the control logic of the vehicle, the link, and the cloud. Even if the observation system detects problems such as "ECU overload" or "wasted link bandwidth," manual analysis and manual issuance of optimization commands are still required, resulting in a response lag of more than 2 hours, and the optimization action is passive and slow. This application designs an adaptive closed-loop mechanism of "observation-analysis-optimization." Abnormal events and root cause localization results generated by the intelligent analysis layer directly trigger the adaptive optimization layer, which automatically generates and issues commands according to the preset optimization strategy to adjust the vehicle-side data acquisition strategy, communication link switching, and cloud computing power scaling in real time. The optimization effect is then continuously verified by the observation layer. In this way, the optimization response latency is compressed from hours to milliseconds (e.g., link switching <100 ms, cloud expansion <200 ms), realizing the leap from "observable" to "self-healing" of the system, and significantly improving the efficiency of dynamic resource adaptation.

[0120] 5. The efficiency of cloud resource utilization

[0121] In existing technologies, cloud resource configuration is static and cannot be elastically adjusted according to the ebb and flow of vehicle access volume. Insufficient resources during peak periods lead to service degradation, while idle resources during off-peak periods result in waste, leading to overall low resource utilization. This application, based on real-time monitoring of regional vehicle access volume and cloud service load by an observation cloud, automatically triggers cloud resource orchestration services for elastic scaling—automatically expanding computing nodes when access volume surges and automatically shrinking and reclaiming resources when access volume decreases. Simultaneously, service routing optimization directs new requests to low-load nodes, balancing cluster pressure. In this way, cloud computing resource utilization can be improved by more than 30%, effectively reducing annual bandwidth and computing costs for automakers while ensuring service quality during peak periods.

[0122] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and / or operation of possible implementations of systems, methods, and / or computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0123] Exemplary embodiments have been disclosed in this application, and while specific terminology has been used, it is used only and should be interpreted in a general illustrative sense and is not intended to be limiting. In some embodiments, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this application as set forth by the appended claims.

Claims

1. An observation cloud-based vehicle cloud full-link observation and adaptive optimization method, characterized in that, include: During the data acquisition phase, vehicle-side observation data is obtained from the lightweight acquisition agent deployed on the vehicle, which dynamically adjusts the acquisition strategy based on the vehicle driving scenario. Link quality data of the communication link is also obtained, as well as the operational status data of the cloud service. During the data fusion phase, the vehicle-side observation data, the link quality data, and the operational status data are processed by time-series formatting, and a full-link topology map representing the dependencies between the vehicle-side, the link, and the cloud is constructed and updated based on the formatted data. During the intelligent analysis phase, based on the full-link topology map and the pre-built dynamic baseline corresponding to the vehicle driving scenario, anomaly detection and trend prediction are performed on the formatted data. When an anomaly is predicted, the root cause node is located by analyzing the propagation path of the fault in the full-link topology map. During the optimization execution phase, based on the location results of the root cause nodes and the preset optimization strategy, optimization instructions are generated and issued to adjust the operating status of at least one link in the vehicle, communication link, or cloud, so as to form an adaptive closed loop of observation-analysis-optimization.

2. The optimization method according to claim 1, characterized in that, The data acquisition phase involves obtaining vehicle-mounted observation data, specifically including: The lightweight acquisition agent is instructed to dynamically adjust the acquisition frequency of different types of vehicle-side indicators based on the driving scene identification results obtained from the observation cloud platform and / or the abnormal status of indicators obtained through real-time analysis. The strategy for dynamically adjusting the data collection frequency includes: automatically increasing the data collection frequency when core functions are running or when indicator values ​​are abnormal, and automatically decreasing the data collection frequency when non-core functions are running or when system resource utilization exceeds a preset threshold.

3. The optimization method according to claim 1, characterized in that, In the data fusion phase, a full-link topology map representing the dependencies between the vehicle, the link, and the cloud is constructed and updated, specifically including: By utilizing the service topology automatic discovery capability of the observation cloud, and combining the service call relationships and communication endpoint information parsed from the vehicle-side observation data and the operation status data, an initial topology map containing vehicle-side functional domains, communication link nodes, and cloud microservice nodes is automatically generated. Based on real-time collected link quality data and node health status data, the connection status and health indicators between nodes in the initial topology diagram are dynamically updated.

4. The optimization method according to claim 1, characterized in that, The construction methods for the pre-built dynamic baseline corresponding to the vehicle driving scenario include: Acquire time-series indicator data from the vehicle, network, and cloud that are tagged with driving scenarios within a historical time window; Time-series index data under the same driving scenario label are aggregated, and statistical models or machine learning clustering algorithms are used to calculate the normal fluctuation range of each index under different scenarios, so as to generate multiple dynamic baselines corresponding to different driving scenarios. The threshold range of the dynamic baseline is adaptively switched according to the real-time identified vehicle driving scenario.

5. The optimization method according to claim 1, characterized in that, When an anomaly is predicted during the intelligent analysis phase, the root cause node is located by analyzing the propagation path of the fault in the full-link topology diagram. This specifically includes: When at least one node's metric is detected to deviate from its corresponding dynamic baseline, a backtracking analysis is performed starting from that node and along the dependencies in the full-link topology graph. Using a graph neural network model, the temporal correlation, spatial correlation, and causal contribution between abnormal indicators of upstream nodes and fault phenomena of downstream nodes are calculated, and a root cause map representing the fault propagation relationship is constructed. The node with the highest contribution in the root cause graph and no other fault-introduced edges is identified as the root cause node.

6. The optimization method according to claim 1, characterized in that, The intelligent analysis phase also includes: The comprehensive availability score of the vehicle cloud system is calculated based on the operating status of core functions, the degree of deviation of indicators from the dynamic baseline, and historical fault data. When the overall availability score is lower than a preset threshold, the optimization execution phase is automatically triggered.

7. The optimization method according to claim 1, characterized in that, The optimization instructions generated and issued during the optimization execution phase include at least one vehicle-side optimization instruction, which is used to perform any of the following operations: Adjust the data collection strategy of the lightweight data acquisition agent, including reducing the collection frequency of non-core indicators or pausing their collection; Trigger the vehicle-side sensors to perform a self-calibration process; When the vehicle's battery level falls below a preset threshold, data reporting for non-core functions is suspended to ensure communication bandwidth for core functions.

8. The optimization method according to claim 1, characterized in that, The optimization instructions generated and issued during the optimization execution phase include at least one optimization instruction for the communication link, which is used to perform any of the following operations: The system instructs the vehicle-side communication module to switch to a backup communication link when the service quality index of the current communication link falls below a preset switching threshold. The vehicle-mounted communication module is instructed to allocate different transmission bandwidth ratios to service data streams of different priorities.

9. The optimization method according to claim 1, characterized in that, The optimization instructions generated and issued during the optimization execution phase include at least one cloud-based optimization instruction, which is used to perform any of the following operations: Trigger cloud resource orchestration services to elastically scale up or down the computing node cluster that provides specific services; Update the routing policy of the cloud load balancer to direct new service requests to less loaded compute nodes.

10. A vehicle-cloud end-to-end observation and adaptive optimization system based on observation cloud, characterized in that, include: The lightweight acquisition module is used to acquire vehicle-side observation data reported by the lightweight acquisition agent deployed on the vehicle, which is obtained by dynamically adjusting the acquisition strategy based on the vehicle driving scenario, acquire link quality data of the communication link, and acquire cloud service operation status data. The time-series fusion module is used to perform time-series formatting processing on the vehicle-side observation data, the link quality data, and the operating status data, and to construct and update a full-link topology map representing the dependency relationship between the vehicle-side, the link, and the cloud based on the formatted data. The intelligent analysis module is used to perform anomaly detection and trend prediction on the formatted data based on the full-link topology map and a pre-built dynamic baseline corresponding to the vehicle driving scenario. When an anomaly is predicted, the root cause node is located by analyzing the propagation path of the fault in the full-link topology map. The adaptive optimization module is used to generate and issue optimization instructions to adjust the operating status of at least one link in the vehicle, communication link or cloud based on the location results of the root cause node and the preset optimization strategy, so as to form an adaptive closed loop of observation-analysis-optimization.