A method and apparatus for SOVD hidden fault detection
Patent Information
- Application Number
- CN202611229091.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-13
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]有鉴于此,有必要提供一种SOVD隐蔽故障检测方法及装置,用以解决现有技术中SOVD客户端主要关注响应内容(故障码、数据值等),响应时间仅用于超时判断或简单的性能监控,未作为诊断信息本身进行分析,从而无法从时间维度的统计变化中发现这些隐蔽性故障的技术问题
[0016]本发明的有益效果是:接收当前ECU在当前时间窗口内的当前ECU处理时间,并获取车辆的原始RTT序列以及对各ECU进行周期探测的网络往返时间;原始RTT序列包括HTTP请求集合、以及每个HTTP请求对应的ECU编号、时间戳、原始响应时间和API端点编号;将每个HTTP请求的原始响应时间和网络往返时间之间的差值组成的集合作为HTTP请求集合对应的ECU处理时间集合;根据ECU处理时间集合中的每个ECU处理时间和原始RTT序列中每个ECU处理时间对应的ECU编号、API端点编号和时间戳建立响应时间的正常基线;根据正常基线与当前ECU处理时间进行计算,得到综合异常分数,并根据综合异常分数进行故障判断,得到隐蔽故障检测结果;本发明将ECU处理时间作为了诊断信息本身进行分析,确定了响应时间的正常基线,对当前ECU处理时间进行异常检测,实现了从时间维度的统计变化中发现隐蔽性故障的目的。
Smart Images

Figure CN122845474A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fault detection technology, and in particular to a method and apparatus for detecting hidden faults in SOVD (Self-Imposed Vulnerability Defect). Background Technology
[0002] The SOVD (Service-Oriented Vehicle Diagnostics) protocol is based on HTTPS (Hypertext Transfer Security Protocol). Clients invoke the diagnostic service via HTTP requests, and the server returns a response. Each request-response process generates a measurable response time (the time interval from sending a request to receiving a complete response). ISO 17978-3 defines a timeout mechanism, allowing the client to interrupt the wait if the response time exceeds a preset threshold (e.g., 30 seconds).
[0003] Currently, ECUs exhibit slow-response faults, such as memory leaks leading to decreased processing speed, task scheduling deadlocks, and frequent watchdog resets. These faults do not initially trigger any diagnostic fault codes or generate abnormal measurements, but they significantly alter the response time distribution of diagnostic requests. However, the SOVD client primarily focuses on the response content (fault codes, data values, etc.), using response time only for timeout checks or simple performance monitoring, and is not analyzed as diagnostic information itself. Therefore, these hidden faults cannot be detected through statistical changes over time.
[0004] Therefore, there is an urgent need to propose a method and device for detecting hidden faults in SOVD, which can solve the technical problem that in the existing technology, SOVD clients mainly focus on response content (fault codes, data values, etc.), and the response time is only used for timeout judgment or simple performance monitoring, without being analyzed as diagnostic information itself, thus failing to discover these hidden faults from statistical changes in the time dimension. Summary of the Invention
[0005] In view of this, it is necessary to provide a method and device for detecting hidden faults in SOVD, in order to solve the technical problem that in the existing technology, SOVD clients mainly focus on response content (fault codes, data values, etc.), and the response time is only used for timeout judgment or simple performance monitoring, and is not analyzed as diagnostic information itself, so these hidden faults cannot be discovered from the statistical changes in the time dimension.
[0006] To address the aforementioned problems, firstly, this invention provides a method for detecting hidden faults in SOVD, applied to SOVD clients; Receive the current ECU processing time within the current time window, and obtain the vehicle's original RTT sequence and the network round-trip time for periodic probing of each ECU; the original RTT sequence includes a set of HTTP requests, and the ECU number, timestamp, original response time, and API endpoint number corresponding to each HTTP request; The set of differences between the original response time and the network round-trip time of each HTTP request is taken as the ECU processing time set corresponding to the HTTP request set; A normal baseline for response time is established based on the ECU processing time in the ECU processing time set and the ECU number, API endpoint number, and timestamp corresponding to the ECU processing time in the original RTT sequence; A comprehensive anomaly score is calculated based on the normal baseline and the current ECU processing time. Fault judgment is then performed based on the comprehensive anomaly score to obtain the hidden fault detection result.
[0007] In one possible implementation, the original RTT sequence generation process involves recording the request sending time when an HTTP request is sent to the corresponding ECU based on the diagnostic information, recording the complete response reception time when the corresponding ECU returns an HTTP response, determining the original response time based on the difference between the complete response reception time and the request sending time, and saving the original RTT sequence of the current HTTP request.
[0008] In one possible implementation, the network round-trip time acquisition process involves periodically sending ICMP Echo requests to each ECU, and determining the network round-trip time upon receiving a response from the corresponding ECU; wherein the detection period for sending the ICMP Echo request to each ECU is 2 to 5 times the diagnostic request period.
[0009] In one possible implementation, the normal baseline includes an ECU-level baseline; establishing a normal baseline for response time based on the ECU processing time in the ECU processing time set and the ECU number, API endpoint number, and timestamp corresponding to each ECU processing time in the original RTT sequence includes: Based on the ECU number and the API endpoint number, the ECU processing time of all API endpoints under the same ECU is input into a Gaussian mixture model for multimodal distribution fitting, and the weight, mean, and variance of each mode are output; the modes in the multimodal distribution include light request peak, medium request peak, and heavy routine peak; An ECU-level baseline is determined based on the weights, mean, and variance of each mode.
[0010] In one possible implementation, the normal baseline includes an API endpoint-level baseline; the step of establishing a normal baseline for response time based on the ECU processing time in the ECU processing time set and the ECU number, API endpoint number, and timestamp corresponding to the ECU processing time in the original RTT sequence further includes: The ECU processing time for the same API endpoint is calculated based on the API endpoint number and kernel density estimation to obtain the baseline mean, standard deviation and fast / slow boundary of the corresponding API endpoint. The API endpoint-level baseline is determined based on the baseline mean, the standard deviation, and the fast / slow boundary for each API endpoint.
[0011] In one possible implementation, the normal baseline includes a time-dependent baseline; the step of establishing a normal baseline for response time based on the normal baseline for each ECU processing time in the ECU processing time set and the ECU number, API endpoint number, and timestamp corresponding to each ECU processing time in the original RTT sequence further includes: After aligning the ECU processing times of different API endpoints within the same time window under the same ECU based on the ECU number, the API endpoint number, and the timestamp, the Pearson correlation coefficient between the ECU processing times is calculated to obtain the response time Pearson correlation coefficient matrix, and the response time Pearson correlation coefficient matrix is used as the time correlation baseline.
[0012] In one possible implementation, the step of calculating a comprehensive anomaly score based on the normal baseline and the current ECU processing time includes: The current ECU processing time is simulated and calculated based on the ECU-level baseline to obtain the KL divergence; when the KL divergence is greater than a preset divergence threshold, the global degradation anomaly score of the ECU is calculated based on the KL divergence. For each API endpoint, calculate the deviation multiple between the current ECU processing time and the API endpoint-level baseline. When the deviation multiple is greater than a preset standard deviation, calculate the endpoint anomaly score based on the deviation multiple. Based on the current ECU processing time, calculate the correlation coefficient matrix between endpoints within the current window, and calculate the norm difference between the correlation coefficient matrix and the time correlation baseline; when the norm difference is greater than a preset difference threshold, calculate the resource competition anomaly score based on the norm difference; The weighted sum of the ECU global degradation anomaly score, the endpoint anomaly score, and the resource contention anomaly score is used to obtain the comprehensive anomaly score.
[0013] In one possible implementation, the step of determining the fault based on the comprehensive anomaly score to obtain the hidden fault detection result includes: An exponentially weighted moving average is applied to the comprehensive anomaly score to obtain a smoothed value sequence; Calculate the first derivative based on the smoothed value sequence; Fault judgment is performed based on all first derivatives to obtain the hidden fault detection results.
[0014] In one possible implementation, the step of determining the fault based on all first derivatives to obtain the hidden fault detection result includes: When the first derivative is continuously positive within a preset window, the fault type is determined to be a trend warning, triggering an active diagnosis process. The active diagnosis process involves controlling the SOVD client to increase the diagnostic sampling frequency of the corresponding ECU, and collecting and diagnosing data for subsequent time based on the diagnostic sampling frequency to determine a new comprehensive anomaly score. When the new comprehensive anomaly score of multiple consecutive windows is greater than the anomaly threshold, the hidden fault detection result is determined to be a hidden fault, and the API endpoint with the hidden fault is reported through the SOVD client.
[0015] Secondly, the present invention also provides an SOVD hidden fault detection device, comprising: The data acquisition module is used to receive the current ECU processing time within the current time window, and to acquire the vehicle's original RTT sequence and the network round-trip time for periodic detection of each ECU; the original RTT sequence includes a set of HTTP requests, and the ECU number, timestamp, original response time and API endpoint number corresponding to each HTTP request; The set determination module is used to take the set of differences between the original response time and the network round-trip time of each HTTP request as the ECU processing time set corresponding to the HTTP request set; The baseline determination module is used to establish a normal baseline for response time based on the ECU processing time of each ECU in the ECU processing time set and the ECU number, API endpoint number, and timestamp corresponding to the ECU processing time in the original RTT sequence; The score calculation module is used to calculate a comprehensive anomaly score based on the normal baseline and the current ECU processing time, and to make a fault judgment based on the comprehensive anomaly score to obtain the hidden fault detection result.
[0016] The beneficial effects of this invention are as follows: It receives the current ECU processing time within the current time window and obtains the vehicle's original RTT sequence and the network round-trip time for periodic detection of each ECU. The original RTT sequence includes a set of HTTP requests, and the ECU number, timestamp, original response time, and API endpoint number corresponding to each HTTP request. The set of differences between the original response time and network round-trip time of each HTTP request is used as the ECU processing time set corresponding to the HTTP request set. A normal baseline for the response time is established based on the ECU processing time in the ECU processing time set and the ECU number, API endpoint number, and timestamp corresponding to each ECU processing time in the original RTT sequence. A comprehensive anomaly score is calculated based on the normal baseline and the current ECU processing time, and fault judgment is performed based on the comprehensive anomaly score to obtain the hidden fault detection result. This invention analyzes the ECU processing time as diagnostic information itself, determines the normal baseline for the response time, and performs anomaly detection on the current ECU processing time, achieving the goal of discovering hidden faults from statistical changes in the time dimension. Attached Figure Description
[0017] Figure 1 A schematic flowchart of an embodiment of the SOVD hidden fault detection method provided by the present invention; Figure 2 This is a schematic diagram of an embodiment of the HTTP request recording process provided by the present invention; Figure 3 A schematic diagram of an embodiment of the multi-level baseline modeling diagram provided by the present invention; Figure 4 For the present invention Figure 1 A schematic flowchart of an embodiment of step S104; Figure 5 A schematic diagram of an embodiment of the proactive early warning process provided by the present invention; Figure 6 This is a schematic diagram of an embodiment of the SOVD concealed fault detection device provided by the present invention. Detailed Implementation
[0018] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.
[0019] like Figure 1 As shown, a specific embodiment of the present invention discloses a method for detecting hidden faults in SOVD, which is applied to an SOVD client; S101. Receive the current ECU processing time within the current time window, and obtain the vehicle's original RTT sequence and the network round-trip time for periodic probing of each ECU; the original RTT sequence includes the HTTP request set, and the ECU number, timestamp, original response time and API endpoint number corresponding to each HTTP request.
[0020] In this embodiment of the invention, when the vehicle is driving normally, the SOVD client automatically initiates a round of diagnostic probes for each ECU (electronic control unit) at fixed intervals (e.g., every 5 minutes or every 10 minutes) to obtain the network round-trip time of each ECU. If the vehicle is not currently initiating any diagnostic requests, there is no data to collect at that moment, and the collection action is in an idle state. When the vehicle's dashboard warning light illuminates, or when the vehicle is brought in for repair, the repair tools or the cloud will initiate temporary HTTP requests. At this time, the collection module will record the information of the HTTP requests. The SOVD Client or the vehicle gateway can record a precise timestamp for each HTTP request, forming the original RTT sequence of response times. That is, the original RTT sequence is the data of all HTTP requests in the historical data already stored by the collection module. To diagnose the vehicle's current condition, it is also necessary to collect the current ECU processing time within the current time window (e.g., every 10 minutes or every 500 requests), which can also be represented as the current ECU's actual processing time.
[0021] S102. The set of differences between the original response time and the network round-trip time of each HTTP request is taken as the ECU processing time set corresponding to the HTTP request set.
[0022] In this embodiment of the invention, the ECU processing time for each HTTP request can be determined based on the difference between the original response time and the network round-trip time for each HTTP request. Thus, based on the ECU processing times for all HTTP requests, a set of ECU processing times is obtained. Pure network latency is separated from the total response time.
[0023] S103. Establish a normal baseline for response time based on the ECU processing time in the ECU processing time set and the ECU number, API endpoint number, and timestamp corresponding to the ECU processing time in the original RTT sequence.
[0024] In order to accurately detect anomalies in this embodiment of the invention, it is necessary to establish a normal baseline for response time based on the processing time of each ECU in the ECU processing time set, and the ECU number, API endpoint number and timestamp corresponding to the processing time of each ECU in the original RTT sequence.
[0025] S104. Calculate the comprehensive anomaly score based on the normal baseline and the current ECU processing time, and make a fault judgment based on the comprehensive anomaly score to obtain the hidden fault detection result.
[0026] In this embodiment of the invention, the current ECU processing time can be sample data collected within the current time window (e.g., 10 minutes or every 500 requests). This allows for comparison of the current ECU processing time with the ECU-level baseline, API endpoint-level baseline, and time-related baseline within the normal baseline. A comprehensive anomaly score is obtained based on the comparison results of each baseline. Fault judgment is then performed based on the comprehensive anomaly scores of all time windows to obtain hidden fault detection results and achieve predictive early warning.
[0027] The SOVD hidden fault detection method provided in this invention can be applied to an SOVD hidden fault detection system. The SOVD hidden fault detection can be a software system running on a terminal device. The terminal device can be a tablet computer, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), mobile phone, etc. This application embodiment does not impose any restrictions on the specific type of terminal device. The system can be set up on an SOVD client. The SOVD client can transmit data with multiple ECUs. After receiving data from multiple ECUs, the SOVD client performs SOVD hidden fault detection, obtains the hidden fault detection result, and then transmits the hidden fault detection result through an output interface. For example, it can transmit the result to a display screen for displaying the fault data, or transmit the data to other settings for fault processing.
[0028] Compared with existing technologies, this embodiment provides the following method: receiving the current ECU processing time within the current time window, obtaining the vehicle's original RTT sequence, and the network round-trip time for periodic detection of each ECU; the original RTT sequence includes a set of HTTP requests, and the ECU number, timestamp, original response time, and API endpoint number corresponding to each HTTP request; the set of differences between the original response time and network round-trip time of each HTTP request is used as the ECU processing time set corresponding to the HTTP request set; a normal baseline for response time is established based on the ECU processing time in the ECU processing time set and the ECU number, API endpoint number, and timestamp corresponding to each ECU processing time in the original RTT sequence; a comprehensive anomaly score is calculated based on the normal baseline and the current ECU processing time, and fault judgment is performed based on the comprehensive anomaly score to obtain the hidden fault detection result; this invention analyzes the ECU processing time as diagnostic information itself, determines the normal baseline for response time, and performs anomaly detection on the current ECU processing time, achieving the purpose of discovering hidden faults from statistical changes in the time dimension.
[0029] In some embodiments of the present invention, the original RTT sequence generation process is as follows: when sending an HTTP request to the corresponding ECU based on the diagnostic information, the request sending time is recorded; when receiving the HTTP response returned by the corresponding ECU, the complete response receiving time is recorded; the original response time is determined based on the difference between the complete response receiving time and the request sending time; and the original RTT sequence of the current HTTP request is saved.
[0030] In the embodiments of the present invention, as follows Figure 2 As shown, Figure 2To record the HTTP request process, when the SOVD client sends an HTTP request to the corresponding ECU based on the diagnostic information, it can record the sending time T_req. After receiving the HTTP request, the ECU returns complete response data to the SOVD client. When the SOVD client receives the complete response data, it records the complete response reception time T_resp. Then, the difference between the complete response reception time and the request sending time is used to determine the original response time RTT_raw = T_resp - T_req. Obvious outliers (such as outliers >5 seconds due to network retransmissions) are then removed, retaining only valid samples. The data is stored separately according to the ECU identifier and API endpoint, resulting in the stored data for the current HTTP request. This stored data can be stored in the SOVD client and may include the original response time RTT_raw, ECU number ECU_ID, API endpoint, and timestamp. The timestamp can be the "event occurrence moment" accurate to milliseconds / microseconds. Under the SOVD protocol and standard HTTP communication mechanism, each HTTP request strictly corresponds to and only corresponds to the HTTP response returned by a specific ECU. This is a one-to-one relationship. In the automotive Ethernet, each ECU supporting SOVD is assigned a unique IP address (or has a unique routing port configured at the gateway). For example, when the diagnostic information is "left front radar ECU", the SOVD client determines the target URL of the HTTP request to be https: / / 192.168.1.10 / ... (the IP address of this ECU) based on the one-to-one relationship in the "left front radar ECU". This data packet, through the vehicle switch, will only be routed to the ECU with IP address 192.168.1.10. Only this ECU receives and processes the request, and finally returns a response along the same path. Other ECUs (such as 192.168.1.11) will not receive this request packet at the physical and network layers, let alone respond. Each record in the original RTT sequence corresponds to one and only one specific ECU.
[0031] In some embodiments of the present invention, the process of obtaining network round-trip time involves periodically sending ICMP Echo requests to each ECU, and determining the network round-trip time upon receiving a response from the corresponding ECU; wherein the detection period for sending ICMP Echo requests to each ECU is 2 to 5 times the diagnostic request period.
[0032] In this embodiment of the invention, upon startup, the SOVD client first obtains a list of all target ECUs supporting the SOVD protocol within the current diagnostic domain through a vehicle topology discovery mechanism, and establishes a one-to-one mapping between ECU identifiers (ECU IDs) and IP addresses. For each ECU in the list (e.g., central gateway ECU, cockpit domain controller, left front radar ECU, etc.), the SOVD client allocates an independent probe state machine in memory to store the real-time diagnostic request frequency and the most recent ICMP probe result for that ECU. To achieve a balance between obtaining sufficiently accurate network latency data and avoiding increased bus congestion, the SOVD client periodically sends ICMP Echo requests to each ECU to measure the network round-trip time (RTT_net). The probe period for the SOVD client to send ICMP Echo requests to each ECU is 2 to 5 times the diagnostic request period to avoid additional congestion. When the SOVD client intercepts a SOVD HTTP request destined for a specific ECU, it uses that specific ECUID as an index to find the most recent valid measurement RTT_net before the current time among all measured network round-trip times. The raw response time RTT_raw is then determined as the SOVD client's response time RTT_sovd. For each diagnostic request, the SOVD response time RTT_sovd includes: network latency (uplink + downlink) + ECU internal processing time. Approximately: RTT_sovd ≈ RTT_net + T_proc, where T_proc is the ECU processing time (including request parsing, operation execution, and response construction). Therefore, the decoupled processing time is: T_proc = RTT_sovd - RTT_net (note that RTT_net already includes round-trip network latency, but diagnostic request packets are slightly larger; a correction factor can be added in engineering practice; the default factor is 1.0). When network jitter is severe, multiple pings can be used to obtain the moving average RTT_net_avg to improve robustness. For example, the SOVD Client periodically sends ICMP Echo requests to the target ECU, with a cycle 2-5 times the diagnostic request cycle. The round-trip time (RTT_net_instant) of the ICMP response is recorded, and the average of this time using a sliding window is used to obtain RTT_net_avg. The decoupling processing time is calculated as: T_proc = RTT_sovd - RTT_net_avg × γ, where γ is the path correction coefficient (diagnostic request packets are slightly larger than ICMP packets; γ defaults to 1.0~1.2 and can be calibrated offline). When RTT_net_avg exceeds three times the baseline network latency, the network is marked as abnormal during that period, and detection is paused or its weight is reduced.
[0033] To accurately detect anomalies, a normal baseline for response time needs to be established. A normal baseline is as follows: Figure 3 As shown, Figure 3For multi-level baseline modeling, normal baselines may include ECU-level baselines, API endpoint-level baselines, and time-dependent baselines. In some embodiments of the present invention, the construction of ECU-level baselines includes step S103, which includes: Based on the ECU number and API endpoint number, the ECU processing time of all API endpoints under the same ECU is input into the Gaussian mixture model for multimodal distribution fitting, and the weight, mean, and variance of each mode are output. The modes in the multimodal distribution include light request peak, medium request peak, and heavy routine peak.
[0034] In this embodiment of the invention, the Gaussian Mixture Model (GMM) is a classic probability density model, commonly used in statistics and machine learning for cluster analysis and fitting complex data distributions. The number of components is 3. Based on the timestamp stored for each HTTP request in the original RTT sequence, the original response time, the ECU number corresponding to the returned HTTP response, and the API endpoint number, the ECU processing time of all API endpoints under the same ECU is collected. Then, the ECU processing time is input into the Gaussian Mixture Model. Since different request types have different loads, the GMM can fit multimodal distributions (such as light request peaks, medium request peaks, and heavy routine peaks), storing the weight, mean, and variance of each mode.
[0035] An ECU-level baseline is determined based on the weights, mean, and variance of each mode.
[0036] In this embodiment of the invention, after obtaining the weight, mean, and variance of each mode, a baseline GMM curve is constructed for the weight, mean, and variance of each mode. This curve is a standard normal mixture curve, i.e., an ECU-level baseline.
[0037] In some embodiments of the present invention, the construction of the API endpoint-level baseline further includes step S103: The ECU processing time for the same API endpoint is calculated based on the API endpoint number and kernel density estimation, yielding the baseline mean, standard deviation, and fast / slow boundary for the corresponding API endpoint.
[0038] In this embodiment of the invention, because the data attributes are singular, kernel density estimation (KDE) is used instead of ground truth model (GMM) to calculate the ECU processing time for the same API endpoint. Since a single endpoint distribution is usually unimodal, kernel density estimation (KDE) is used to calculate the baseline mean μ and standard deviation σ of this endpoint. 2 And the fast and slow dividing lines (10%, 50%, 90th percentile). Among them, KDE is the specific method for "building" the baseline model, and the final output baseline model is the probability density curve fitted by KDE.
[0039] The API endpoint-level baseline is determined based on the baseline mean, standard deviation, and fast / slow boundary for each API endpoint.
[0040] In this embodiment of the invention, a baseline model is established separately for each API endpoint of each ECU (or even each DID / routine ID). Based on the baseline mean, standard deviation and fast / slow boundary of each API endpoint, the API endpoint-level baseline of each API endpoint of each ECU is determined, i.e., the baseline model.
[0041] In some embodiments of the present invention, step S103 further includes the construction of the time-related baseline: After aligning the ECU processing times of different API endpoints within the same time window under the same ECU based on ECU number, API endpoint number, and timestamp, the Pearson correlation coefficient between ECU processing times is calculated to obtain the response time Pearson correlation coefficient matrix, which is then used as the time correlation baseline.
[0042] In this embodiment of the invention, the ECU processing time of different API endpoints within the same time window under the same ECU is determined based on the ECU number and API endpoint number. Then, the ECU processing times are aligned based on timestamps, and the Pearson correlation coefficient between ECU processing times is calculated. For example, if it is found that "API_A" is fast and "API_B" is also fast, indicating a strong correlation, the Pearson correlation coefficient is calculated. Based on the Pearson correlation coefficients of all ECU processing times, a response time Pearson correlation coefficient matrix is constructed, forming a "time fingerprint" matrix. The response time Pearson correlation coefficient matrix is used as the time correlation baseline, and the time correlation baseline serves as the "time fingerprint" in normal mode. For example, under normal conditions, the response time for reading fault codes and reading freeze frames should be highly correlated (both are lightweight operations). If the correlation disappears on a certain day, it indicates that one of the API paths is abnormal.
[0043] Quantitative references for normal states are provided through ECU-level baselines, API endpoint-level baselines, and time-correlation baselines, providing a benchmark for subsequent offset detection and supporting fine-grained anomaly localization.
[0044] In some embodiments of the present invention, such as Figure 4 As shown, step S104 includes: S401. Simulate and calculate the current ECU processing time based on the ECU-level baseline to obtain the KL divergence; when the KL divergence is greater than the preset divergence threshold, calculate the ECU global degradation anomaly score based on the KL divergence.
[0045] In this embodiment of the invention, an actual distribution histogram (or empirical distribution curve) is plotted based on the current ECU processing time. The actual distribution histogram is compared with the ECU-level baseline to obtain the KL divergence, which measures the "area / morphological difference" between these two curves. KL divergence: D_KL(P_cur||P_base)=Σ_xP_cur(x)·ln(P_cur(x) / P_base(x)), where P_cur is the actual distribution histogram and P_base is the ECU-level baseline. That is, KL divergence is a dynamically calculated difference value, measuring the morphological distance between the current data distribution curve and this baseline curve. When the KL divergence D_KL> a preset divergence threshold θ_ECU (e.g., 0.5), the overall performance of the ECU is determined to have degraded. Then, the KL divergence can be calculated to obtain the ECU global degradation anomaly score S_ECU, calculated as: S_ECU=min(D_KL / 0.5,0.1). If the KL divergence ≤ the preset divergence threshold, then the ECU global degradation anomaly score S_ECU is indeed 0.
[0046] S402. For each API endpoint, calculate the deviation multiple between the current ECU processing time and the API endpoint-level baseline. When the deviation multiple is greater than the preset standard deviation multiple, calculate the endpoint anomaly score based on the deviation multiple.
[0047] In this embodiment of the invention, for each API endpoint, the deviation factor z-score between the current ECU processing time and the API endpoint-level baseline is calculated. The formula is z-score=(μ_cur-μ_base) / σ_base, where μ_cur is the mean of the current ECU processing time, and the API endpoint-level baseline includes the baseline mean μ (i.e., μ_base) and the standard deviation σ. 2 (i.e., σ_base). If z-score > a preset multiple of standard deviation (for example, if the preset multiple of standard deviation is 3, it means exceeding 3 times the standard deviation), the endpoint is determined to be abnormal. The largest z value among all endpoints is taken: z_max = max_k(z_k). The endpoint abnormality score is calculated based on z_max, and the endpoint-level offset abnormality score is: S_API = min(z_max / 3, 1.0). If z-score ≤ the preset multiple of standard deviation, then the endpoint abnormality score is indeed 0.
[0048] S403. Based on the current ECU processing time, calculate the correlation coefficient matrix between endpoints within the current window, and calculate the norm difference between the correlation coefficient matrix and the time correlation baseline; when the norm difference is greater than the preset difference threshold, calculate the resource competition anomaly score based on the norm difference.
[0049] In this embodiment of the invention, based on the current ECU processing time, the correlation coefficient matrix between endpoints within the current window is calculated. Then, the temporal correlation baseline is subtracted from the correlation coefficient matrix to obtain the Frobenius norm difference, ΔCorr=||Corr_cur-Corr_base||_F, where Corr_base is the temporal correlation baseline and Corr_cur is the correlation coefficient matrix. If the norm difference is greater than a preset difference threshold, the resource contention anomaly score S_Corr=min(ΔCorr / θ_Corr,1.0) is calculated based on the norm difference. If the norm difference is ≤ the preset difference threshold, the resource contention anomaly score is indeed 0. The preset difference threshold is the 99.9th quantile of the normal difference.
[0050] S404. The weighted sum of the ECU global degradation anomaly score, endpoint anomaly score, and resource contention anomaly score is obtained to obtain the comprehensive anomaly score.
[0051] In this embodiment of the invention, the global degradation anomaly score S_ECU, the endpoint anomaly score S_API, and the resource contention anomaly score S_Corr are weighted and summed to obtain the comprehensive anomaly score S_anomaly. The weights w1, w2, and w3 of the global degradation anomaly score, endpoint anomaly score, and resource contention anomaly score are 0.4, 0.4, and 0.2, respectively. That is, S_anomaly = 0.4 × S_ECU + 0.4 × S_API + 0.2 × S_Corr.
[0052] A single outlier score may be affected by short-term fluctuations, thus requiring trend analysis. In some embodiments of the present invention, step S104 includes: An exponentially weighted moving average is applied to the overall anomaly scores to obtain a smoothed value sequence; Calculate the first derivative based on the smoothed value sequence; Fault judgment is performed based on all first derivatives to obtain the hidden fault detection results.
[0053] In the embodiments of the present invention, as follows Figure 5 As shown, Figure 5For the proactive early warning process, after the offset detection engine outputs anomaly scores, an exponentially weighted moving average (EWMA) is applied to the comprehensive anomaly score to obtain a smoothed value: S_smooth(t) = α·S_raw(t) + (1-α)·S_smooth(t-1), α = 0.3. Here, S_raw(t) is the "comprehensive anomaly score S_anomaly" within the t-th time window without any historical smoothing, and S_smooth(t-1) is the smoothed value at time t-1. This exponentially weighted moving average acts as a "low-pass filter," with a weight α = 0.3, meaning that the instantaneous judgment (SrawSraw) at the current moment only accounts for 30% of the final decision weight. (1-α) indicates that the long-term trend S_smooth(t-1) at the previous moment accounts for 70% of the final decision weight. This ensures that the final score no longer closely follows each instantaneous fluctuation (such as network spikes or occasional load peaks), but rather reflects the overall trend over a recent period. Ensure that subsequent alarms are based on "chronic illness" rather than "sneezing". Then, calculate the smoothed value for each time window using the formula to obtain a smoothed value sequence, such as S_smooth(1), S_smooth(2), S_smooth(3), ..., S_smooth(t). Next, perform trend analysis on each window based on the smoothed value sequence. Specifically, the first derivative (slope) slope(t) = (S_smooth(t) - S_smooth(tW)) / W, where W = 10 windows, to obtain the first derivative (slope) of the t-th window.
[0054] In some embodiments of the present invention, fault judgment is performed based on all first derivatives to obtain hidden fault detection results, including: When the first derivative remains positive within a preset window, the fault type is determined to be a trend warning, triggering an active diagnostic process. The active diagnostic process involves controlling the SOVD client to increase the diagnostic sampling frequency of the corresponding ECU, and collecting and diagnosing data for subsequent time periods based on the diagnostic sampling frequency to determine a new comprehensive anomaly score.
[0055] In this embodiment of the invention, the strategy linkage determines whether the first derivative is continuously positive within a preset window. For example, if the first derivative is continuously positive for more than 10 windows, even if the current score is below the threshold, the fault type is determined to be a trend warning, and the time to reach the threshold is estimated as T_arrive=(θ-S_smooth(t)) / slope_mean, where slope_mean is the average value (i.e., average rate of rise) of the first derivative of the smoothed sequence within the past W consecutive time windows that triggered the warning. Then, a prompt message "ECU response time continues to deteriorate, and it is expected to reach the abnormal threshold after T_arrive hours" can be sent, triggering an active diagnosis process. The active diagnosis process involves controlling the SOVD client extension interface to send a suggestion to the Client to increase the diagnostic sampling frequency of the relevant ECU / API endpoints (e.g., from 1Hz to 10Hz) to obtain more granular data to confirm the root cause. Then, data is collected and diagnosed for subsequent time according to the diagnostic sampling frequency to obtain a new comprehensive abnormal score calculated from the data of subsequent time. The process of calculating the new comprehensive abnormal score is the same as described above, and will not be repeated in this embodiment of the invention. If the condition that "the first derivative is continuously positive within a preset number of windows" is not met, no processing will be performed temporarily.
[0056] When the new comprehensive anomaly scores of multiple consecutive windows are all greater than the anomaly threshold, the hidden fault detection result is determined to be a hidden fault, and the API endpoint with the hidden fault is reported through the SOVD client.
[0057] In this embodiment of the invention, after adjusting the diagnostic sampling frequency to obtain a new comprehensive anomaly score, if the new comprehensive anomaly score for multiple consecutive windows is greater than the anomaly threshold (e.g., the new comprehensive anomaly score for three consecutive windows is greater than the anomaly threshold), then the hidden fault detection result is determined to be a hidden fault, and the API endpoint with the hidden fault is reported through the SOVD client. If a specific endpoint anomaly is detected simultaneously, it is further recommended to check the specific service of that endpoint (e.g., "Routine 0x1234 execution time is abnormal, it is recommended to check code efficiency"). The SOVD response time diagnostic service extension includes, for example: adding an API: GET, / diagnostics / response_time / stats?target={ecu_id}&window={minutes}, which returns the statistical information of the current window (sample size, mean, variance, KL divergence, z-score, anomaly score, trend slope). Adding an API: GET, / diagnostics / response_time / timeseries?target={ecu_id}&hours={h}, which returns the historical T_proc sequence (for offline analysis). A new Identification type, ResponseTimeDriftWarning, has been added, which includes the following fields: ecu_id, anomaly score, offset type (global degradation / endpoint anomaly / resource contention), list of affected endpoints, and recommended actions (such as "check ECU CPU load" or "check routine 0x1234 code efficiency").
[0058] Furthermore, for vehicles that are continuously under high load (such as those undergoing OTA downloads), the baseline model can temporarily relax the threshold to avoid false alarms.
[0059] This invention upgrades passive detection to proactive early warning, forming a closed loop of detection → early warning → enhanced sampling → reporting, thereby achieving predictive maintenance.
[0060] This invention extends SOVD diagnostics from "response content analysis" to "response time distribution analysis," detecting hidden ECU degradation through statistical distribution offset. Specifically: Phase 1: Collect the response time RTT_raw for each SOVD request and store it by ECU and API endpoint.
[0061] Phase 2: Decouple the network latency RTT_net by measuring the network latency with ICMP assistance, and obtain the ECU processing time T_proc = RTT_sovd - RTT_net.
[0062] Phase 3: Establish a three-level baseline, namely, ECU-level Gaussian mixture model, API endpoint-level kernel density estimation, and inter-endpoint correlation coefficient matrix.
[0063] Phase 4: Calculate the KL divergence, z-score, and correlation difference of the current distribution, and weight them together to form a comprehensive anomaly score.
[0064] Phase 5: Perform EWMA smoothing and slope trend detection on abnormal scores. If the scores continue to deteriorate, trigger a trend warning and actively increase the sampling frequency. If the threshold is exceeded, report hidden faults.
[0065] The above five stages are connected in sequence to form a complete closed loop of "acquisition → decoupling → baseline → detection → early warning".
[0066] To better implement the SOVD hidden fault detection method in the embodiments of the present invention, correspondingly, the embodiments of the present invention also provide an SOVD hidden fault detection device, such as... Figure 6 As shown, the SOVD concealed fault detection device 600 includes: The data acquisition module 601 is used to receive the current ECU processing time within the current time window, and to acquire the vehicle's original RTT sequence and the network round-trip time for periodic detection of each ECU; the original RTT sequence includes a set of HTTP requests, and the ECU number, timestamp, original response time and API endpoint number corresponding to each HTTP request; The set determination module 602 is used to take the set of differences between the original response time and the network round-trip time of each HTTP request as the ECU processing time set corresponding to the HTTP request set; The baseline determination module 603 is used to establish a normal baseline for the response time based on the ECU processing time of each ECU in the ECU processing time set and the ECU number, API endpoint number and timestamp corresponding to the ECU processing time of each ECU in the original RTT sequence; The score calculation module 604 is used to calculate the comprehensive abnormal score based on the normal baseline and the current ECU processing time, and to make fault judgments based on the comprehensive abnormal score to obtain the hidden fault detection results.
[0067] The SOVD concealed fault detection device 600 provided in the above embodiments can realize the technical solutions described in the above SOVD concealed fault detection method embodiments. The specific implementation principles of each module or unit can be found in the corresponding content in the above SOVD concealed fault detection method embodiments, and will not be repeated here.
[0068] The SOVD hidden fault detection method and device provided by the present invention have been described in detail above. Specific examples have been used to illustrate the principle and implementation of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core idea of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation and application scope based on the idea of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for detecting hidden faults in SOVD, characterized in that, Applied to SOVD clients; Receive the current ECU processing time within the current time window, and obtain the vehicle's original RTT sequence and the network round-trip time for periodic probing of each ECU; the original RTT sequence includes a set of HTTP requests, and the ECU number, timestamp, original response time, and API endpoint number corresponding to each HTTP request; The set of differences between the original response time and the network round-trip time of each HTTP request is taken as the ECU processing time set corresponding to the HTTP request set; A normal baseline for response time is established based on the ECU processing time in the ECU processing time set and the ECU number, API endpoint number, and timestamp corresponding to the ECU processing time in the original RTT sequence; A comprehensive anomaly score is calculated based on the normal baseline and the current ECU processing time. Fault judgment is then performed based on the comprehensive anomaly score to obtain the hidden fault detection result.
2. The SOVD concealed fault detection method according to claim 1, characterized in that, The original RTT sequence generation process is as follows: when sending an HTTP request to the corresponding ECU based on the diagnostic information, the request sending time is recorded; when receiving the HTTP response returned by the corresponding ECU, the complete response reception time is recorded; the original response time is determined based on the difference between the complete response reception time and the request sending time; and the original RTT sequence of the current HTTP request is saved.
3. The SOVD concealed fault detection method according to claim 1, characterized in that, The process of obtaining the network round-trip time involves periodically sending ICMP Echo requests to each ECU, and determining the network round-trip time upon receiving a response from the corresponding ECU; wherein the detection period for sending the ICMP Echo request to each ECU is 2 to 5 times the diagnostic request period.
4. The SOVD concealed fault detection method according to claim 1, characterized in that, The normal baseline includes an ECU-level baseline; establishing a normal baseline for response time based on the ECU processing time in the ECU processing time set and the ECU number, API endpoint number, and timestamp corresponding to the ECU processing time in the original RTT sequence includes: Based on the ECU number and the API endpoint number, the ECU processing time of all API endpoints under the same ECU is input into a Gaussian mixture model for multimodal distribution fitting, and the weight, mean, and variance of each mode are output; the modes in the multimodal distribution include light request peak, medium request peak, and heavy routine peak; An ECU-level baseline is determined based on the weights, mean, and variance of each mode.
5. The SOVD concealed fault detection method according to claim 4, characterized in that, The normal baseline includes an API endpoint-level baseline; the step of establishing a normal baseline for response time based on the ECU processing time in the ECU processing time set and the ECU number, API endpoint number, and timestamp corresponding to the ECU processing time in the original RTT sequence further includes: The ECU processing time for the same API endpoint is calculated based on the API endpoint number and kernel density estimation to obtain the baseline mean, standard deviation and fast / slow boundary of the corresponding API endpoint. The API endpoint-level baseline is determined based on the baseline mean, the standard deviation, and the fast / slow boundary for each API endpoint.
6. The SOVD concealed fault detection method according to claim 5, characterized in that, The normal baseline includes a time-dependent baseline; the step of establishing a normal baseline for response time based on the ECU processing time in the ECU processing time set and the ECU number, API endpoint number, and timestamp corresponding to the ECU processing time in the original RTT sequence further includes: After aligning the ECU processing times of different API endpoints within the same time window under the same ECU based on the ECU number, the API endpoint number, and the timestamp, the Pearson correlation coefficient between the ECU processing times is calculated to obtain the response time Pearson correlation coefficient matrix, and the response time Pearson correlation coefficient matrix is used as the time correlation baseline.
7. The SOVD concealed fault detection method according to claim 6, characterized in that, The calculation of the comprehensive anomaly score based on the normal baseline and the current ECU processing time includes: The current ECU processing time is simulated and calculated based on the ECU-level baseline to obtain the KL divergence; when the KL divergence is greater than a preset divergence threshold, the global degradation anomaly score of the ECU is calculated based on the KL divergence. For each API endpoint, calculate the deviation multiple between the current ECU processing time and the API endpoint-level baseline. When the deviation multiple is greater than a preset standard deviation, calculate the endpoint anomaly score based on the deviation multiple. Based on the current ECU processing time, calculate the correlation coefficient matrix between endpoints within the current window, and calculate the norm difference between the correlation coefficient matrix and the time correlation baseline; when the norm difference is greater than a preset difference threshold, calculate the resource competition anomaly score based on the norm difference; The weighted sum of the ECU global degradation anomaly score, the endpoint anomaly score, and the resource contention anomaly score is used to obtain the comprehensive anomaly score.
8. The SOVD concealed fault detection method according to claim 1, characterized in that, The step of determining the fault based on the comprehensive anomaly score to obtain the hidden fault detection result includes: An exponentially weighted moving average is applied to the comprehensive anomaly score to obtain a smoothed value sequence; Calculate the first derivative based on the smoothed value sequence; Fault judgment is performed based on all first derivatives to obtain the hidden fault detection results.
9. The SOVD concealed fault detection method according to claim 8, characterized in that, The process of determining faults based on all first derivatives to obtain hidden fault detection results includes: When the first derivative is continuously positive within a preset window, the fault type is determined to be a trend warning, triggering an active diagnosis process. The active diagnosis process involves controlling the SOVD client to increase the diagnostic sampling frequency of the corresponding ECU, and collecting and diagnosing data for subsequent time based on the diagnostic sampling frequency to determine a new comprehensive anomaly score. When the new comprehensive anomaly score of multiple consecutive windows is greater than the anomaly threshold, the hidden fault detection result is determined to be a hidden fault, and the API endpoint with the hidden fault is reported through the SOVD client.
10. A SOVD concealed fault detection device, characterized in that, include: The data acquisition module is used to receive the current ECU processing time within the current time window, and to acquire the vehicle's original RTT sequence and the network round-trip time for periodic detection of each ECU; the original RTT sequence includes a set of HTTP requests, and the ECU number, timestamp, original response time and API endpoint number corresponding to each HTTP request; The set determination module is used to take the set of differences between the original response time and the network round-trip time of each HTTP request as the ECU processing time set corresponding to the HTTP request set; The baseline determination module is used to establish a normal baseline for response time based on the ECU processing time of each ECU in the ECU processing time set and the ECU number, API endpoint number, and timestamp corresponding to the ECU processing time in the original RTT sequence; The score calculation module is used to calculate a comprehensive anomaly score based on the normal baseline and the current ECU processing time, and to make a fault judgment based on the comprehensive anomaly score to obtain the hidden fault detection result.