Software time synchronization method and system for multi-sensor data fusion

By extracting the multi-dimensional feature vectors of sensor nodes and using machine learning algorithms to identify abnormal nodes and dynamically adjust clock parameters, the time asynchrony problem caused by clock drift and network delay in distributed systems is solved, and high-precision time synchronization and coordination is achieved.

CN120611178APending Publication Date: 2025-09-09SHENZHEN YOUBIKANG TECH CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510706510.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies cannot effectively solve the time asynchrony problem caused by clock drift and network delay factors in distributed systems. In particular, under the influence of clock stability and network delay fluctuations of sensor networks in dynamic environments, synchronization accuracy and coordination are insufficient.

Method used

The principal component analysis method is used to extract the temperature, load and timing drift frequency feature vectors of sensor nodes. Confidence assessment models and multiple machine learning algorithms (such as k-means, isolation forest, Prophet, Hilbert cross-correlation phase detection, etc.) are combined to identify abnormal nodes and network bottlenecks, and dynamically adjust clock parameters to achieve high-precision synchronization.

Benefits of technology

It significantly improves the timestamp reliability and synchronization resource allocation efficiency of the sensor network, reduces the misjudgment caused by hardware status fluctuations, and improves the system's synchronization stability and task execution success rate in extreme environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611178A_ABST
    Figure CN120611178A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of software time synchronization, and discloses a software time synchronization method and system for multi-sensor data fusion, and the method comprises the steps: extracting temperature, load and drift frequency characteristics through principal component analysis based on the working state and historical drift data of a sensor, and constructing a confidence evaluation model to calculate the credibility of a timestamp; identifying an abnormal node group by using k-means and an isolated forest algorithm, analyzing a phase deviation fluctuation and network delay interaction effect, extracting a nonlinear drift feature in combination with a Prophet algorithm, and calculating a phase correlation value by using Hilbert cross-correlation; and dynamically adjusting node clock parameters and generating a calibration timestamp according to the network influence weight and the stability evaluation result. According to the method, the problem of time desynchrony caused by clock drift and network delay factors in a distributed system is effectively solved, and the overall time consistency and reliability of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software time synchronization, and in particular to a software time synchronization method and system for multi-sensor data fusion. Background Art

[0002] In rapidly developing fields such as the Internet of Things, industrial automation, and intelligent transportation, multi-sensor collaboration has become the core foundation for intelligent decision-making and precise control. With the expansion of distributed systems and the increasing heterogeneity of nodes, the demand for high-precision time synchronization is becoming increasingly urgent to ensure the accuracy of cross-node data fusion and the coordination of task execution. However, sensor networks in dynamic environments face complex challenges. On the one hand, the operating state of sensor nodes (such as temperature fluctuations and load changes) can significantly affect their clock stability, leading to frequent nonlinear drift. On the other hand, random fluctuations in network latency and differences in node hardware performance further exacerbate the risk of timestamp deviation.

[0003] In one existing technology, the implementation process of time synchronization is as follows: the system first selects a central reference clock node, and all sensor nodes periodically send time synchronization requests to this node. When the reference node receives the request, it immediately returns the current standard timestamp. The sensor node estimates the network transmission delay by measuring the round-trip time between the request and the response, and corrects the local clock based on a preset deviation calculation algorithm. For example, after sending the request, node A estimates the network delay through the round-trip time and directly overwrites the local clock value based on the timestamp returned by the reference node. To compensate for clock drift, the system collects historical deviation data, constructs a linear prediction model, and periodically applies uniform compensation parameters to all nodes. Anomaly detection uses a fixed threshold rule. If the network delay of a node exceeds the preset value or the clock deviation continues to exceed the range, it is marked as an anomaly and synchronization is suspended.

[0004] In practical applications, the synchronization architecture that relies on a single reference node is unable to perceive heterogeneous differences between nodes. For example, different hardware may have different response characteristics to temperature changes, and a unified linear model cannot describe the dynamic fluctuations of nonlinear drift. At the same time, static threshold rules cannot identify the correlation between network latency and hardware status (such as instantaneous load surges), resulting in short-term anomalies being mistakenly judged as persistent failures. In addition, the communication bottleneck of the reference node will directly affect the synchronization accuracy of the entire network. When the network is congested, the deviation estimation of all nodes will be disturbed. Therefore, this technology cannot effectively solve the time synchronization problem caused by clock drift and network latency in distributed systems. Summary of the Invention

[0005] The present invention provides a software time synchronization method and system for multi-sensor data fusion to solve the time asynchrony problem caused by clock drift and network delay factors in a distributed system.

[0006] In a first aspect, in order to solve the above technical problems, the present invention provides a software time synchronization method for multi-sensor data fusion, comprising:

[0007] Obtain working status data and historical clock drift records from sensor nodes, use principal component analysis to extract feature vectors containing temperature, load, and timing drift frequency, and obtain node status feature sets;

[0008] According to the node state feature set, a pre-established confidence evaluation model is used to calculate the confidence of each node timestamp to obtain a timestamp confidence distribution;

[0009] The timestamp confidence distribution is grouped using a k-means algorithm, and abnormal node groups are identified using an isolation forest algorithm to obtain abnormal node groups;

[0010] Analyze the phase fluctuation frequency of the phase offset value in the abnormal node group and the timing drift frequency of the timing data, and determine whether it is a high dynamic deviation node to obtain a stability assessment result;

[0011] Analyzing the interaction effect data of the phase offset value and the network delay in the abnormal node group, and calculating the delay distribution characteristics of the network. If the correlation strength between the interaction effect data and the delay distribution characteristics exceeds a preset threshold, the node is determined to be a network bottleneck affecting node, and a network influence weight set is obtained;

[0012] Extracting network delay data and hardware status data from the abnormal node group, and performing time deviation node judgment to obtain a time deviation node list;

[0013] Obtaining the clock drift data of the node from the deviation node list, extracting the nonlinear drift characteristics using the prophet algorithm, and obtaining the nonlinear drift parameters;

[0014] According to the nonlinear drift parameter, a Hilbert cross-correlation phase detection algorithm is used to calculate a phase correlation value between the deviation node and the reference clock data;

[0015] Obtaining the sensitivity of the confidence assessment model to timestamp distribution through the correlation between the phase correlation value and the hardware status data, and determining the parameter adjustment amplitude of the confidence assessment model using the timing distribution characteristics of the phase offset value to obtain the comprehensive linkage parameter;

[0016] According to the phase correlation value, comprehensive linkage parameter, stability assessment result and network impact weight set, the clock parameter value of the deviation node is adjusted, and if the linkage between the network delay data and the hardware status data exceeds expectations, a calibration timestamp set is generated.

[0017] In an optional embodiment, the calculating the confidence of each node timestamp using a pre-established confidence assessment model based on the node state feature set to obtain a timestamp confidence distribution includes:

[0018] Acquire feature data of a node timestamp from the node state feature set, and input the data into a pre-established confidence assessment model to obtain a confidence score;

[0019] Comparing the confidence score with a preset confidence threshold, and when the confidence score is lower than the confidence threshold, marking it as a confidence deviation node, to obtain a confidence deviation node set;

[0020] Extracting feature data of node timestamps in the confidence deviation node set, and analyzing timestamp confidence distribution using the OPTICS density clustering algorithm;

[0021] Wherein, the confidence assessment model is constructed by gradient boosting decision tree algorithm;

[0022] The characteristic data of the node timestamp include characteristic vectors of temperature, load and drift frequency.

[0023] In an optional embodiment, analyzing the phase fluctuation frequency of the phase offset value and the timing drift frequency of the timing data in the abnormal node group, and determining whether the node is a high dynamic deviation node to obtain a stability assessment result includes:

[0024] Acquire time series data of phase offset values ​​and clock drift from the abnormal node group, calculate the phase fluctuation frequency and timing drift frequency using Fourier transform, and obtain a frequency distribution result;

[0025] According to the frequency distribution result, a multi-band coherence analysis algorithm is used to calculate the coherence coefficient of the phase fluctuation frequency and the timing drift frequency. If the coherence coefficient is higher than a preset coherence threshold, it is determined to be a high dynamic state and a high dynamic flag is obtained;

[0026] Extracting deviation features from abnormal nodes based on the high dynamic identification, and using a K-means clustering algorithm to determine a set of high dynamic deviation nodes;

[0027] The stability evaluation value is calculated by using the mean square error through the high dynamic deviation node set and the time series data, and is used as the stability evaluation result.

[0028] In an optional embodiment, the analysis of the interaction effect data of the phase offset value and the network delay in the abnormal node group and the calculation of the network delay distribution characteristics are performed. If the correlation strength between the interaction effect data and the delay distribution characteristics exceeds a preset threshold, the node is determined to be a network bottleneck affecting node, and a network influence weight set is obtained, including:

[0029] Acquire network delay data and phase offset values ​​from the abnormal node group, and calculate the interaction effect of the network delay data and the phase offset value based on a sliding window cross-correlation analysis method to obtain interaction effect data;

[0030] According to the interaction effect data, the skewness and kurtosis of the network delay distribution are calculated using the moment estimation method to obtain the delay distribution characteristics;

[0031] Calculate the typical correlation coefficient between the delay distribution characteristics and the interaction effect data using a typical correlation analysis method. When the typical correlation coefficient exceeds a preset typical correlation threshold, use the DBSCAN algorithm to classify the abnormal node group and identify the network bottleneck affecting node set.

[0032] According to the network bottleneck influencing node set and the interaction effect data, a weighted average algorithm is adopted to generate a network influence weight set to obtain a network influence weight set.

[0033] In an optional embodiment, extracting network delay data and hardware status data from the abnormal node group, and performing time deviation node judgment to obtain a time deviation node list includes:

[0034] When the network delay data exceeds the preset delay threshold, it is marked as a delay abnormal node;

[0035] When the hardware status data deviates from the preset normal range, it is marked as an abnormal status node;

[0036] A fast set operation algorithm based on a bitmap is used to compare the intersection of the delay abnormal node and the state abnormal node, determine that the node in the intersection is a time deviation node, and generate a time deviation node list.

[0037] In an optional embodiment, the sensitivity of the confidence assessment model to the timestamp distribution is obtained by the correlation between the phase correlation value and the hardware status data, and the parameter adjustment amplitude of the confidence assessment model is determined by using the timing distribution characteristics of the phase offset value to obtain the comprehensive linkage parameter, including:

[0038] Using a moving average method to extract the time series distribution characteristics of the phase correlation value;

[0039] Obtaining hardware status data of the time-deviation node, analyzing the correlation between the phase correlation value and the node status using the Pearson correlation coefficient in combination with the timing distribution characteristics, and using the correlation as the sensitivity of the confidence assessment model to the timestamp distribution;

[0040] When the sensitivity exceeds a preset sensitivity threshold, a linear regression algorithm is used to adjust the parameters of the confidence assessment model, an adjustment range is obtained, and an optimized model parameter set is obtained;

[0041] The optimized model parameter set and the hardware status data are synthesized by a principal component weighted synthesis method to generate a comprehensive linkage parameter set.

[0042] In an optional embodiment, adjusting the clock parameter value of the deviation node according to the phase correlation value, the comprehensive linkage parameter, the stability assessment result, and the network impact weight set, and generating a calibration timestamp set if the linkage between the network delay data and the hardware status data exceeds expectations, includes:

[0043] Calculating an initial clock compensation amount using an incremental PID algorithm according to the phase correlation value and the network influence weight set;

[0044] Verifying feature linkage anomalies using a chi-square test based on the comprehensive linkage parameter set; and when feature linkage anomalies are present, performing a secondary correction on the initial clock compensation amount using an L2 regularized logistic regression model based on the stability assessment result to obtain a final clock compensation amount.

[0045] When there is no characteristic linkage anomaly, the initial clock compensation amount is used as the final clock compensation amount;

[0046] A final calibration timestamp set is generated according to the final clock compensation amount.

[0047] In a second aspect, the present invention provides a software time synchronization system for multi-sensor data fusion, comprising:

[0048] The data acquisition module is used to obtain working status data and historical clock drift records from sensor nodes, and uses principal component analysis to extract feature vectors containing temperature, load and timing drift frequency to obtain the node status feature set;

[0049] A confidence analysis module is used to calculate the confidence of each node timestamp based on the node state feature set and adopt a pre-established confidence evaluation model to obtain a timestamp confidence distribution;

[0050] An abnormal node analysis module is used to group the timestamp confidence distribution using a k-means algorithm and identify abnormal node groups using an isolation forest algorithm to obtain abnormal node groups;

[0051] A stability evaluation module is used to analyze the phase fluctuation frequency of the phase offset value in the abnormal node group and the timing drift frequency of the timing data, and to determine whether the node is a high dynamic deviation node to obtain a stability evaluation result;

[0052] a network impact analysis module configured to analyze the interaction effect data between the phase offset values ​​and the network delay in the abnormal node group and calculate the delay distribution characteristics of the network. If the correlation strength between the interaction effect data and the delay distribution characteristics exceeds a preset threshold, the node is determined to be a network bottleneck node and a network impact weight set is obtained.

[0053] Deviation node analysis module, used to extract network delay data and hardware status data from the abnormal node group, and perform time deviation node judgment to obtain a time deviation node list;

[0054] A clock drift analysis module is used to obtain the clock drift data of the node from the deviation node list, extract the nonlinear drift characteristics using the prophet algorithm, and obtain the nonlinear drift parameters;

[0055] A phase correlation calculation module, configured to calculate a phase correlation value between the deviation node and the reference clock data using a Hilbert cross-correlation phase detection algorithm according to the nonlinear drift parameter;

[0056] A linkage parameter analysis module is used to obtain the sensitivity of the confidence assessment model to the timestamp distribution through the correlation between the phase correlation value and the hardware status data, and to determine the parameter adjustment amplitude of the confidence assessment model using the timing distribution characteristics of the phase offset value to obtain the comprehensive linkage parameter;

[0057] A time calibration module is used to adjust the clock parameter value of the deviation node according to the phase correlation value, the comprehensive linkage parameter, the stability assessment result and the network impact weight set, and generate a calibration timestamp set if the linkage between the network delay data and the hardware status data exceeds expectations.

[0058] In a third aspect, the present invention also provides an electronic device comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements a software time synchronization method for multi-sensor data fusion as described above.

[0059] In a fourth aspect, the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the software time synchronization methods for multi-sensor data fusion described above.

[0060] Compared with the prior art, the present invention has the following beneficial effects:

[0061] (1) Principal component analysis extracts multidimensional features such as temperature, load, and timing drift frequency, and combines this with a gradient boosting decision tree to construct a dynamic confidence assessment model, which can quantify the reliability of node timestamps in real time. Compared with traditional static threshold judgment, this method dynamically adjusts weights based on environmental context, significantly reducing misjudgments caused by hardware status fluctuations and significantly improving the accuracy and adaptability of confidence assessment.

[0062] (2) A hierarchical analysis strategy combining K-means clustering and the isolation forest algorithm is used to accurately locate the spatiotemporal patterns of abnormal nodes. Through standardized confidence distribution, distance measurement, and abnormal path segmentation, high-dynamic deviation nodes and network bottleneck nodes are effectively identified, significantly improving the efficiency of synchronous resource allocation and the overall fault tolerance of the system.

[0063] (3) The Prophet algorithm extracts the nonlinear characteristics of clock drift and combines it with the Hilbert cross-correlation phase detection algorithm to accurately capture high-frequency dynamic offsets. A support vector machine is used to quantify drift trends and generate compensation parameters, significantly reducing synchronization errors and improving long-term operational stability in extreme environments or load mutation scenarios.

[0064] This invention achieves a closed-loop time synchronization optimization framework through four major technological breakthroughs: multi-dimensional feature fusion, hierarchical management of abnormal nodes, nonlinear drift compensation, and dynamic linkage calibration. In complex scenarios such as the Industrial Internet of Things and intelligent transportation, system time consistency errors are precisely controlled to a very small range, significantly improving the success rate of node collaborative task execution. This effectively addresses the synchronization misalignment issues caused by traditional methods due to static modeling, a single reference source, and linear compensation, providing core technical support for the highly reliable operation of distributed systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 This is a flowchart of a software time synchronization method for multi-sensor data fusion provided by the first embodiment of the present invention;

[0066] Figure 2 This is a schematic diagram of the structure of a software time synchronization system for multi-sensor data fusion provided by the second embodiment of the present invention. DETAILED DESCRIPTION

[0067] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0068] Reference Figure 1 The first embodiment of the present invention provides a software time synchronization method for multi-sensor data fusion, comprising the following steps:

[0069] S11, obtains working status data and historical clock drift records from sensor nodes, uses principal component analysis to extract feature vectors containing temperature, load and timing drift frequency, and obtains node status feature set;

[0070] S12, based on the node state feature set, using a pre-established confidence evaluation model, calculating the confidence of each node timestamp to obtain a timestamp confidence distribution;

[0071] S13, grouping the timestamp confidence distribution using a k-means algorithm, and identifying abnormal node groups using an isolation forest algorithm to obtain abnormal node groups;

[0072] S14, analyzing the phase fluctuation frequency of the phase offset value and the timing drift frequency of the timing data in the abnormal node group, and determining whether the node is a high dynamic deviation node to obtain a stability assessment result;

[0073] S15, analyzing the interaction effect data of the phase offset values ​​and the network delay in the abnormal node group, and calculating the delay distribution characteristics of the network. If the correlation strength between the interaction effect data and the delay distribution characteristics exceeds a preset threshold, the node is determined to be a network bottleneck affecting node, and a network influence weight set is obtained;

[0074] S16, extracting network delay data and hardware status data from the abnormal node group, and performing time deviation node judgment to obtain a time deviation node list;

[0075] S17, obtaining clock drift data of the node from the deviation node list, extracting nonlinear drift features using a prophet algorithm, and obtaining nonlinear drift parameters;

[0076] S18, calculating a phase correlation value between the deviation node and the reference clock data using a Hilbert cross-correlation phase detection algorithm according to the nonlinear drift parameter;

[0077] S19, obtaining the sensitivity of the confidence assessment model to the timestamp distribution through the correlation between the phase correlation value and the hardware status data, and determining the parameter adjustment amplitude of the confidence assessment model using the time series distribution characteristics of the phase offset value to obtain the comprehensive linkage parameter;

[0078] S20, adjusting the clock parameter value of the deviation node according to the phase correlation value, comprehensive linkage parameter, stability assessment result and network impact weight set, and generating a calibration timestamp set if the linkage between the network delay data and the hardware status data exceeds expectations.

[0079] In step S11, the working status data and historical clock drift records are obtained from the sensor nodes, and the principal component analysis method is used to extract the feature vector containing temperature, load and timing drift frequency to obtain the node status feature set.

[0080] Specifically, operating status data, including temperature, load, and historical clock drift records, is periodically collected from sensor nodes. Temperature is monitored in real time using built-in sensors, load is calculated by computing resource usage, and clock drift records contain timestamps and deviation values. Data collected at different frequencies is unified into a structured format (such as JSON) using a predefined protocol (such as MQTT) and stored in a distributed database (such as HBase) to ensure timestamp alignment. Subsequently, principal component analysis (PCA) is applied to the standardized dataset, using temperature, load, and clock timing drift frequency as input features. Principal components are extracted through covariance matrix decomposition, retaining those dimensions that explain a sufficient proportion of variance. Based on predefined weight assignment rules (e.g., a weight of 0.4 for temperature, 0.35 for load, and 0.25 for timing drift frequency), the principal components are weighted and fused with the original features to generate a weighted feature vector. If the weight deviation exceeds a threshold (e.g., 0.1), the weight ratio is dynamically adjusted. Finally, a clustering algorithm (such as K-means) is used to remove redundant data and identify the core feature set representing the node status. This step effectively compresses the data size and retains key information through dimensionality reduction and feature fusion, providing input features with high signal-to-noise ratio for subsequent timestamp confidence assessment.

[0081] In step S12, based on the node state feature set, a pre-established confidence evaluation model is used to calculate the confidence of each node timestamp to obtain a timestamp confidence distribution.

[0082] In a specific embodiment, the confidence of each node timestamp is calculated based on the node state feature set using a pre-established confidence evaluation model to obtain a timestamp confidence distribution, including:

[0083] Acquire feature data of a node timestamp from the node state feature set, and input the data into a pre-established confidence assessment model to obtain a confidence score;

[0084] Comparing the confidence score with a preset confidence threshold, and when the confidence score is lower than the confidence threshold, marking it as a confidence deviation node, to obtain a confidence deviation node set;

[0085] Extracting feature data of node timestamps in the confidence deviation node set, and analyzing timestamp confidence distribution using the OPTICS density clustering algorithm;

[0086] Wherein, the confidence assessment model is constructed by gradient boosting decision tree algorithm;

[0087] The characteristic data of the node timestamp include characteristic vectors of temperature, load and drift frequency.

[0088] Specifically, feature data corresponding to each node's timestamp is extracted from the node status feature set. This feature data consists of a feature vector composed of three dimensions: temperature, load, and timing drift frequency. This feature vector is then fed into a pre-trained confidence assessment model. This model, based on an ensemble learning method using gradient boosting decision trees, iterates through multiple rounds of decision tree construction and corrects prediction residuals, ultimately outputting a confidence score for each timestamp.

[0089] The confidence assessment model training process first requires constructing a multidimensional training dataset containing temperature, load, and timing drift frequency. Specifically, the temperature sampling values ​​of each sensor node (e.g., degrees Celsius data collected per second), real-time load indicators (CPU occupancy, memory usage), and timing drift frequency calculated via a sliding window Fourier transform (reflecting the periodic fluctuation characteristics of clock deviation per unit time) are extracted from historical operation logs. After aligning these three types of raw data by timestamp, principal component analysis is used to reduce the feature dimensionality, removing redundant information while retaining principal components that explain more than 85% of the variance. This results in a standardized feature vector that serves as the model input.

[0090] Training adopts a phased, progressive construction strategy. Initially, the base learner is set as a shallow decision tree with 10 leaf nodes. Huber loss is used as the loss function to enhance robustness against outliers. In the first iteration, the model calculates split points on the input feature vector, assesses the importance of each feature dimension using the Gini coefficient, and prioritizes the time series drift frequency as the root node split feature. Each decision tree uses recursive partitioning to divide the training sample into 20-30 feature space subsets, generating preliminary confidence predictions.

[0091] Starting from the second round of iteration, a residual correction mechanism is introduced. The difference between the first-round prediction result and the true confidence label (a 0-1 standardized score annotated by experts) is calculated as the residual vector. To address this residual, a new decision tree is created by limiting the maximum depth to 5 layers to prevent overfitting. The weighted quantile algorithm of the XGBoost framework is used to quickly locate the optimal split point, focusing on correcting the prediction deviation of the previous model under high temperature (>60°C) and high load (CPU>80%) conditions. The output of each new tree is scaled by a learning rate of 0.1, and the overall prediction value is updated in a cumulative manner. After 200 rounds of iteration, the early stopping mechanism is triggered when the mean absolute error of the validation set decreases by less than 0.001 for five consecutive rounds.

[0092] During the model optimization phase, feature interaction detection was introduced. Shapley value analysis revealed that the product of temperature and load has a nonlinear effect on the confidence score. Therefore, a temperature-load coupling factor was added to the feature engineering process. The final model includes 180 decision trees with a depth of 4-6 layers. It achieved a Pearson correlation coefficient of 0.92 on the test set and maintained a detection response latency of less than 200ms for sudden clock drift (such as a sudden temperature jump of 10°C).

[0093] If the confidence score is lower than the preset confidence threshold, the system marks the corresponding node as a confidence deviation node and aggregates it into a confidence deviation node set. Subsequently, the confidence scores and associated feature vectors of all timestamps are extracted from the confidence deviation node set, and their distribution characteristics are analyzed using the OPTICS density clustering algorithm.

[0094] In the implementation of the OPTICS density clustering algorithm, the calculation of core distance and reachability distance relies on the node's timestamp confidence score and its associated multidimensional feature vector dataset. Specifically, the system extracts the normalized feature vector corresponding to each timestamp from the set of deviation nodes (including the principal component features of temperature, load, and time series drift frequency) as input spatial data, and presets the neighborhood radius ε = 0.3 (the Euclidean space unit after feature normalization) and the minimum number of neighborhood points MinPts = 8 density parameters.

[0095] For each data point p, the algorithm first uses a KD tree to quickly retrieve all sample points within its ε neighborhood. When the neighborhood contains at least MinPts points, p is marked as a core point. The core distance, coreDist(p), is then defined as the actual distance from the point to its MinPtsth nearest neighbor. For example, in a three-dimensional feature space, if the Euclidean distance between p's eighth nearest neighbor, q, and p is 0.25, then coreDist(p) = 0.25. For non-core points, the core distance is marked as undefined.

[0096] The calculation of the reachable distance reachDist(p, o) occurs during the orderly expansion of the core point o. When processing the neighborhood of the core point o, for each point p (including o itself) in its ε neighborhood, the actual distance distance(p, o) from p to o is calculated. At this time, the reachable distance takes the larger value of coreDist(o) and distance(p, o). For example, when coreDist(o) = 0.3 and p is 0.15 away from o, reachDist(p, o) = 0.3; if p is 0.35 away from o, then reachDist(p, o) = 0.35. This calculation method ensures that within the density area of ​​the core point o, the reachable distance of all points is at least equal to the minimum density requirement of the area (i.e., coreDist(o)).

[0097] Through iterative processing, the algorithm generates a sorted list of reachable distances for all core points, maintaining a priority queue to dynamically update the minimum reachable distance. The resulting reachability map intuitively displays the density structure of the confidence distribution—cluster boundaries are identified where the reachable distance suddenly increases, while consistently low reachable distance intervals correspond to high-density clusters, and outliers exceeding the core distance threshold are identified as isolated nodes. This hierarchical analysis based on density reachability effectively distinguishes between persistent low-confidence clusters caused by hardware aging (reachable distance fluctuations less than 0.1) and transient outliers caused by sudden network interference (reachable distance jumps exceeding 0.4).

[0098] This step dynamically quantifies the reliability of timestamps, screens potential abnormal nodes, and provides a data basis for subsequent abnormal group identification and calibration, thereby improving the system's detection accuracy and response efficiency for clock deviations.

[0099] In step S13, the timestamp confidence distribution is grouped using the k-means algorithm, and the abnormal node group is identified using the isolation forest algorithm to obtain the abnormal node group.

[0100] Specifically, the confidence score of each node and its corresponding temperature, load, and drift frequency feature vectors are extracted from the set of deviation nodes generated in step S12 as input data. First, the K-means clustering algorithm is used to group the timestamp confidence distribution. This algorithm uses Euclidean distance to measure the similarity between nodes. Through iterative optimization, the data is divided into a preset number of clusters, ensuring that the confidence scores and feature vectors of nodes within a cluster are as similar as possible and that there is significant difference between clusters. After clustering is completed, each node is assigned to a specific cluster to form a preliminary node grouping. The preset number of clusters is dynamically calculated through the following process: the system first performs principal component analysis on the standardized multi-dimensional feature data (the principal components of temperature, load, and timing drift frequency). When the cumulative variance contribution rate is ≥85%, the effective number of dimensions is determined to be 3; then the candidate clusters from k=2 to k=5 are traversed, and the average silhouette coefficient of each k value is calculated. When k=3, the silhouette coefficient reaches a peak of 0.68 (an increase of 15% compared to k=2, and a decrease to 0.63 when k=4). Combined with the 3-dimensional spatial characteristics of the principal component, the system automatically locks the optimal number of clusters to 3.

[0101] The isolation forest algorithm is then applied to the clustered node groups to identify anomalous node clusters. The isolation forest algorithm constructs multiple binary isolation trees by randomly selecting feature dimensions and recursively partitioning the data space. Because anomalous nodes are sparsely distributed, they typically require fewer splits to be isolated to leaf nodes. The algorithm calculates the average path length of all nodes in the isolation tree and generates an anomaly score. If the anomaly score exceeds a preset threshold, the node is considered an anomaly. Finally, the nodes that meet the conditions are aggregated into an abnormal node group, where the preset score threshold is dynamically calibrated through the following process: the system loads the verification set (including 200 normal nodes and 50 known abnormal nodes clearly marked in historical data) in the initialization phase, first runs the isolation forest algorithm on the normal nodes to generate an abnormal score distribution, and calculates its mean μ = 0.25 and standard deviation σ = 0.08; then, according to the 3σ principle, the initial threshold is set to μ + 3σ = 0.49, and then the threshold is verified with known abnormal nodes (for example, 45 / 50 abnormal nodes are detected exceeding the threshold), and finally the optimal threshold is determined to be 0.52 through the ROC curve (at this time TPR = 92%, FPR = 5%), and is recalibrated based on the new data distribution every 24 hours during operation.

[0102] This step clusters potential deviation patterns and combines them with anomaly detection to accurately locate outlier nodes in the group, providing a highly reliable anomaly dataset for subsequent stability assessment and network bottleneck analysis, thereby enhancing the system's ability to handle clock synchronization issues in complex dynamic environments.

[0103] In step S14, the phase fluctuation frequency of the phase offset value and the timing drift frequency of the timing data in the abnormal node group are analyzed, and a high dynamic deviation node is judged to obtain a stability evaluation result.

[0104] In a specific embodiment, the analyzing the phase fluctuation frequency of the phase offset value and the timing drift frequency of the timing data in the abnormal node group, and determining whether the node is a high dynamic deviation node, and obtaining a stability assessment result, includes:

[0105] Acquire time series data of phase offset values ​​and clock drift from the abnormal node group, calculate the phase fluctuation frequency and timing drift frequency using Fourier transform, and obtain a frequency distribution result;

[0106] According to the frequency distribution result, a multi-band coherence analysis algorithm is used to calculate the coherence coefficient of the phase fluctuation frequency and the timing drift frequency. If the coherence coefficient is higher than a preset coherence threshold, it is determined to be a high dynamic state and a high dynamic flag is obtained;

[0107] Extracting deviation features from abnormal nodes based on the high dynamic identification, and using a K-means clustering algorithm to determine a set of high dynamic deviation nodes;

[0108] The stability evaluation value is calculated by using the mean square error through the high dynamic deviation node set and the time series data, and is used as the stability evaluation result.

[0109] Specifically, the phase offset and clock drift time series data of each node are extracted from the abnormal node group as input. The phase offset value contains the phase deviation record of the node at different time points, while the time series data records the cumulative deviation between the node clock and the reference clock. First, a Fourier transform is performed on both types of time series data to convert the time domain signals into frequency domain signals. The phase fluctuation frequency and timing drift frequency are extracted to form the phase offset frequency distribution and timing drift frequency distribution, and the resulting frequency distribution is obtained.

[0110] Based on the frequency distribution results, a multi-band coherence analysis algorithm is used to divide the frequency into several sub-bands. The coherence coefficient between the phase offset frequency and the timing drift frequency is calculated within each sub-band. If the average coherence coefficient across all sub-bands exceeds a preset coherence threshold, the node is determined to be in a high-dynamic state and a high-dynamic flag is generated. The preset coherence threshold is dynamically generated by taking the upper limit of the 95% confidence interval based on the coherence coefficient distribution of normal nodes in the same frequency band in historical data. For example, if the average coherence coefficient of normal nodes in the 0.5-0.8Hz frequency band is 0.72 with a standard deviation of 0.08, the threshold is automatically set to 0.72 + 1.96 × 0.08 ≈ 0.88. Exceeding this value triggers a high-dynamic determination.

[0111] When forming a set of high dynamic deviation nodes, the system first extracts the time series data of the phase offset from the nodes marked as high dynamic state (for example, 1800 phase offset values ​​collected at 0.1 second intervals for each node in the last 30 minutes), and calculates three key statistics based on the time series data: the mean of the phase offset (arithmetic mean, reflecting the overall deviation level), the variance of the phase offset (using an unbiased estimation formula to calculate the fluctuation intensity), and the maximum fluctuation amplitude (the absolute maximum value of the difference between adjacent sampling points). These three statistics are used to form a three-dimensional feature vector (such as [2.3ms, 0.8ms 2 , 1.5ms]), the system dynamically groups nodes using the K-means clustering algorithm. The specific process is as follows: In the initialization phase, the k-means++ algorithm is used to select three initial cluster centers (corresponding to high, medium, and low dynamic levels), for example, [3.0, 1.2, 2.0], [1.5, 0.5, 0.8], and [0.8, 0.1, 0.3]. During the iterative optimization process, the Euclidean distance from each node's feature vector to each cluster center is calculated, and nodes are assigned to the closest cluster. The cluster center is then updated to the mean of the node features within the cluster. After 10 iterations, the cluster center is terminated when the change in cluster center is less than 1e-4. Finally, the cluster with the largest feature mean (for example, the cluster with the center [2.9, 1.1, 1.9], containing 15% of the nodes) is selected as the set of nodes with high dynamic deviation. For the nodes in this set, the system extracts their original phase offset time series data (for example, 1800 samples of node A) and calculates the stability assessment value using the root mean square error formula.

[0112] This step uses frequency domain analysis and feature clustering to accurately identify unstable nodes in highly dynamic environments and quantify their clock fluctuation characteristics, providing key input for subsequent network impact weight calculation and dynamic calibration, thereby enhancing the system's adaptability to complex scenarios.

[0113] In step S15, the interaction effect data of the phase offset value and the network delay in the abnormal node group is analyzed, and the delay distribution characteristics of the network are calculated. If the correlation strength between the interaction effect data and the delay distribution characteristics exceeds a preset threshold, it is determined to be a network bottleneck affecting node, and a network impact weight set is obtained.

[0114] In a specific embodiment, the analysis of the interaction effect data of the phase offset value and the network delay in the abnormal node group and the calculation of the network delay distribution characteristics are performed. If the correlation strength between the interaction effect data and the delay distribution characteristics exceeds a preset threshold, the node is determined to be a network bottleneck affecting node, and a network influence weight set is obtained, including:

[0115] Acquire network delay data and phase offset values ​​from the abnormal node group, and calculate the interaction effect of the network delay data and the phase offset value based on a sliding window cross-correlation analysis method to obtain interaction effect data;

[0116] According to the interaction effect data, the skewness and kurtosis of the network delay distribution are calculated using the moment estimation method to obtain the delay distribution characteristics;

[0117] Calculate the typical correlation coefficient between the delay distribution characteristics and the interaction effect data using a typical correlation analysis method. When the typical correlation coefficient exceeds a preset typical correlation threshold, use the DBSCAN algorithm to classify the abnormal node group and identify the network bottleneck affecting node set.

[0118] According to the network bottleneck influencing node set and the interaction effect data, a weighted average algorithm is adopted to generate a network influence weight set to obtain a network influence weight set.

[0119] Specifically, in step S15, network delay data and phase offset values ​​for each node in the abnormal node cluster are extracted as input data. The network delay data records the communication delay between the node and the reference node, while the phase offset value represents the phase deviation between the node's local clock and the reference clock. First, a sliding window cross-correlation analysis method is used to traverse the time series data with a fixed time window length. The cross-correlation coefficient between the network delay and phase offset values ​​within each window is calculated to generate an interaction effect data series. This data series reflects the dynamic correlation between the two over different time periods.

[0120] When analyzing the correlation between network delay distribution characteristics and interaction effect data, the system first extracts network delay time series data for each node from the abnormal node cluster (for example, delay values ​​collected every minute for 10 hours, resulting in 600 data points) and phase offset interaction effect series (dynamic correlations between delay and phase offset calculated using a sliding window, such as correlation coefficients calculated within 30-second windows). The characteristics of the delay distribution are analyzed using statistical methods: skewness quantifies the asymmetry of the distribution by calculating the cubic mean of the data. For example, positive skewness results when the delay data is concentrated to the right of the mean (e.g., most delay values ​​are above average). Kurtosis measures the sharpness of the distribution by calculating the fourth mean. Kurtosis means that delay values ​​are concentrated near the mean and exhibit extreme fluctuations (e.g., occasional delay spikes of hundreds of milliseconds). These two metrics together form a feature vector describing the network delay pattern (e.g., a node has a skewness of 1.8 and a kurtosis of 5.2).

[0121] Subsequently, the system uses the canonical correlation analysis method to explore the correlation between the delay distribution characteristics and the interaction effect. This method constructs two groups of variables (the skewness and kurtosis of the delay as one group, and the mean and variance of the correlation coefficient of the interaction effect as another group) to find the linear combination that best explains the correlation between the two. For example, if the right-skewed characteristic of the delay distribution (skewness > 1.5) and the high volatility of the interaction effect (correlation coefficient variance > 0.2) show a strong correlation (canonical correlation coefficient > 0.75), it is determined that the network delay has a significant dynamic impact on the phase offset. This threshold is not a fixed value, but is dynamically generated based on the statistical distribution of normal interaction scenarios in historical data (such as taking the top 1% of the correlation coefficient values ​​as the critical point) to ensure that it can adapt to different network environments.

[0122] Finally, the DBSCAN density clustering algorithm is applied to the node groups that meet the association conditions. This algorithm automatically identifies high-density areas based on the distribution density of nodes in the four-dimensional feature space (skewness, kurtosis, correlation coefficient mean, and variance). For example, if the delay skewness of a group of nodes is concentrated in the range of 1.7-2.3, the kurtosis is 5.0-6.5, and the correlation coefficient variance is continuously higher than 0.15, the algorithm will classify them into the same cluster. The common characteristics of these nodes are: the delay distribution is severely right-skewed during periods of network congestion (such as during hourly data synchronization), and the dynamic correlation with the phase offset fluctuates violently. Such nodes are marked as network bottleneck-affected node sets, and their clock parameters will be adjusted first in subsequent calibrations, such as targeted increases in compensation or reductions in synchronization frequency, thereby effectively alleviating the global time inaccuracy problem caused by network fluctuations.

[0123] Finally, based on the interaction effect data of the network bottleneck node set, a weighted average algorithm is used to calculate the influence weight of each node. The weight allocation is based on the correlation coefficient strength of the nodes in the interaction effect data and the density level in the clustering results to generate a network influence weight set.

[0124] This step uses dynamic correlation analysis and density clustering to accurately locate the key nodes that cause network congestion and quantify their impact on global synchronization performance. This provides a basis for subsequent dynamic adjustment of clock parameters, thereby optimizing the system's time synchronization stability under high load or network fluctuation scenarios.

[0125] In step S16, network delay data and hardware status data are extracted from the abnormal node group, and time deviation node judgment is performed to obtain a time deviation node list.

[0126] In a specific embodiment, extracting network delay data and hardware status data from the abnormal node group, and performing time deviation node judgment to obtain a time deviation node list includes:

[0127] When the network delay data exceeds the preset delay threshold, it is marked as a delay abnormal node;

[0128] When the hardware status data deviates from the preset normal range, it is marked as an abnormal status node;

[0129] A fast set operation algorithm based on a bitmap is used to compare the intersection of the delay abnormal node and the state abnormal node, determine that the node in the intersection is a time deviation node, and generate a time deviation node list.

[0130] Specifically, the system extracts network latency data and hardware status data from each node in the abnormal node cluster as input. Network latency data records the communication delay between the node and the reference node, while hardware status data includes parameters such as CPU utilization, memory usage, and temperature. The system also sets preset latency thresholds and normal hardware status ranges. For example, the latency threshold is set to the upper limit of communication latency, and the normal CPU utilization range is set to no more than a set percentage.

[0131] The system traverses all nodes in the abnormal node group and checks each node's network latency data to see if it exceeds the preset latency threshold. If so, the node is marked as having abnormal latency. The system also checks the hardware status data to see if CPU usage, memory utilization, and temperature are outside the normal range. If any of these parameters exceed the standard, the node is marked as abnormal.

[0132] Subsequently, a fast bitmap-based set operation algorithm was used to process the two sets of abnormal nodes. Delay abnormal nodes and status abnormal nodes were mapped into binary bitmaps, with each bit corresponding to the node's status (1 for abnormal and 0 for normal). A bitwise AND operation on the bitmaps was performed to filter out nodes that existed in both sets—that is, nodes with both abnormal delays and hardware statuses—to generate a list of time deviation nodes.

[0133] This step uses dual anomaly detection and efficient set operations to accurately locate time deviation nodes caused by network delays and hardware loads, avoiding misjudgment due to a single factor and providing a highly reliable target node set for subsequent clock calibration, thereby improving the system's judgment efficiency and processing accuracy for complex anomaly scenarios.

[0134] In step S17, the clock drift data of the node is obtained from the deviation node list, and the prophet algorithm is used to extract the nonlinear drift characteristics to obtain the nonlinear drift parameters.

[0135] Specifically, clock drift data for each node is extracted from the deviation node list. This data consists of a sequence of node timestamps and their corresponding clock deviation values, reflecting the cumulative deviation between the node's local clock and the reference clock. Clock drift data is stored in a historical record, for example, recording the deviation value every 5 seconds, forming a continuous time series.

[0136] The Prophet algorithm is used to model clock drift data. This algorithm decomposes the time series into trend, seasonal, and residual terms to identify nonlinear drift characteristics. First, the algorithm automatically detects change points in the time series and divides the overall trend into multiple piecewise linear or nonlinear intervals, capturing the dynamics of the drift rate. Second, for periodic fluctuations (such as the daily drift period caused by temperature changes), the algorithm fits the seasonal component using a Fourier series. Finally, the trend and seasonal terms are superimposed to generate a complete fitted model.

[0137] Trend and seasonal parameters are extracted from the model as nonlinear drift features. Trend parameters describe the direction and magnitude of long-term changes in clock bias, such as exponential growth of bias with runtime. Seasonal parameters characterize periodic fluctuations, such as the amplification of periodic bias caused by daily workload peaks. Furthermore, the change points in the algorithm's output mark the moment of a sudden trend change, identifying hardware status changes or environmental disturbances.

[0138] This step analyzes the complex pattern of clock drift, quantifies the nonlinear dynamic characteristics, and provides accurate drift prediction parameters for subsequent phase offset compensation, thereby enhancing the system's ability to correct nonlinear clock deviations.

[0139] In step S18, a Hilbert cross-correlation phase detection algorithm is used according to the nonlinear drift parameter to calculate a phase correlation value between the deviation node and the reference clock data.

[0140] Specifically, the node's trend parameter, seasonal parameter, and change point location are extracted from the nonlinear drift parameters generated in step S17, while the reference clock's timing phase data is simultaneously acquired as input. The trend parameter describes the long-term trend of the node's clock deviation, such as the rate at which the deviation increases or decreases over time. The seasonal parameter characterizes periodic fluctuations, such as the periodic phase shift caused by temperature fluctuations at fixed times of the day. The change point location marks the time when the trend suddenly changes, such as when hardware failure or environmental interference causes a sudden change in clock deviation. The reference clock data consists of a sequence of standard timestamps and their corresponding ideal phase values.

[0141] Using the Hilbert cross-correlation phase detection algorithm, the node's clock deviation timing data and reference clock phase data are first Hilbert transformed, converting the original signals into analytical signals containing instantaneous phase information. For example, the phase fluctuation signal of a certain period in the node deviation data is transformed to obtain its real and imaginary components, and then the instantaneous phase angle is extracted. Subsequently, the cross-correlation function of the two analytical signals is calculated, and the correlation coefficients at different time offsets are iterated. The offset that maximizes the correlation coefficient is found as the phase correlation value. This value reflects the degree of phase deviation between the node clock and the reference clock. For example, the phase correlation value of a node during high-temperature periods is significantly lower than that during normal-temperature periods, indicating that its phase synchronization is significantly affected by temperature.

[0142] This step quantifies the dynamic phase correlation between the node and the reference clock, providing accurate phase deviation measurement for the subsequent dynamic compensation algorithm, thereby achieving high-precision correction of nonlinear drift and ensuring the stability of multi-node time synchronization in complex environments.

[0143] In step S19, the sensitivity of the confidence assessment model to the timestamp distribution is obtained through the correlation between the phase correlation value and the hardware status data, and the timing distribution characteristics of the phase offset value are used to determine the parameter adjustment amplitude of the confidence assessment model to obtain the comprehensive linkage parameter.

[0144] In a specific embodiment, the sensitivity of the confidence assessment model to the timestamp distribution is obtained by the correlation between the phase correlation value and the hardware status data, and the parameter adjustment amplitude of the confidence assessment model is determined by using the timing distribution characteristics of the phase offset value to obtain the comprehensive linkage parameter, including:

[0145] Using a moving average method to extract the time series distribution characteristics of the phase correlation value;

[0146] Obtaining hardware status data of the time-deviation node, analyzing the correlation between the phase correlation value and the node status using the Pearson correlation coefficient in combination with the timing distribution characteristics, and using the correlation as the sensitivity of the confidence assessment model to the timestamp distribution;

[0147] When the sensitivity exceeds a preset sensitivity threshold, a linear regression algorithm is used to adjust the parameters of the confidence assessment model, an adjustment range is obtained, and an optimized model parameter set is obtained;

[0148] The optimized model parameter set and the hardware status data are synthesized by a principal component weighted synthesis method to generate a comprehensive linkage parameter set.

[0149] Specifically, the phase correlation value sequence obtained from step S18 and the hardware status data from step S16 are used as input. The phase correlation value indicates the degree of phase synchronization between the node clock and the reference clock, and the hardware status data includes CPU usage, temperature, and memory occupancy parameters. First, the time series data of the phase correlation value is smoothed using a moving average method. For example, the moving average is calculated over a fixed time window (e.g., 10 minutes) to eliminate short-term noise interference, extract the long-term trend characteristics of the phase offset, and form a time series distribution feature sequence.

[0150] Align the time series distribution feature sequence with the hardware status data by timestamp, and analyze the linear correlation between the two using the Pearson correlation coefficient. For example, the correlation coefficient between temperature and phase correlation values ​​is calculated. If the coefficient is close to negative 1, it indicates that increased temperature leads to decreased phase synchronization. This correlation coefficient is defined as the confidence assessment model's sensitivity to the timestamp distribution and is used to quantify the impact of hardware status on timestamp credibility.

[0151] If the sensitivity exceeds a preset sensitivity threshold (e.g., 0.7), a linear regression algorithm is used to adjust the parameters of the confidence assessment model. The regression model uses hardware status data as the independent variable and time series distribution characteristics as the dependent variable. The model weights are updated by minimizing the prediction residuals. For example, the weight allocation to high-temperature nodes is reduced, thereby correcting the bias of the confidence assessment model. The adjusted parameters form the optimized model parameter set.

[0152] Finally, the optimized model parameter set is fused with the hardware status data using a weighted principal component synthesis method. Principal components (such as the main influencing factors of temperature and load) are extracted and weighted based on their variance contributions. For example, the temperature principal component has a weight of 0.6, and the load principal component has a weight of 0.4. This weighted synthesis generates a comprehensive linkage parameter set that reflects the dynamic relationship between hardware status and model parameters and guides subsequent clock calibration strategies.

[0153] This step dynamically perceives the relationship between hardware status and phase synchronization, adaptively optimizes the confidence assessment model, and ensures that the timestamp credibility assessment is more in line with the actual operating environment, thereby improving the robustness of multi-sensor data fusion in complex scenarios.

[0154] In step S20, the clock parameter value of the deviation node is adjusted according to the phase correlation value, comprehensive linkage parameter, stability assessment result and network impact weight set. If the linkage between the network delay data and the hardware status data exceeds expectations, a calibration timestamp set is generated.

[0155] In a specific embodiment, adjusting the clock parameter value of the deviation node according to the phase correlation value, the comprehensive linkage parameter, the stability assessment result, and the network impact weight set, and generating a calibration timestamp set if the linkage between the network delay data and the hardware status data exceeds expectations, includes:

[0156] Calculating an initial clock compensation amount using an incremental PID algorithm according to the phase correlation value and the network influence weight set;

[0157] Verifying feature linkage anomalies using a chi-square test based on the comprehensive linkage parameter set; and when feature linkage anomalies are present, performing a secondary correction on the initial clock compensation amount using an L2 regularized logistic regression model based on the stability assessment result to obtain a final clock compensation amount.

[0158] When there is no characteristic linkage anomaly, the initial clock compensation amount is used as the final clock compensation amount;

[0159] A final calibration timestamp set is generated according to the final clock compensation amount.

[0160] Specifically, the phase correlation value from step S18, the comprehensive linkage parameter set generated in step S19, the stability assessment result from step S14, and the network influence weight set from step S15 are obtained as input data. The phase correlation value quantifies the degree of phase deviation between the node and the reference clock. The comprehensive linkage parameter set contains the dynamic association weights between hardware status and model parameters. The stability assessment result reflects the stability of the node's clock fluctuations. The network influence weight set describes the degree of influence of the node on network performance.

[0161] An incremental PID algorithm dynamically adjusts the compensation based on the phase correlation value and a set of network influence weights. The algorithm generates a base compensation based on the current phase deviation (proportional term), the historical accumulated deviation (integral term), and the rate of change of the deviation (differential term). This is then combined with network influence weights (e.g., compensation for high-weight nodes is prioritized) to form the initial clock compensation. For example, if a node's phase correlation value is low due to network congestion, its network influence weight is high, and the algorithm will automatically increase the compensation to quickly correct the deviation.

[0162] The comprehensive linkage parameter set is input into the chi-square test module to detect whether the linkage between hardware state parameters (such as temperature and load) and timing distribution characteristics deviates from the expected pattern. If the test statistic exceeds the critical value, it is determined that there is a feature linkage anomaly (for example, a sudden increase in node load in a high-temperature environment causes an abnormal phase shift). At this time, combined with the stability assessment results (such as the node's mean square error stability score), an L2 regularized logistic regression model is used to correct the initial clock compensation. The regression model uses the stability score to constrain the compensation adjustment range to avoid overcorrection, while suppressing noise interference through regularization, and outputs the final clock compensation value.

[0163] In the chi-square test module, the system first extracts the joint distribution data of hardware status data (such as temperature and load) and the timing distribution characteristics of model parameter adjustments (such as phase offset sensitivity and compensation change rate) from the integrated linkage parameter set, and constructs a multidimensional contingency table. For example, the temperature is divided into three ranges: "low temperature (<40°C)", "medium temperature (40-60°C)", and "high temperature (>60°C)". The model parameter adjustment range is divided into three levels: "fine adjustment (±5%)", "medium adjustment (±10%)", and "emphasis (±15%)". The actual frequency of occurrence of different adjustment ranges in each temperature range is counted to form a 3x3 observation frequency matrix.

[0164] The system calculates the expected distribution of parameter adjustments for each temperature range based on historical normal operating data. For example, in low-temperature environments, the expected distribution of fine-tuning is 60%, medium-tuning is 30%, and emphasis is 10%. In medium-temperature environments, the distributions are 50%, 35%, and 15%, and in high-temperature environments, the distributions are 40%, 40%, and 20%. By comparing the overall deviation between the observed and expected frequencies, a chi-square statistic is calculated: if the actual frequency of "emphasis" adjustments in a high-temperature node cluster reaches 45% (far exceeding the expected 20%), and the frequency of "fine-tuning" in the low-temperature node cluster drops to 40% (below the expected 60%), the statistic will increase significantly. When the calculated value is compared with a dynamic threshold (such as a critical value of 5.991, set based on 2 degrees of freedom and a significance level of 0.05), if the statistic exceeds the threshold (for example, reaching 7.8), it is determined that there is a feature linkage anomaly, triggering the secondary correction mechanism. This threshold is not fixed. The distribution characteristics will be recalculated every 6 hours based on the latest 1,000 sets of normal data. For example, when the ambient temperature rises seasonally, causing the proportion of high-temperature scenes to increase by 30%, the system will automatically increase the expected emphasis ratio of the high-temperature range from 20% to 25%, ensuring that the inspection standard dynamically adapts to changes in the operating environment.

[0165] If no feature linkage anomalies are detected, the initial clock offset is used as the final value. Based on the final clock offset, the local clocks of the deviating nodes are numerically adjusted to generate a set of calibrated timestamps synchronized with the reference clock. For example, if a node requires a 2-millisecond offset, its subsequent timestamps will be incremented by 2 milliseconds to ensure consistency with the global clock.

[0166] This step adaptively responds to network fluctuations and hardware status changes through multi-source data fusion and dynamic compensation mechanisms, significantly improving the accuracy and robustness of time synchronization in complex environments and providing a highly consistent timing benchmark for distributed systems.

[0167] Reference Figure 2 The second embodiment of the present invention provides a software time synchronization system for multi-sensor data fusion, including:

[0168] The data acquisition module is used to obtain working status data and historical clock drift records from sensor nodes, and uses principal component analysis to extract feature vectors containing temperature, load and drift frequency to obtain the node status feature set;

[0169] A confidence analysis module is used to calculate the confidence of each node timestamp based on the node state feature set and adopt a pre-established confidence evaluation model to obtain a timestamp confidence distribution;

[0170] An abnormal node analysis module is used to group the timestamp confidence distribution using a k-means algorithm and identify abnormal node groups using an isolation forest algorithm to obtain abnormal node groups;

[0171] A stability evaluation module is used to analyze the phase fluctuation frequency of the phase offset value in the abnormal node group and the timing drift frequency of the timing data, and to determine whether the node is a high dynamic deviation node to obtain a stability evaluation result;

[0172] a network impact analysis module configured to analyze the interaction effect data between the phase offset values ​​and the network delay in the abnormal node group and calculate the delay distribution characteristics of the network. If the correlation strength between the interaction effect data and the delay distribution characteristics exceeds a preset threshold, the node is determined to be a network bottleneck node and a network impact weight set is obtained.

[0173] Deviation node analysis module, used to extract network delay data and hardware status data from the abnormal node group, and perform time deviation node judgment to obtain a time deviation node list;

[0174] A clock drift analysis module is used to obtain the clock drift data of the node from the deviation node list, extract the nonlinear drift characteristics using the prophet algorithm, and obtain the nonlinear drift parameters;

[0175] A phase correlation calculation module, configured to calculate a phase correlation value between the deviation node and the reference clock data using a Hilbert cross-correlation phase detection algorithm according to the nonlinear drift parameter;

[0176] A linkage parameter analysis module is used to obtain the sensitivity of the confidence assessment model to the timestamp distribution through the correlation between the phase correlation value and the hardware status data, and to determine the parameter adjustment amplitude of the confidence assessment model using the timing distribution characteristics of the phase offset value to obtain the comprehensive linkage parameter;

[0177] A time calibration module is used to adjust the clock parameter value of the deviation node according to the phase correlation value, the comprehensive linkage parameter, the stability assessment result and the network impact weight set, and generate a calibration timestamp set if the linkage between the network delay data and the hardware status data exceeds expectations.

[0178] It should be noted that the software time synchronization device for multi-sensor data fusion provided in an embodiment of the present invention is used to execute all the process steps of the software time synchronization method for multi-sensor data fusion in the above embodiment. The working principles and beneficial effects of the two correspond one to one, so they will not be repeated here.

[0179] An embodiment of the present invention further provides an electronic device. The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a software time synchronization program for multi-sensor data fusion. When the processor executes the computer program, the steps of the above-mentioned software time synchronization method for multi-sensor data fusion are implemented, such as Figure 1 Alternatively, when the processor executes the computer program, the functions of the modules / units in the above-mentioned device embodiments are realized, such as the software time synchronization module for multi-sensor data fusion.

[0180] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0181] The electronic device may be a computing device such as a desktop computer, notebook, PDA, or smart tablet. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will appreciate that the aforementioned components are merely examples of electronic devices and do not constitute a limitation of the electronic device. The electronic device may include more or fewer components than those described above, or a combination of certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, and the like.

[0182] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the electronic device, connecting various parts of the entire electronic device using various interfaces and lines.

[0183] The memory can be used to store the computer programs and / or modules, and the processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created based on the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory can include a high-speed random access memory and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0184] Wherein, if the module / unit integrated in the electronic device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0185] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive effort.

[0186] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A software time synchronization method for multi-sensor data fusion, characterized in that: include: Obtain working status data and historical clock drift records, use principal component analysis to extract feature vectors, and obtain node status feature sets; According to the node state feature set, a pre-established confidence evaluation model is used to calculate the confidence of each node timestamp to obtain a timestamp confidence distribution; The timestamp confidence distribution is grouped using a k-means algorithm, and abnormal node groups are identified using an isolation forest algorithm to obtain abnormal node groups; Analyze the phase fluctuation frequency of the phase offset value in the abnormal node group and the timing drift frequency of the timing data, and perform high dynamic deviation node judgment to obtain a stability assessment result; Analyzing the interaction effect data of the phase offset value and the network delay in the abnormal node group, and calculating the delay distribution characteristics of the network. If the correlation strength between the interaction effect data and the delay distribution characteristics exceeds a preset threshold, the node is determined to be a network bottleneck affecting node, and a network influence weight set is obtained; Extracting network delay data and hardware status data from the abnormal node group, and performing time deviation node judgment to obtain a time deviation node list; Obtaining the clock drift data of the node from the deviation node list, extracting the nonlinear drift characteristics using the prophet algorithm, and obtaining the nonlinear drift parameters; Calculating the phase correlation value between the deviation node and the reference clock data using the Hilbert cross-correlation phase detection algorithm according to the nonlinear drift parameter; Obtaining the sensitivity of the confidence assessment model to timestamp distribution and the time series distribution characteristics of the phase correlation value through the correlation between the phase correlation value and the hardware status data, determining the parameter adjustment amplitude of the confidence assessment model, and obtaining the comprehensive linkage parameter; According to the phase correlation value, comprehensive linkage parameter, stability assessment result and network impact weight set, the clock parameter value of the deviation node is adjusted, and if the linkage between the network delay data and the hardware status data exceeds expectations, a calibration timestamp set is generated.

2. The software time synchronization method for multi-sensor data fusion according to claim 1, characterized in that: The method of calculating the confidence of each node timestamp using a pre-established confidence evaluation model based on the node state feature set to obtain a timestamp confidence distribution includes: Acquire feature data of a node timestamp from the node state feature set, and input the data into a pre-established confidence assessment model to obtain a confidence score; Comparing the confidence score with a preset confidence threshold, and when the confidence score is lower than the confidence threshold, marking it as a confidence deviation node, to obtain a confidence deviation node set; Extracting feature data of node timestamps in the confidence deviation node set, and analyzing timestamp confidence distribution using the OPTICS density clustering algorithm; Wherein, the confidence assessment model is constructed by gradient boosting decision tree algorithm; The characteristic data of the node timestamp include characteristic vectors of temperature, load and drift frequency.

3. The software time synchronization method for multi-sensor data fusion according to claim 1, characterized in that: The analysis of the phase fluctuation frequency of the phase offset value and the timing drift frequency of the timing data in the abnormal node group, and the determination of the high dynamic deviation node to obtain the stability assessment result includes: Acquire time series data of phase offset values ​​and clock drift from the abnormal node group, calculate the phase fluctuation frequency and timing drift frequency using Fourier transform, and obtain a frequency distribution result; According to the frequency distribution result, a multi-band coherence analysis algorithm is used to calculate the coherence coefficient of the phase fluctuation frequency and the timing drift frequency. If the coherence coefficient is higher than a preset coherence threshold, it is determined to be a high dynamic state and a high dynamic flag is obtained; Extracting deviation features from abnormal nodes based on the high dynamic identification, and using a K-means clustering algorithm to determine a set of high dynamic deviation nodes; The stability evaluation value is calculated by using the mean square error through the high dynamic deviation node set and the time series data, and is used as the stability evaluation result.

4. The software time synchronization method for multi-sensor data fusion according to claim 1, characterized in that: The analysis of the interaction effect data of the phase offset value and the network delay in the abnormal node group and the calculation of the network delay distribution characteristics are performed. If the correlation strength between the interaction effect data and the delay distribution characteristics exceeds a preset threshold, the node is determined to be a network bottleneck affecting node, and a network influence weight set is obtained, including: Acquire network delay data and phase offset values ​​from the abnormal node group, and calculate the interaction effect of the network delay data and the phase offset value based on a sliding window cross-correlation analysis method to obtain interaction effect data; According to the interaction effect data, the skewness and kurtosis of the network delay distribution are calculated using the moment estimation method to obtain the delay distribution characteristics; Calculating the typical correlation coefficient between the delay distribution characteristics and the interaction effect data by a typical correlation analysis method; and when the typical correlation coefficient exceeds a preset typical correlation threshold, using the DBSCAN algorithm to classify the abnormal node group and identify the network bottleneck affecting node set; A network influence weight set is generated by adopting a weighted average algorithm according to the network bottleneck influence node set and the interaction effect data, thereby obtaining a network influence weight set.

5. The software time synchronization method for multi-sensor data fusion according to claim 1, characterized in that: The network delay data and hardware status data are extracted from the abnormal node group, and time deviation node judgment is performed to obtain a time deviation node list, including: When the network delay data exceeds the preset delay threshold, it is marked as a delay abnormal node; When the hardware status data deviates from the preset normal range, it is marked as an abnormal status node; A fast set operation algorithm based on a bitmap is used to compare the intersection of the delay abnormal node and the state abnormal node, determine that the node in the intersection is a time deviation node, and generate a time deviation node list.

6. The software time synchronization method for multi-sensor data fusion according to claim 1, characterized in that: The sensitivity of the confidence assessment model to the timestamp distribution is obtained by the correlation between the phase correlation value and the hardware status data, and the parameter adjustment amplitude of the confidence assessment model is determined by using the timing distribution characteristics of the phase offset value to obtain the comprehensive linkage parameter, including: Using a moving average method to extract the time series distribution characteristics of the phase correlation value; Obtaining hardware status data of the time-deviation node, analyzing the correlation between the phase correlation value and the node status using the Pearson correlation coefficient in combination with the timing distribution characteristics, and using the correlation as the sensitivity of the confidence assessment model to the timestamp distribution; When the sensitivity exceeds a preset sensitivity threshold, a linear regression algorithm is used to adjust the parameters of the confidence assessment model, an adjustment range is obtained, and an optimized model parameter set is obtained; The optimized model parameter set and the hardware status data are synthesized by a principal component weighted synthesis method to generate a comprehensive linkage parameter set.

7. The software time synchronization method for multi-sensor data fusion according to claim 1, characterized in that: The step of adjusting the clock parameter value of the deviation node according to the phase correlation value, the comprehensive linkage parameter, the stability assessment result, and the network impact weight set, and generating a calibration timestamp set if the linkage between the network delay data and the hardware status data exceeds expectations, includes: Calculating an initial clock compensation amount using an incremental PID algorithm according to the phase correlation value and the network influence weight set; Verifying feature linkage anomalies using a chi-square test based on the comprehensive linkage parameter set; and when feature linkage anomalies are present, performing a secondary correction on the initial clock compensation amount using an L2 regularized logistic regression model based on the stability assessment result to obtain a final clock compensation amount. When there is no characteristic linkage anomaly, the initial clock compensation amount is used as the final clock compensation amount; A final calibration timestamp set is generated according to the final clock compensation amount.

8. A software time synchronization system for multi-sensor data fusion, characterized in that: include: The data acquisition module is used to obtain working status data and historical clock drift records from sensor nodes, and uses principal component analysis to extract feature vectors containing temperature, load and timing drift frequency to obtain the node status feature set; A confidence analysis module is used to calculate the confidence of each node timestamp based on the node state feature set and adopt a pre-established confidence evaluation model to obtain a timestamp confidence distribution; An abnormal node analysis module is used to group the timestamp confidence distribution using a k-means algorithm and identify abnormal node groups using an isolation forest algorithm to obtain abnormal node groups; A stability evaluation module is used to analyze the phase fluctuation frequency of the phase offset value in the abnormal node group and the timing drift frequency of the timing data, and to determine whether the node is a high dynamic deviation node to obtain a stability evaluation result; a network impact analysis module configured to analyze the interaction effect data between the phase offset values ​​and the network delay in the abnormal node group and calculate the delay distribution characteristics of the network. If the correlation strength between the interaction effect data and the delay distribution characteristics exceeds a preset threshold, the node is determined to be a network bottleneck node and a network impact weight set is obtained. Deviation node analysis module, used to extract network delay data and hardware status data from the abnormal node group, and perform time deviation node judgment to obtain a time deviation node list; A clock drift analysis module is used to obtain the clock drift data of the node from the deviation node list, extract the nonlinear drift characteristics using the prophet algorithm, and obtain the nonlinear drift parameters; A phase correlation calculation module, configured to calculate a phase correlation value between the deviation node and the reference clock data using a Hilbert cross-correlation phase detection algorithm according to the nonlinear drift parameter; A linkage parameter analysis module is used to obtain the sensitivity of the confidence assessment model to the timestamp distribution through the correlation between the phase correlation value and the hardware status data, and to determine the parameter adjustment amplitude of the confidence assessment model using the timing distribution characteristics of the phase offset value to obtain the comprehensive linkage parameter; A time calibration module is used to adjust the clock parameter value of the deviation node according to the phase correlation value, the comprehensive linkage parameter, the stability assessment result and the network impact weight set, and generate a calibration timestamp set if the linkage between the network delay data and the hardware status data exceeds expectations.

9. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the software time synchronization method for multi-sensor data fusion according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the software time synchronization method for multi-sensor data fusion according to any one of claims 1 to 7.

Citation Information

Cited By

  • Clock compensation method and device for fusing terminal and electric energy meter, equipment and medium

    CN121027971A

  • Three-phase electric power parameter wireless synchronous measurement and vector analysis method and system

    CN121069008A

  • EMMC performance difference test method and system based on three-level hardware timestamps

    CN121506223A

  • Self-adaptive temperature control method and system for graphene electric heater

    CN121677032A

  • Power-saving protection device control method and system based on multiple sensors

    CN122315938A