Intelligent driving state analysis method and system based on multi-source heterogeneous data fusion
By aligning and trusting data from vehicle sensors and roadside equipment with timestamps, and combining environmental variables for dynamic weighting and evidence fusion, the problem of credibility assessment of multi-source heterogeneous data in complex environments is solved, thereby improving the reliability and prediction accuracy of driving status analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DUOPAI (SHENZHEN) CLOUD TECH CO LTD
- Filing Date
- 2026-01-20
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies struggle to scientifically assess the credibility and dynamically assign weights to multi-source heterogeneous data in complex and ever-changing environments, resulting in low reliability and predictive accuracy of driving status analysis results.
By acquiring real-time data from vehicle-mounted sensors and roadside equipment, timestamp alignment and field mapping are performed to construct a multi-source input dataset. Trust is assessed using evidence theory and a Bayesian trust evaluation model, dynamic weight allocation is performed in conjunction with environmental variables, and multi-source evidence fusion is performed using weighted evidence combination rules. Finally, iterative calculations and temporal smoothing filtering are performed in a state evolution Bayesian network.
It significantly improves the temporal consistency and availability of multi-source data, enhances the system's adaptability in complex environments, improves the reliability and prediction accuracy of driving status analysis, and reduces the misjudgment rate.
Smart Images

Figure CN122067397A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent transportation technology, and in particular to a method and system for intelligent analysis of driving status based on the fusion of multi-source heterogeneous data. Background Technology
[0002] In modern intelligent transportation systems, intelligent analysis of vehicle status has become a key pillar for improving road safety and traffic efficiency. With the rapid development of the Internet of Things (IoT) and edge computing technologies, the real-time collection and processing of massive amounts of traffic data using edge nodes deployed on the roadside and in vehicles has gradually become a hot research topic in the industry. Technological advancements in this field provide important strategic support for the digital transformation and sustainable development of urban traffic management.
[0003] In existing technologies, traffic status analysis typically relies on the simple aggregation and processing of data from vehicle-mounted sensors (such as speed and location) and roadside sensing devices (such as traffic flow and video streams). This approach often employs pre-defined fixed weights or simple weighted average algorithms to directly fuse data from different sources to calculate a traffic index. However, this method ignores the quality differences and reliability fluctuations of multi-source heterogeneous data under various environmental conditions (such as severe weather or signal obstruction). For example, when there is a logical contradiction between speed readings from vehicle-mounted sensors and traffic flow statistics from roadside cameras, existing technologies lack an effective trust assessment mechanism to distinguish the credibility of each data source. Because the trust weights of different information channels cannot be dynamically adjusted according to the real-time environment, the system often struggles to balance differences when facing conflicts in multi-source information, thus failing to construct a unified and accurate judgment basis.
[0004] In summary, existing technologies struggle to scientifically assess the reliability and dynamically assign weights to multi-source heterogeneous data in complex and ever-changing environments, resulting in low reliability and predictive accuracy of driving status analysis results. Summary of the Invention
[0005] This invention provides a method and system for intelligent analysis of driving status based on multi-source heterogeneous data fusion, so as to achieve accurate analysis and prediction of driving status.
[0006] Firstly, to address the aforementioned technical problems, this invention provides a method for intelligent analysis of vehicle status based on multi-source heterogeneous data fusion, comprising: Real-time speed data from vehicle-mounted sensors and traffic flow information from roadside equipment are acquired respectively. The real-time speed data and traffic flow information are then time-stamped and mapped to obtain a multi-source input dataset. Retrieve historical calibration records corresponding to the vehicle-mounted sensors and the roadside equipment, perform evidence theory calculations on the historical calibration records, construct confidence functions and likelihood functions, and calculate the source confidence score. Real-time environmental variable data is acquired, and the environmental variable data and the source trust score are input into a preset trust assessment Bayesian model to perform conditional probability inference and obtain an environmental adaptability vector. The numerical deviation characteristics between the real-time speed data and the traffic flow information are calculated, and dynamic weight allocation is performed in combination with the environmental adaptability vector to obtain a fusion weight matrix; The dominant weight vector is extracted from the fusion weight matrix, and the multi-source evidence fusion calculation is performed on the multi-source input dataset using the weighted evidence combination rule to obtain the comprehensive driving status index. The comprehensive driving status index is input into a preset state evolution Bayesian network, and the state evolution is deduced and iteratively calculated by combining the environmental variable data and the real-time speed data to obtain a refined state prediction sequence. The refined state prediction sequence is subjected to temporal smoothing filtering, and the final driving state classification result is determined by matching the filtered sequence features with a preset decision threshold.
[0007] Secondly, the present invention provides a vehicle status intelligent analysis system based on multi-source heterogeneous data fusion, comprising: The multi-source data processing module acquires real-time speed data from vehicle-mounted sensors and traffic flow information from roadside equipment, and performs timestamp alignment and field mapping on the real-time speed data and traffic flow information to obtain a multi-source input dataset. The source-end trust assessment module retrieves historical calibration records corresponding to the vehicle-mounted sensors and the roadside equipment, performs evidence theory calculations on the historical calibration records, constructs a confidence function and a likelihood function, and calculates the source-end trust score. The environmental adaptability inference module acquires real-time environmental variable data, inputs the environmental variable data and the source trust score into a preset trust assessment Bayesian model, performs conditional probability inference, and obtains an environmental adaptability vector. The dynamic weight allocation module calculates the numerical deviation characteristics between the real-time speed data and the traffic flow information, and performs dynamic weight allocation in conjunction with the environmental adaptability vector to obtain a fusion weight matrix. The multi-source evidence fusion module extracts the dominant weight vector from the fusion weight matrix and performs multi-source evidence fusion calculation on the multi-source input dataset using weighted evidence combination rules to obtain a comprehensive driving status index. The state evolution prediction module inputs the comprehensive driving state index into a preset state evolution Bayesian network, and combines the environmental variable data and the real-time speed data to perform state evolution inference and iterative calculation to obtain a refined state prediction sequence. The state classification decision module performs time-series smoothing filtering on the refined state prediction sequence and determines the final driving state classification result by matching the filtered sequence features with a preset decision threshold.
[0008] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention acquires data from vehicle-mounted sensors and roadside equipment separately, and uses a multimodal data synchronization protocol to perform microsecond-level timestamp alignment and field mapping to construct a unified format multi-source input dataset. This effectively overcomes the timing deviation and data fragmentation problems caused by asynchronous sampling frequencies of heterogeneous devices in the prior art, significantly improves the consistency and availability of multi-source data in the time dimension, and provides a solid and unified data foundation for subsequent high-precision fusion analysis, thereby ensuring the initial input quality of the analysis system.
[0009] (2) This invention innovatively introduces evidence theory to calculate historical calibration records to establish source-end baseline scores, and further utilizes a Bayesian trust assessment model to fuse real-time environmental variables (such as weather and road conditions) to infer an environmentally adaptive trust vector. This dynamic trust assessment mechanism overcomes the shortcomings of existing technologies that often use static weights or ignore environmental interference (such as the impact of rain and fog on visual devices), significantly enhancing the system's adaptability in complex and ever-changing environments. It can objectively and in real time identify the credibility of each data source, ensuring that subsequent analysis relies only on high-quality data sources.
[0010] (3) This invention calculates the numerical deviation characteristics between real-time speed and traffic flow, combines them with an environmentally adaptive trust vector for dynamic weight allocation, and utilizes weighted evidence combination rules for deep fusion. This effectively solves the "one-vote veto" paradox or weighted average distortion problem that occurs in existing technologies when faced with serious conflicts between multi-source data (such as logical contradictions between speed and traffic flow), significantly improves the robustness of the system in the event of sensor failure or abnormal data fluctuations, and can extract the comprehensive index that is closest to the real road conditions from contradictory information, greatly reducing the misjudgment rate.
[0011] (4) This invention maps comprehensive driving state indicators to a state evolution Bayesian network, performs iterative deduction in conjunction with variable environmental signals, and uses a time-series smoothing filter for noise suppression to ultimately determine the driving state classification. This closed-loop mechanism based on evolutionary reasoning and filtering denoising breaks through the limitations of existing technologies that only focus on the current instantaneous state and ignore the evolutionary trend. It significantly improves the prediction accuracy and anti-interference ability of future traffic congestion trends, and can filter out false alarms caused by occasional noise such as temporary parking, providing traffic management departments with more stable and forward-looking decision support. Attached Figure Description
[0012] Figure 1 This is a schematic diagram of the intelligent vehicle status analysis method based on multi-source heterogeneous data fusion provided in the first embodiment of the present invention; Figure 2 This is a schematic diagram of the intelligent vehicle status analysis system based on multi-source heterogeneous data fusion provided in the second embodiment of the present invention. Detailed Implementation
[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] Reference Figure 1 The first embodiment of the present invention provides a method for intelligent analysis of driving status based on multi-source heterogeneous data fusion, including the following steps: S11, acquire real-time speed data from vehicle-mounted sensors and traffic flow information from roadside equipment respectively, and perform timestamp alignment and field mapping on the real-time speed data and traffic flow information to obtain a multi-source input dataset; S12, retrieve the historical calibration records corresponding to the vehicle-mounted sensor and the roadside equipment, perform evidence theory calculations on the historical calibration records, construct a confidence function and a likelihood function, and calculate the source confidence score; S13, acquire real-time environmental variable data, input the environmental variable data and the source trust score into a preset trust evaluation Bayesian model, perform conditional probability inference, and obtain an environmental adaptability vector; S14, calculate the numerical deviation characteristics between the real-time speed data and the traffic flow information, and perform dynamic weight allocation in combination with the environmental adaptability vector to obtain the fusion weight matrix; S15, extract the dominant weight vector from the fusion weight matrix, and use the weighted evidence combination rule to perform multi-source evidence fusion calculation on the multi-source input dataset to obtain the comprehensive driving status index. S16, input the comprehensive driving status index into a preset state evolution Bayesian network, and combine the environmental variable data and the real-time speed data to perform state evolution deduction and iterative calculation to obtain a refined state prediction sequence; S17, perform time-series smoothing filtering on the refined state prediction sequence, and determine the final driving state classification result by matching the filtered sequence features with a preset decision threshold.
[0015] In step S11, real-time speed data from vehicle-mounted sensors and traffic flow information from roadside equipment are acquired respectively. The real-time speed data and traffic flow information are then time-stamped and mapped to obtain a multi-source input dataset.
[0016] Specifically, to eliminate spatiotemporal discrepancies between heterogeneous data acquisition devices, the core of this step lies in executing a preset multimodal data synchronization protocol. First, according to the acquisition specifications defined by this protocol, real-time speed data is acquired via onboard sensors (e.g., CAN bus interface or high-frequency GPS module) at a first sampling frequency (e.g., 10Hz), while traffic flow information is acquired via roadside equipment (e.g., roadside units (RSU) or smart cameras) at a second sampling frequency (e.g., 1Hz). It is worth noting that the first sampling frequency (e.g., 10Hz) is determined based on vehicle dynamic response characteristics and the Nyquist sampling theorem, ensuring complete capture of instantaneous acceleration and deceleration changes of vehicles without spectral aliasing; while the second sampling frequency (e.g., 1Hz) is determined based on the update cycle requirements of macroscopic traffic flow statistics and the transmission bandwidth limitations of roadside equipment.
[0017] It should be noted that the real-time speed data is the raw speed signal with timestamps collected from the vehicle-mounted sensors, forming the basic component of the multi-source input dataset; the speed value mentioned in step S141 refers to the numerical feature extracted from the real-time speed data and obtained after normalization and other preprocessing during the dynamic weight allocation process for collaborative analysis with traffic flow information. Its main purpose is to eliminate dimensions and facilitate the calculation of consistency deviations across source data; the speed deviation value mentioned in step S152 refers to the feature used as weighted evidence during the multi-source evidence fusion process that can reflect changes in driving status. This feature does not directly refer to the difference between speed and flow, but may include, but is not limited to, dynamic parameters such as acceleration calculated based on real-time speed data, the difference between the current speed value and the expected speed value established based on historical data, or the statistical dispersion (such as standard deviation) of speed values of different vehicles within the same time period. These features together constitute the evidence reflecting the dynamics of driving status.
[0018] In one implementation, the absolute value of the timestamp deviation between the real-time speed data and the traffic flow information at the nearest neighbor time is calculated according to the timing alignment mechanism of the multimodal data synchronization protocol. If the timestamp deviation exceeds a preset synchronization threshold (e.g., 200ms), timing resampling alignment processing is triggered. The synchronization threshold (e.g., 200ms) is set by reverse calculation based on the principle that the vehicle's displacement under the maximum design speed limit (e.g., 60km / h) on urban roads does not exceed the lane-level positioning accuracy range (e.g., 3.5 meters) to ensure strict correspondence of data in the spatial dimension.
[0019] Preferably, in the resampling stage of the protocol, a linear interpolation algorithm is used to upsample the low-frequency traffic flow information. The specific execution logic of this algorithm is as follows: retrieve the target aligned time point (i.e., the timestamp of the real-time speed data). Two consecutive valid data points in the traffic flow time series Using the formula Calculate the flow rate at the target time point. Indicates the time point at which the target is aligned. The estimated flow rate value; and These represent the target time points on the timeline. The two most recent valid sampling times before and after; and Corresponding to time respectively and The actual flow rate values collected.
[0020] This approach leverages the physical assumption of continuous traffic flow changes over short periods, effectively filling the temporal gaps caused by differences in sampling frequencies. Optionally, if the protocol detects extremely low data transmission latency, it can also be configured to directly select the time distance using nearest neighbor interpolation. The minimum effective traffic data is used as the alignment value.
[0021] Furthermore, according to the format standardization rules of the multimodal data synchronization protocol, field mapping and integrity verification are performed. Based on a predefined unified data dictionary, heterogeneous feature fields in the original data stream (such as control bits in hexadecimal messages or key-value pairs in JSON messages) are parsed and mapped to standardized data structures. It is worth noting that if data is missing during verification, the protocol will invoke a sliding window mean-filling algorithm to calculate the average value within a preset window size (e.g., 5 sampling points) before and after that time point to complete the data. The window size (e.g., 5 sampling points) is determined based on the autocorrelation function (ACF) decay characteristics of historical traffic flow data, selecting a lag order where the correlation coefficient remains above 0.9, thereby ensuring data integrity while maintaining the authenticity of statistical characteristics.
[0022] In step S12, historical calibration records corresponding to the vehicle-mounted sensor and the roadside equipment are retrieved. Evidence theory calculations are performed on the historical calibration records to construct a confidence function and a likelihood function, and the source-end confidence score is calculated, including: S121, based on the device identification information of the vehicle-mounted sensor and the roadside equipment, retrieve and extract matching historical calibration records, and classify the historical calibration records into sensor calibration records and equipment calibration records; S122, perform basic probability allocation calculations on the sensor calibration record and the device calibration record respectively to obtain the source confidence function value and the source likelihood function value; S123, if the source confidence function value is greater than the preset confidence threshold, then increase the calculation weight of the source likelihood function value to generate a weighted likelihood function value; S124, the source confidence function value and the weighted likelihood function value are averaged and fused to obtain the source confidence score.
[0023] Specifically, to ensure the accuracy of the traceability data, the unique hardware identifiers (e.g., physical addresses or universally unique identifiers) of the vehicle-mounted sensors and roadside equipment are used for indexing and retrieval in a pre-built distributed database. The extraction time range is set to a preset past time window (e.g., the most recent 30 days). This time window length is based on the aging characteristic curves of the equipment components and the regular maintenance cycle, aiming to select the effective observation period that best reflects the current performance trend of the equipment. Subsequently, a classification operation is performed according to the equipment attribute tags, strictly dividing the extracted data into sensor calibration records belonging to the vehicle-mounted end and equipment calibration records belonging to the roadside end, establishing a data foundation for subsequent independent evaluations based on different physical characteristics.
[0024] In step S122, when calculating the basic probability allocation, the DS evidence theory is introduced as the core algorithm framework for handling uncertainty. For each type of calibration record, a Gaussian distribution model is used to fit the historical calibration error, and the probability of the error falling within the standard deviation range is mapped to a basic probability allocation value, denoted as . Based on this allocation value, two core indicators are calculated. One is the source confidence function value (Bel), calculated using the following formula: .in, This represents the assumption that "the device is trustworthy." The recognition framework includes the proposition All subsets within; Indicates assignment to that specific subset The basic probability allocation value (i.e., quality of evidence) characterizes the degree of certainty with which evidence supports the proposition that "the device is credible." Another indicator is the source likelihood function value (Pl), calculated using the formula... This represents the highest probability that the evidence does not contradict the proposition. This two-dimensional evaluation method effectively overcomes the limitation of traditional Bayesian probability in distinguishing between "don't know" and "uncertainty".
[0025] Regarding the weight enhancement process in step S123, to strengthen the positive influence of high-reputation devices in the fusion decision, a confidence threshold based on statistical significance (e.g., 0.85) is set. This threshold is determined by performing cumulative probability density statistical analysis on the reputation distribution of a large number of historical normal devices, selecting the quantile where the cumulative probability reaches 85%. When the calculated source confidence function value... When the value exceeds this threshold, it indicates that the device exhibits extremely high stability in historical records. In this case, an exponential gain function is used to dynamically weight and boost the synchronously calculated source likelihood function value. For example, using the formula... Generate the weighted likelihood function value, where The confidence threshold is... This is an adjustment coefficient (e.g., 0.5). It is worth explaining in detail that the settings of each parameter in the above formula are based on strict scientific principles: confidence threshold... For example, 0.85 is determined based on experimental statistics. This is achieved by constructing a reputation sample set containing thousands of historically operating devices, performing cumulative probability density analysis (CDF) on their source confidence function values, and selecting the quantile with a cumulative probability of 85% as the threshold; while the adjustment coefficient... For example, 0.5 is determined based on a grid search algorithm, that is, within a preset interval during the model training phase. Iterative testing is conducted to select parameter values that maximize the F1 score of the final fusion result on the validation set, thereby achieving the best balance between trust enhancement and overfitting risk.
[0026] Finally, in step S124, when calculating the source-end trust score, the final trust quantification is performed. A weighted arithmetic mean is used to combine the source confidence function values representing "certain trust". The weighted likelihood function value representing the "potential trust ceiling" The fusion process is then performed. The calculation formula can be expressed as follows: Weighting factors are typically set. For example, a weighting of 0.6 slightly emphasizes definitive evidence. This weighting factor is based on the risk aversion principle in driving safety assessment, meaning that proven reliability has higher decision-making value than potential probability. The final calculated value is the source-end trust score of the device, used to measure the overall credibility level of the data source at the current moment.
[0027] In step S13, real-time environmental variable data is acquired, and the environmental variable data and the source-end trust score are input into a preset trust assessment Bayesian model to perform conditional probability inference, thereby obtaining an environmental adaptability vector, including: S131, acquire the current weather and road condition signals, perform feature analysis and quantification, and extract weather condition factors and road condition factors as the environmental variable data; S132, Using the weather condition factor and the road surface condition factor as condition nodes, calculate the conditional probability distribution of the source trust score under different environmental conditions; S133, Determine the probability node values of the vehicle-mounted sensor and the roadside equipment based on the conditional probability distribution; S134, if the probability node value exceeds the preset reliability threshold, the probability node value and the source trust score are weighted and fused to generate an adjusted trust score. S135, Based on the adjusted trust score, a combined environmental adaptability vector is generated.
[0028] Specifically, for step S131, in order to transform unstructured physical environment signals into mathematical features that can be used for probabilistic reasoning, multidimensional meteorological signals (such as rainfall and light intensity) and road surface condition images are first collected. For the meteorological signals (rainfall rate), continuous data is obtained using a 1Hz sampling frequency, and a 5-minute moving average is taken as the feature value. For the road surface condition images, the Canny edge detection algorithm is used to extract texture features, and the road surface friction level is obtained through support vector machine classification.
[0029] Subsequently, feature parsing and Min-Max normalization are performed to map the physical signal to... Dimensionless factors within the interval. For example, weather condition factors (0 represents no interference, 1 represents extreme interference) and road surface condition factors (0 represents high friction, 1 represents low friction). It is worth explaining in detail that the quantization mapping function is constructed based on the sensor signal-to-noise ratio (SNR) attenuation characteristic curve. For example, based on experimental data of signal attenuation from millimeter-wave radar at different rainfall rates, a nonlinear mapping relationship between rainfall intensity and factor values is established to accurately characterize the degree of environmental interference on the sensor's physical performance.
[0030] In steps S132 and S133, a Bayesian belief network is introduced as the core inference model, and a conditional probability table is constructed based on historical data. First, a directed acyclic graph (DAG) is constructed, with the weather condition factor and the road surface condition factor set as parent nodes (condition nodes), and "data source reliability" set as a child node.
[0031] During the training and construction of the conditional probability table, a multi-source heterogeneous dataset covering a preset historical period (e.g., the past 12 months) is collected. This dataset contains paired samples of "environmental characteristics - device performance". Maximum likelihood estimation is used to perform statistical analysis on the samples. For each specific environmental combination... For example, "heavy rain + slippery conditions" and equipment reliability status. (For example, "high reliability"), count the joint frequency of its occurrence in historical data. and the total frequency of this environmental combination. Using formulas Calculate the original conditional probability.
[0032] It is worth explaining in detail that, in order to address the "zero probability" problem that data sparsity may cause (i.e., certain extreme environmental combinations have not appeared in historical data, resulting in a probability of 0), the Laplacian smoothing technique must be introduced during training. This is achieved by adding 1 to the numerator (…). Add the total number of states to the denominator. The revised calculation formula is as follows: This process ensures that the model can still provide a non-zero prior probability when facing unknown or rare environments, thus guaranteeing the robustness of the inference.
[0033] Based on the pre-trained conditional probability table, online Bayesian inference is performed to calculate the posterior probability distribution under the current real-time environment input. Subsequently, the mathematical expectation method is used to extract statistical features from the posterior probability distribution, and scalarized probability node values are calculated. These values range from [0,1], with values closer to 1 indicating higher reliability of the data source under the current environment, consistent with the dimensions of the source trust score. This value intuitively quantifies the expected performance of the data source under the current environmental pressure.
[0034] For the fusion decision in step S134, to prevent noise from being introduced by low-confidence environmental inferences, a reliability threshold based on statistical analysis (e.g., 0.6) is set. This threshold is determined based on receiver operating characteristic curve analysis. By calculating the maximum point of the Youden index, a critical value that optimally balances the false alarm rate and the false negative rate is selected to distinguish the boundary between "controllable environmental impact" and "uncontrollable environmental impact." If the calculated probability node value exceeds this threshold (e.g., roadside equipment maintains a node value of 0.75 in rainy weather), it indicates that the environmental impact is within an acceptable range. In this case, a weighted linear fusion algorithm is used to fuse this probability node value with the source-end confidence score obtained in step S12 to generate an adjusted confidence score. The fusion weights are... (For example, 0.4) is dynamically set based on the information entropy of Bayesian inference. The smaller the entropy value, the higher the certainty, and the greater the weight assigned. Conversely, if the probability node value is lower than the threshold, the original score is kept unchanged or the weight is reduced to avoid the decision-making risk caused by unreliable environmental inference.
[0035] Finally, in step S135, the adjusted trust scores for each dimension (vehicle-mounted and roadside) after the aforementioned environmental adaptability correction are encapsulated according to a preset tensor order and combined to generate a multi-dimensional environmental adaptability vector. The preset tensor order is vehicle-mounted trust scores first, followed by roadside trust scores, a sequence that corresponds one-to-one with the data source dimensions for subsequent dynamic weight allocation. This vector not only contains the device's historical reputation information but also incorporates real-time corrections to perception performance based on the current physical environment, thus providing a robust and environmentally-aware reference benchmark for subsequent dynamic weight allocation.
[0036] In step S14, the numerical deviation characteristics between the real-time speed data and the traffic flow information are calculated, and dynamic weight allocation is performed in conjunction with the environmental adaptability vector to obtain a fusion weight matrix, including: S141, extract the speed value from the real-time speed data and the flow value from the traffic flow information, and perform a normalization difference calculation on the speed value and the flow value to obtain the numerical deviation characteristics; S142, Based on the numerical deviation characteristics and the environmental adaptability vector, the weight coefficients of each data source are calculated using a weighted average method; S143, if the weight coefficient is greater than the preset allocation threshold, the weight coefficient is corrected using the environmental variable data to obtain a corrected weight set, and the weight set is matrix-combined to obtain a fused weight matrix.
[0037] Specifically, regarding step S141, for the heterogeneous data collected by vehicle-mounted sensors and roadside equipment, the first step is to address the problem of incompatible units (speed is in km / h, flow rate is in veh / h) leading to incomparability. A max-min normalization algorithm is used to map the speed values to... For the interval, the denominator parameter is set to the design maximum speed of that road segment (e.g., 80 km / h); similarly, the flow rate value is mapped to... The denominator parameter for the interval is set to the saturation capacity of that lane (e.g., 1800 pcu / h). It is worth noting that the saturation capacity setting is based on the definition of the theoretical maximum traffic volume of a single lane on urban arterial roads in the *Road Capacity Manual*. When performing normalized difference calculations, when the trends of speed and flow are linearly correlated, the absolute difference is used to calculate the numerical deviation characteristic; when the trends are non-linearly correlated, Euclidean distance is used. A larger characteristic value (e.g., close to 1) indicates a more severe contradiction between "high flow rate and low flow rate" or "low flow rate and high flow rate," suggesting a significant logical conflict between data sources.
[0038] Subsequently, in step S142, to identify high-confidence sources in the conflict data, a variant algorithm of the inverse distance weighting strategy is introduced. The environmental adaptability vector generated in step S13 is used as the basic trust benchmark, and dynamically adjusted in conjunction with numerical deviation characteristics. If the numerical deviation characteristic is small (e.g., less than 0.2), the weight coefficients of each data source mainly depend on the trust score in its environmental adaptability vector; if the numerical deviation characteristic is large (significant contradiction), a nonlinear penalty is applied to data sources with lower environmental adaptability scores. This is achieved through a weighted average formula. Calculate the weighting coefficients for each data source. Among them, This represents the score of the corresponding data source in the environmental adaptability vector. Numerical deviation characteristics, The conflict sensitivity factor is set to 0.5. This factor is based on sensitivity analysis of historical datasets containing labeled fault tags. Specifically, it is set within a preset interval (e.g., ...). The system iterates through the values of the factor and calculates the false alarm rate (i.e., the severe oscillation of weights due to occasional noise) and the false alarm rate (i.e., the failure to identify real sensor faults) under different values. The value that minimizes the weighted sum of the two is selected as the final setpoint. Finally, for the correction and matrix construction in step S143, a preset allocation threshold (e.g., 0.6) is set to prevent a single source from becoming overly dominant. This threshold is determined based on statistical analysis of historical fusion errors using the "inflection point method." A curve showing the relationship between "single-source weight ratio" and "fusion result accuracy" is plotted, and the inflection point where the accuracy gain begins to significantly slow down and tends to saturate is selected as the threshold. Exceeding this point means that further increasing the weight will lead to overfitting. If the weight coefficient of a certain data source exceeds this threshold, an attenuation coefficient is calculated using interference factors (such as strong light or heavy rain) in the environmental variable data to correct the weight. It is worth noting that this attenuation coefficient is not a fixed value but is dynamically calculated based on a signal-to-noise ratio attenuation model. The formula is used... Calculate the attenuation coefficient, where The normalized intensity of the current environmental disturbance factors. This is a dimensionless sensitivity constant based on the physical characteristics of the sensor (e.g., for a vision sensor, it is set according to the photoelectric conversion curve of its photosensitive element). Finally, a diagonal weight matrix is constructed using the corrected weight set. The independence assumptions between information sources are explicitly stated in matrix form, thus obtaining the final fusion weight matrix.
[0039] In step S15, the dominant weight vector is extracted from the fusion weight matrix, and multi-source evidence fusion calculation is performed on the multi-source input dataset using the weighted evidence combination rule to obtain a comprehensive driving status index, including: S151, Perform a numerical scan on the fusion weight matrix, identify and select the continuous subsequence with the largest average value among the diagonal elements of the fusion weight matrix, and generate the dominant weight vector; S152, using the dominant weight vector, the velocity deviation value and flow density value in the multi-source input dataset are weighted and filtered to determine preliminary weighted evidence; S153, the preliminary weighted evidence is fused with the environmental variable data to obtain the evidence strength. If the evidence strength exceeds a preset strength threshold, the acceleration features of the real-time speed data and the lane occupancy rate features of the traffic flow information are extracted, and the preliminary weighted evidence is expanded to obtain expanded weighted evidence. S154, apply the weighted evidence combination rule to perform fusion operation on the extended weighted evidence to obtain the comprehensive driving status index.
[0040] Specifically, for step S151, in order to identify the most critical information source set from the diagonal weight matrix containing a large amount of redundant information, a sliding window scanning algorithm is employed. A window of fixed length (e.g., window size) is set. The algorithm performs a sliding scan along the diagonal of the matrix with a step size of 1, calculating the arithmetic mean of the weighted elements within the window. This step size is based on the principle of "maximizing window coverage," ensuring that all consecutive groups of elements are scanned. The group of elements corresponding to the window position with the largest average value is selected as the "consecutive high-value group," and this group is extracted to generate the dominant weight vector. It is worth elaborating on the window size (e.g.) The setting of ) is based on the consensus protocol of distributed systems, that is, at least 3 independent sources of information are required to form the smallest effective voting unit, thereby avoiding the misleading of a single high-weight random noise.
[0041] Subsequently, in step S152, to construct a computable basic evidence body, feature vector alignment construction is first performed. The speed deviation values (from the vehicle-mounted end) and traffic density values (from the roadside end) in the multi-source input dataset are encapsulated into a single dimension (i.e., ...) according to the source order corresponding to the dominant weight vector. (Dimensional) Basic Feature Vector For example, if the weight vectors correspond in the order of [vehicle-mounted, roadside], then the feature vectors are constructed as follows: Next, the Hadamard product operation is performed, i.e. The weights are applied element-wise to the feature values. During this process, a screening threshold is set (e.g., 0.6, based on statistical significance). (Confirmed). Only elements in the resulting vector whose magnitude exceeds this threshold are retained and normalized, thus being identified as preliminary weighted evidence.
[0042] Regarding the evidence expansion in step S153, a weighted linear fusion algorithm is used to calculate the evidence strength in order to verify the robustness of the preliminary evidence in the physical environment. The specific formula is as follows: ; in, The L2 norm of preliminary weighted evidence; The normalized road surface friction coefficient; Normalized weather visibility; The fusion weighting coefficients are set (e.g., 0.5, 0.3, 0.2 respectively). These weighting coefficients are determined based on the Analytic Hierarchy Process (AHP), which calculates that direct evidence (the data itself) is slightly more important than indirect evidence (environmental factors) by constructing a judgment matrix.
[0043] It should be noted that the initial weighted evidence vector needs to be adjusted during the calculation. Calculate its relative norm (e.g., the ratio relative to the historical maximum norm) and the road surface friction coefficient. and weather visibility Normalization is performed to ensure that all feature values are dimensionless and have similar scales, and then a weighted sum is performed to calculate the strength of evidence.
[0044] Finally, in step S154, in order to synthesize the multidimensional conflicting evidence into a single decision indicator, the Dempster synthesis rule is applied. The orthogonal sum of each evidence body (i.e., each component in the extended weighted evidence) is calculated using the formula... Calculate the conflict coefficient K. Wherein, and Each represents a subset of the hypothesis in the identification framework, representing two independent bodies of evidence (e.g., different feature components in extended weighted evidence). and The basic probability allocation value; and To identify focal elements in the frame; conditions Represents a subset of hypotheses With the hypothesis subset The intersection of the two hypotheses is an empty set, meaning they are logically mutually exclusive (e.g., one piece of evidence points to "congestion," while the other points to "smooth traffic"). K is the conflict coefficient, which is equal to the sum of the probability products of all mutually exclusive hypotheses, used to quantify the overall degree of conflict between multiple sources of evidence. Then, normalization is performed using 1 / (1-K), ultimately outputting a scalar value in the interval [0,1] as a comprehensive driving status index. The closer this index value is to 1, the higher the certainty of the current driving status (e.g., congestion).
[0045] In step S16, the comprehensive driving state index is input into a preset state evolution Bayesian network, and state evolution inference and iterative calculation are performed in combination with the environmental variable data and the real-time speed data to obtain a refined state prediction sequence, including: S161, map the comprehensive driving status index to the input layer of the preset state evolution Bayesian network, initialize the network state nodes, and determine the initial probability distribution; S162, Combining the acceleration characteristics of the environmental variable data and the real-time speed data, perform vector operations on the initial probability distribution to obtain the intermediate optimized probability; S163, if the information entropy of the intermediate optimization probability exceeds the preset environmental variability threshold, then an extended probability distribution is generated by combining the lane occupancy characteristics of the traffic flow information and the braking response characteristics of the real-time speed data. S164, using the extended probability distribution, time slice iteration is performed in the state evolution Bayesian network to deduce a refined state prediction sequence.
[0046] It should be noted that the state evolution Bayesian network constructed in this step preferably adopts a dynamic Bayesian network (DBN) architecture, whose topology includes logically coupled observation layers, hidden state layers, and control layers. The hidden state layer is defined as having dimensions of... (For example The system consists of a discrete state vector sequence, used to characterize the hierarchical state of traffic flow from "smooth" to "congested"; the control layer consists of a discrete state vector sequence. For example The environmental feature vectors constitute the structure. The evolutionary relationships between nodes at each layer are described through matrix operations, utilizing a... A state transition probability matrix of dimension represents the natural evolution probability of the hidden states at adjacent time steps, and utilizes a... The environmental impact weight matrix quantifies the corrective weights of control variables such as road surface friction and visibility on the state distribution, thereby enabling accurate prediction of nonlinear traffic flow trends.
[0047] Specifically, for step S161, in order to capture the dynamic evolution of traffic flow over time, a dynamic Bayesian network is introduced as the core evolutionary model. This model contains a series of state nodes arranged in time slices. First, a state space discretization mapping is performed, transforming the continuous scalar form of the comprehensive driving state index (within a certain range) output from step S15 into a single value. ) mapped to 3D discrete state vector (e.g.) These correspond to: smooth traffic, mostly smooth traffic, light congestion, moderate congestion, and severe congestion, respectively. It is worth explaining in detail that the mapping rule is based on the service level grading standard in the "Urban Traffic Operation Evaluation Specification": [The text then lists various traffic levels and their corresponding levels, which are not directly related to the preceding sentences and are omitted from the translation.] The interval is divided into five subintervals that are left-closed and right-open (i.e., Based on the interval where the index value falls, an initial probability distribution vector is generated. For example, if the index value is 0.75 (falling into the fourth interval), then the initial vector can be initialized as follows: This reflects the distribution characteristics dominated by "moderate congestion".
[0048] Subsequently, in step S162, to incorporate the influence of environmental physical constraints on the state transition, a correction operation based on vector dot product is performed. First, a... Environmental feature vector Its elements include: normalized road friction coefficient, normalized weather visibility factor, and normalized vehicle acceleration characteristics. Then, using the formula... Calculate the intermediate optimization probability. It is important to define this explicitly. For one This is a 5x3 (5 rows, 3 columns) environmental impact weight matrix. Each row of the matrix corresponds to a traffic state, and each column corresponds to an environmental factor. The matrix element values represent the inducing weight of a specific environmental factor on a specific traffic state (for example, the element values at the intersection of "slippery road surface" and "congestion" states are usually large positive values). This matrix is obtained by training it through multivariate regression analysis on historical multi-source traffic data.
[0049] For the extended processing in step S163, it is necessary to determine whether the current state has high uncertainty. The information entropy of the intermediate optimization probability is then calculated. ,like If the threshold for a variable environment is exceeded (e.g., 0.8 bits), a "variable environment signal" is displayed. This threshold is based on a rigorous information theory interpretation: in a binary or multivariate distribution, entropy increases when the probability of the dominant state decreases and its probability quality disperses to other states. A 0.8-bit threshold corresponds to a dominant state probability decreasing to approximately 0.8 (with other probabilities dispersed), indicating that the system's judgment of the current state has become significantly ambiguous (not absolutely certain), thus requiring the introduction of additional evidence. At this point, additional observation nodes are activated: the lane occupancy rate of the traffic flow information and the braking response time of the real-time speed data are extracted. Using a Bayesian update rule, these two features are input into the network as new evidence, and the posterior probability is calculated to generate an expanded probability distribution. This eliminates uncertainty.
[0050] Finally, in step S164, an evolutionary deduction along the time dimension is performed. This is done using the extended probability distribution. As of the present moment The state confidence is combined with the state transition probability matrix in the dynamic Bayesian network to execute the forward inference algorithm. Future moments are calculated slice by slice according to a preset time step (e.g., 5 minutes). The state probability distribution is used to form a refined state prediction sequence. Here, the prediction step size (e.g., 5 minutes) is set based on the "hysteresis loop" period in traffic flow theory. Physics research shows that when traffic flow transitions from free flow to congested flow (or vice versa), the changes in density and speed are not synchronous, but rather there is a relaxation time based on driver reaction delay and the transmission of fleet fluctuations. This time window has a statistical average of about 5 minutes in urban arterial road scenarios. Therefore, selecting this step size ensures that the model fully captures the dynamic process of traffic flow state phase transition, avoiding noise sensitivity due to too short a step size or information loss due to too long a step size.
[0051] It should be noted that the length of the predicted sequence (i.e., the number of iterations) It can be pre-configured according to the management needs of the actual application scenario. For example, when providing short-term early warnings for traffic management centers, it can be set... =6, which means projecting the state sequence for the next 30 minutes (in 5-minute increments); this can be set when providing a reference for path planning. =12, which means predicting the state sequence for the next hour; in addition, a dynamic termination condition can be set, such as stopping the iteration when the entropy value of the state probability distribution of three consecutive time slices is lower than the stability threshold, and using the sequence generated at this time as the final refined state prediction sequence.
[0052] In step S17, the refined state prediction sequence is subjected to temporal smoothing filtering, and the final driving state classification result is determined by matching the filtered sequence features with a preset decision threshold, including: S171, Perform time-series smoothing and low-pass filtering on the refined state prediction sequence to obtain a stable smooth sequence; S172, calculate the vehicle speed variability based on the stable smooth sequence. If the vehicle speed variability exceeds a preset variability threshold, calculate the congestion probability value and obtain the classification probability distribution. S173, compare the classification probability distribution with the preset decision threshold, and determine the final driving state classification result including congestion state based on the comparison result.
[0053] Specifically, for step S171, in order to eliminate high-frequency noise in the prediction sequence caused by occasional traffic events (such as temporary parking, lane change interference), a sliding window moving average algorithm is first used. A value covering a specific time span (e.g., ...) is set. A convolution operation is performed on the sequence data within a time window of 5 minutes to obtain a preliminary smoothed sequence. It is worth noting that the window size (e.g., 5 minutes) is set based on the statistical average of urban traffic signal control cycles (typically multiples of 90-120 seconds), aiming to smooth out the periodic fluctuations caused by traffic light starts and stops. Subsequently, to further eliminate minor oscillations caused by the environment, a second-order Butterworth low-pass filter is designed and applied, incorporating the road surface friction coefficient data. Preferably, the cutoff frequency of this filter is not a fixed value but is dynamically adjusted based on the road surface friction coefficient. When the friction coefficient is low (slippery road surface), the micro-operational fluctuations of vehicles increase, and the system automatically lowers the cutoff frequency (e.g., from 0.5Hz to 0.2Hz) to enhance the suppression of high-frequency noise, thereby obtaining a more stable and smoothed sequence that reflects macroscopic traffic trends.
[0054] In step S172, to quantify the dispersion of traffic flow and identify potential congestion precursors, the vehicle speed variability of the stable smooth sequence is calculated. Specifically, the coefficient of variation formula is used for calculation: ,in The standard deviation of the sequence is . The variability is the mean. If the calculated variability exceeds a preset variability threshold (e.g., 0.15), it indicates that the traffic flow has shifted from a stable flow to an unstable flow. The variability threshold (e.g., 15%) is set based on the "collapse probability model" in traffic flow theory. Empirical data shows that when the speed variability coefficient exceeds 15%, the probability of traffic flow collapse leading to congestion increases exponentially. Under this triggering condition, a logistic regression model or a preset mapping function is used to convert the variability value into... The congestion probability value of the interval is used to generate a classification probability distribution.
[0055] Finally, in step S173, the final binary or multivariate classification decision is performed. The congestion probability value in the generated classification probability distribution is compared with a preset decision threshold (e.g., 0.65). This decision threshold is set based on various cost-benefit analyses of receiver operating characteristic curves. Considering that misclassifying normal traffic as congestion (false alarm) would cause unnecessary waste of resource scheduling, while omitting congestion (false report) would lead to severe traffic paralysis, the threshold is set to 0.65 (slightly higher than 0.5), reflecting a moderate control of the false alarm rate while ensuring recall. If the congestion probability exceeds this threshold, the current driving state is determined to be "congested"; otherwise, it is determined to be "normal" or "slightly slow," thus outputting the final driving state classification result.
[0056] In summary, this invention, based on the precise spatiotemporal synchronization of multi-source heterogeneous data and the source analysis of evidence theory, combined with an environmentally adaptive Bayesian trust assessment model and a dynamic state evolution inference mechanism, ensures that the performance goals of high reliability of multi-source perception, high accuracy of vehicle state recognition, and strong robustness of traffic trend prediction are continuously met under complex conditions such as sensor data logical conflicts, severe weather interference, and nonlinear evolution of traffic flow. This is achieved through dynamic weight allocation based on numerical deviation characteristics, a deep fusion strategy of multi-source evidence, and a decision optimization mechanism based on temporal smoothing filtering.
[0057] Reference Figure 2 The second embodiment of the present invention provides a vehicle status intelligent analysis system based on multi-source heterogeneous data fusion, comprising: The multi-source data processing module acquires real-time speed data from vehicle-mounted sensors and traffic flow information from roadside equipment, and performs timestamp alignment and field mapping on the real-time speed data and traffic flow information to obtain a multi-source input dataset. The source-end trust assessment module retrieves historical calibration records corresponding to the vehicle-mounted sensors and the roadside equipment, performs evidence theory calculations on the historical calibration records, constructs a confidence function and a likelihood function, and calculates the source-end trust score. The environmental adaptability inference module acquires real-time environmental variable data, inputs the environmental variable data and the source trust score into a preset trust assessment Bayesian model, performs conditional probability inference, and obtains an environmental adaptability vector. The dynamic weight allocation module calculates the numerical deviation characteristics between the real-time speed data and the traffic flow information, and performs dynamic weight allocation in conjunction with the environmental adaptability vector to obtain a fusion weight matrix. The multi-source evidence fusion module extracts the dominant weight vector from the fusion weight matrix and performs multi-source evidence fusion calculation on the multi-source input dataset using weighted evidence combination rules to obtain a comprehensive driving status index. The state evolution prediction module inputs the comprehensive driving state index into a preset state evolution Bayesian network, and combines the environmental variable data and the real-time speed data to perform state evolution inference and iterative calculation to obtain a refined state prediction sequence. The state classification decision module performs time-series smoothing filtering on the refined state prediction sequence and determines the final driving state classification result by matching the filtered sequence features with a preset decision threshold.
[0058] It should be noted that the intelligent driving status analysis system based on multi-source heterogeneous data fusion provided in this embodiment of the invention is used to execute all the process steps of the intelligent driving status analysis method based on multi-source heterogeneous data fusion in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.
[0059] This invention also provides an electronic device. The electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a state evolution prediction program. When the processor executes the computer program, it implements the steps described in the various embodiments of the intelligent vehicle state analysis method based on multi-source heterogeneous data fusion, for example... Figure 1 The step S11 shown. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above system embodiments, such as the state classification decision module.
[0060] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device.
[0061] The electronic device may be a desktop computer, laptop, handheld computer, or smart tablet, etc. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the above components are merely examples of electronic devices and do not constitute a limitation on the electronic device. It may include more or fewer components than described above, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.
[0062] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the electronic device, connecting all parts of the electronic device via various interfaces and lines.
[0063] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0064] Wherein, if the modules / units integrated in the electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0065] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the system embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0066] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A method for intelligent analysis of driving status based on multi-source heterogeneous data fusion, characterized in that, include: Real-time speed data from vehicle-mounted sensors and traffic flow information from roadside equipment are acquired respectively. The real-time speed data and traffic flow information are then time-stamped and mapped to obtain a multi-source input dataset. Retrieve historical calibration records corresponding to the vehicle-mounted sensors and the roadside equipment, perform evidence theory calculations on the historical calibration records, construct confidence functions and likelihood functions, and calculate the source confidence score. Real-time environmental variable data is acquired, and the environmental variable data and the source trust score are input into a preset trust assessment Bayesian model to perform conditional probability inference and obtain an environmental adaptability vector. The numerical deviation characteristics between the real-time speed data and the traffic flow information are calculated, and dynamic weight allocation is performed in combination with the environmental adaptability vector to obtain a fusion weight matrix; The dominant weight vector is extracted from the fusion weight matrix, and the multi-source evidence fusion calculation is performed on the multi-source input dataset using the weighted evidence combination rule to obtain the comprehensive driving status index. The comprehensive driving status index is input into a preset state evolution Bayesian network, and the state evolution is deduced and iteratively calculated by combining the environmental variable data and the real-time speed data to obtain a refined state prediction sequence. The refined state prediction sequence is subjected to temporal smoothing filtering, and the final driving state classification result is determined by matching the filtered sequence features with a preset decision threshold.
2. The intelligent vehicle status analysis method based on multi-source heterogeneous data fusion as described in claim 1, characterized in that, The process of retrieving historical calibration records corresponding to the vehicle-mounted sensors and the roadside equipment, performing evidence theory calculations on the historical calibration records, constructing confidence functions and likelihood functions, and calculating the source-end confidence score includes: Based on the device identification information of the vehicle-mounted sensors and the roadside equipment, the matching historical calibration records are retrieved and extracted, and the historical calibration records are classified into sensor calibration records and device calibration records. Basic probability allocation calculations are performed on the sensor calibration records and the device calibration records respectively to obtain the source confidence function value and the source likelihood function value; If the source confidence function value is greater than the preset confidence threshold, the calculation weight of the source likelihood function value is increased to generate a weighted likelihood function value. The source confidence function value and the weighted likelihood function value are averaged and fused to obtain the source confidence score.
3. The intelligent vehicle status analysis method based on multi-source heterogeneous data fusion as described in claim 1, characterized in that, The process of acquiring real-time environmental variable data, inputting the environmental variable data and the source-end trust score into a preset trust assessment Bayesian model, performing conditional probability inference, and obtaining an environmental adaptability vector includes: The system acquires the current weather and road condition signals, performs feature analysis and quantification, and extracts weather condition factors and road condition factors as the environmental variable data. Using the weather condition factor and the road surface condition factor as condition nodes, calculate the conditional probability distribution of the source trust score under different environmental conditions; Based on the conditional probability distribution, determine the probability node values of the vehicle-mounted sensor and the roadside equipment; If the probability node value exceeds the preset reliability threshold, the probability node value and the source trust score are weighted and fused to generate an adjusted trust score. Based on the adjusted trust scores, an environmental adaptability vector is generated.
4. The intelligent vehicle status analysis method based on multi-source heterogeneous data fusion as described in claim 1, characterized in that, The numerical deviation characteristics between the real-time speed data and the traffic flow information are calculated, and dynamic weight allocation is performed in conjunction with the environmental adaptability vector to obtain a fusion weight matrix, including: Extract the speed value from the real-time speed data and the flow value from the traffic flow information, and perform a normalized difference calculation on the speed value and the flow value to obtain the numerical deviation characteristics; Based on the numerical deviation characteristics and the environmental adaptability vector, the weight coefficients of each data source are calculated using a weighted average method. If the weight coefficient is greater than the preset allocation threshold, the weight coefficient is corrected using the environmental variable data to obtain a corrected weight set, and the weight set is then matrix-combined to obtain a fused weight matrix.
5. The intelligent vehicle status analysis method based on multi-source heterogeneous data fusion as described in claim 1, characterized in that, The process of extracting the dominant weight vector from the fusion weight matrix and performing multi-source evidence fusion calculation on the multi-source input dataset using weighted evidence combination rules to obtain a comprehensive driving status index includes: The fusion weight matrix is numerically scanned to identify and select the continuous subsequence with the largest average value among the diagonal elements of the fusion weight matrix, and a dominant weight vector is generated. Using the dominant weight vector, the velocity deviation value and flow density value in the multi-source input dataset are weighted and filtered to determine preliminary weighted evidence; The preliminary weighted evidence is fused with the environmental variable data to obtain the evidence strength. If the evidence strength exceeds a preset strength threshold, the acceleration features of the real-time speed data and the lane occupancy rate features of the traffic flow information are extracted, and the preliminary weighted evidence is expanded to obtain expanded weighted evidence. The extended weighted evidence is fused using weighted evidence combination rules to obtain a comprehensive driving status index.
6. The intelligent vehicle status analysis method based on multi-source heterogeneous data fusion as described in claim 1, characterized in that, The process involves inputting the comprehensive driving state index into a preset state evolution Bayesian network, combining the environmental variable data and the real-time speed data to perform state evolution deduction and iterative calculation, and obtaining a refined state prediction sequence, including: The comprehensive driving status index is mapped to the input layer of a preset state evolution Bayesian network, and the network state nodes are initialized to determine the initial probability distribution. By combining the acceleration characteristics of the environmental variable data and the real-time speed data, vector operations are performed on the initial probability distribution to obtain the intermediate optimized probability; If the information entropy of the intermediate optimization probability exceeds the preset environmental variability threshold, then an extended probability distribution is generated by combining the lane occupancy characteristics of the traffic flow information and the braking response characteristics of the real-time speed data. Using the extended probability distribution, time-slice iterations are performed in the state evolution Bayesian network to deduce a refined state prediction sequence.
7. The intelligent vehicle status analysis method based on multi-source heterogeneous data fusion as described in claim 1, characterized in that, The step of performing time-series smoothing filtering on the refined state prediction sequence and determining the final driving state classification result by matching a preset decision threshold with the filtered sequence features includes: Perform time-series smoothing and low-pass filtering on the refined state prediction sequence to obtain a stable smooth sequence; The vehicle speed variability is calculated based on the stable and smooth sequence. If the vehicle speed variability exceeds a preset variability threshold, the congestion probability value is calculated to obtain the classification probability distribution. The classification probability distribution is compared with a preset decision threshold, and the final driving state classification result, including congestion state, is determined based on the comparison result.
8. A vehicle status intelligent analysis system based on multi-source heterogeneous data fusion, characterized in that, include: The multi-source data processing module acquires real-time speed data from vehicle-mounted sensors and traffic flow information from roadside equipment, and performs timestamp alignment and field mapping on the real-time speed data and traffic flow information to obtain a multi-source input dataset. The source-end trust assessment module retrieves historical calibration records corresponding to the vehicle-mounted sensors and the roadside equipment, performs evidence theory calculations on the historical calibration records, constructs a confidence function and a likelihood function, and calculates the source-end trust score. The environmental adaptability inference module acquires real-time environmental variable data, inputs the environmental variable data and the source trust score into a preset trust assessment Bayesian model, performs conditional probability inference, and obtains an environmental adaptability vector. The dynamic weight allocation module calculates the numerical deviation characteristics between the real-time speed data and the traffic flow information, and performs dynamic weight allocation in conjunction with the environmental adaptability vector to obtain a fusion weight matrix. The multi-source evidence fusion module extracts the dominant weight vector from the fusion weight matrix and performs multi-source evidence fusion calculation on the multi-source input dataset using weighted evidence combination rules to obtain a comprehensive driving status index. The state evolution prediction module inputs the comprehensive driving state index into a preset state evolution Bayesian network, and combines the environmental variable data and the real-time speed data to perform state evolution inference and iterative calculation to obtain a refined state prediction sequence. The state classification decision module performs time-series smoothing filtering on the refined state prediction sequence and determines the final driving state classification result by matching the filtered sequence features with a preset decision threshold.