Traffic Profile Clustering via Chi-Squared Significance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing traffic prediction methods are inefficient and resource-intensive, often failing to accurately determine cluster labels due to the lack of corresponding traffic profiles, leading to incomplete representation of traffic behavior patterns.
Innovation Solution
A method that uses statistical analysis, specifically the Chi-squared statistic, to calculate significance values for tags and tag combinations within clusters, allowing for efficient and systematic determination of characteristic vectors, which are then used to predict traffic behavior by clustering and labeling traffic profiles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If randomised algorithms are used to find matching labels for clusters, then labels can be assigned automatically, but the method consumes a lot of computational resources and fails to find accurate labels for clusters with few corresponding traffic profiles
Solution Approach 1:
The patent applies preliminary action by pre-calculating significance values for all possible tags using statistical analysis (Chi-squared statistic) before the actual label assignment process. This pre-computation creates a ready-to-use significance database that speeds up the automated labeling process and ensures accurate label selection even for clusters with few traffic profiles, as the significance values are already determined in advance.
Solution Approach 2:
The patent replaces the randomised algorithmic approach with a deterministic statistical method. Instead of using randomised searching to find matching labels, the system uses Chi-squared statistical analysis to calculate significance values objectively. This substitution eliminates the randomness and computational inefficiency while improving label accuracy through mathematically rigorous significance testing.
2Reliability
If manual labelling of clusters is performed, then accurate labels can be assigned, but the process is time-consuming and resource-intensive
Solution Approach 1:
The patent implements self-service by enabling the system to automatically determine characteristic vectors and assign labels without human intervention. The statistical analysis method allows the system to autonomously evaluate tags, calculate significance values, and select the most representative tags for each cluster, eliminating the need for manual labelling while maintaining high accuracy through objective statistical criteria.
Solution Approach 2:
The patent substitutes manual human labor with an automated statistical computing system. The Chi-squared statistic calculation and significance value determination replace the manual evaluation and selection process, dramatically increasing productivity while preserving label accuracy through mathematically sound automated decision-making.
3Adaptability or versatility
If standard clustering algorithms are used to group traffic profiles, then traffic patterns can be analyzed, but the method cannot reliably identify representative tags for clusters with few traffic profiles
Solution Approach 1:
The patent applies parameter changes by introducing significance values as a new evaluation parameter for tag selection. Instead of relying solely on the number of traffic profiles in a cluster, the system calculates statistical significance values for each tag, allowing accurate identification of representative tags even when cluster size is small. This parameter transformation enables reliable tag selection across clusters of varying sizes.
Solution Approach 2:
The patent introduces significance values as an intermediary metric between traffic profile clustering and tag selection. This intermediary statistical measure bridges the gap between cluster formation and representative tag identification, providing an objective criterion for evaluating tag representativeness regardless of cluster size, thus improving measurement precision in tag representation.
Data Source
AI summary
The present subject matter relates to a method of predicting a traffic behaviour in a road system, comprising the following steps carried out by at least one processor connected to a database, which contains a set of traffic profiles: clustering the traffic profiles; calculating a significance value; if the significance value of a tag is higher than a threshold value, assigning this tag as a characteristic vector to this cluster; receiving a request having one or more request tags; determining a cluster by means of matching the request tag/s to the characteristic vectors of the clusters; and outputting a predicted traffic behaviour based on the determined cluster.


