Traffic network congestion identification method based on data fusion
By screening and reliability calculation of traffic data and obtaining adaptive weights, the traditional traffic data fusion method has solved the accuracy and rationality of the accuracy and rationality of traditional traffic data fusion methods, and achieved more accurate vehicle speed acquisition and traffic network congestion identification.
Patent Information
- Application Number
- CN202510299983.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-03-14
AI Technical Summary
The traditional traffic data fusion method obtains the vehicle's passing speed through the weighted average method. The weight is usually given based on empirical values, resulting in insufficient accuracy and rationality of the data, making it difficult to accurately reflect the true status of the vehicle's passing speed, thereby reducing the accuracy of the recognition of congestion periods.
By collecting the speed data of the vehicle when driving on the road section, filtering out the sequence of focus, calculating the credibility and credibility of the trough and credibility of the credibility and credibility of the trough points and credibility of the trough points, calculating the credibility and confidence of the sequence, obtaining the vehicle's passing speed based on the adaptive weight, and then identifying congestion.
It improves the accuracy of vehicle passing speed, accurately identify congestion in the traffic network, and enhances the effectiveness of mitigation measures to mitigate traffic pressure.
Smart Images

Figure CN119811096B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic congestion identification, and in particular to a traffic network congestion identification method based on data fusion. Background Art
[0002] With the continuous advancement of urbanization, traffic problems have become one of the core issues to be solved in modern urban development. Traffic congestion not only leads to a decline in traffic efficiency, but also brings air pollution, energy waste and inconvenience to people's daily travel. Therefore, how to fuse multiple sets of collected data into accurate data, identify congestion in the traffic network through it, and take effective measures to alleviate traffic pressure has become a research hotspot in the field of traffic management.
[0003] Traditional traffic data fusion uses the weighted average method to obtain the speed of vehicles passing through each road. The weight of each data source is usually given based on empirical values, and the weight of each data source in each time period is fixed. There are limitations in the accuracy and rationality of the data, and it is difficult to accurately reflect the actual state of the speed of vehicles passing through each time period, resulting in low accuracy in the subsequent identification of congested time periods. Therefore, when fusing the collected traffic data, it is necessary to obtain the corresponding weights based on the credibility and confidence of each set of data to make the fused data more accurate, thereby improving the accuracy of congestion identification. Summary of the invention
[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a traffic network congestion identification method based on data fusion, and the technical solution adopted is as follows:
[0005] An embodiment of the present invention provides a method for identifying traffic network congestion based on data fusion, the method comprising:
[0006] Collect the speed data of vehicles traveling on a road section within a period of time, obtain the speed sequences of the period of time, and obtain the ordinate and abscissa of each data point;
[0007] The velocity sequence is screened to obtain the focused analysis sequence; a window is established with the trough point and the peak point in the focused analysis sequence as the center, and the trough credibility and the peak credibility are calculated according to the data in the window, and the credible trough point and the credible peak point are obtained;
[0008] The data values in the window corresponding to the credible trough point and the credible peak point are averaged to obtain the sampling mean; the credibility of the sequence is analyzed based on the sampling mean, trough point and peak point calculation;
[0009] A window is established with a data point in the sequence being analyzed as the starting point. The stationarity of the data point is calculated based on the ordinate and abscissa of each data point in the window and the number of data points. The turning point probability of the data point is calculated based on the slope between the data point and the next data point. The data points are screened to obtain key points based on the stationarity and turning point probability of the data points.
[0010] Calculate the similarity of key points between each two focused analysis sequences, and select the two focused analysis sequences with the largest similarity as the control sequence pair; calculate the stability of each sequence in the control sequence pair, and select the sequence with the largest stability as the control sequence;
[0011] The similarity between each focused analysis sequence and the control sequence is calculated and recorded as confidence; the adaptive weight is obtained according to the confidence and credibility of the focused analysis sequence; the vehicle passing speed in the period is obtained according to the adaptive weight and congestion identification is performed.
[0012] Preferably, the speed data of a vehicle traveling on a road section within a period of time is collected to obtain each speed sequence of the period of time, and the ordinate and abscissa of each data point are obtained, including:
[0013] Install a preset number of traffic flow detectors on a road section, use a traffic flow detector to collect vehicle speed data during the period, and form a speed sequence corresponding to the traffic flow detector;
[0014] Get the average of all speed data uploaded at fixed intervals when a vehicle passes through the road section within a period of time, recorded as the speed average. The speed average of all vehicles passing through the road section within a period of time constitutes the speed sequence corresponding to the floating vehicle; the timestamp of each data point in the speed sequence corresponding to the floating vehicle is the time when each vehicle enters the road section; obtain each speed sequence corresponding to the period of time;
[0015] Map each data in the speed sequence corresponding to the traffic flow detector to a two-dimensional rectangular coordinate system, where the ordinate and abscissa of each data point are the speed value and the acquisition time respectively;
[0016] Each data point in the speed sequence corresponding to the floating vehicle is mapped to a two-dimensional rectangular coordinate system, and the ordinate and abscissa of each data point are the mean speed and the moment when the vehicle enters the road section, respectively.
[0017] Preferably, the speed sequence is screened to obtain a focused analysis sequence, including:
[0018] The speed range is set according to the maximum speed limit of the road section, and the number of data values of data points in a speed sequence within the speed range is obtained, which is recorded as the number of normal passing vehicles; the number of data points in the speed sequence with a preset multiple is used as the number threshold, if the number of normal passing vehicles corresponding to each speed sequence in a time period is greater than or equal to the number threshold, then each speed sequence in this time period is a regular analysis sequence; if the number of normal passing vehicles corresponding to each speed sequence in a time period is not greater than or equal to the number threshold, then each speed sequence in this time period is a focused analysis sequence.
[0019] Preferably, the trough credibility and the peak credibility are calculated respectively according to the data in the window, and the credible trough point and the credible peak point are obtained, including:
[0020] With a trough point as the center, a window of a preset time length is established to obtain the number of data points in the window; the horizontal coordinate difference between the latter data point and the previous data point in every two adjacent data points in the window is calculated, recorded as the time difference; the number of data points in the window is multiplied by the inverse of the standard deviation of the corresponding time difference in the window and normalized to obtain the trough credibility of the trough point;
[0021] With a peak point as the center, a window of a preset time length is established to obtain the number of data points in the window; the horizontal coordinate difference between the latter data point and the previous data point in every two adjacent data points in the window is calculated, recorded as the time difference; the reciprocal of the number of data points in the window is multiplied by the reciprocal of the standard deviation of the corresponding time difference in the window and normalized to obtain the peak credibility of the peak point;
[0022] A credible threshold is set. If the peak credibility of a peak point is greater than the threshold, the peak point is a credible peak point. If the trough credibility of a trough point is greater than the threshold, the trough point is a credible trough point.
[0023] Preferably, the credibility of the sequence is analyzed based on the sampling mean, trough points and peak points, including:
[0024] The absolute value of the difference between the sampling mean and the mean of the focused analysis sequence is inverted and multiplied by the adjustment factor to obtain the mean difference term; the mean of the peak credibility of all peak points of the focused analysis sequence and the mean of the trough credibility of all trough points are added to obtain the peak and trough distribution credibility term; the value difference term is multiplied by the peak and trough distribution credibility term to obtain the credibility of the focused analysis sequence.
[0025] Preferably, a window is established with a data point in the focused analysis sequence as the starting point, the stationarity of the data point is calculated according to the ordinate and abscissa of each data point in the window and the number of data points, and the turning point probability of the data point is calculated according to the slope between the data point and the next data point, including:
[0026] Respectively obtain the standard deviation of the ordinate and the standard deviation of the abscissa of the data point in the window corresponding to a data point, and record them as the ordinate standard deviation and the abscissa standard deviation; multiply the ordinate standard deviation, the abscissa standard deviation, the inverse of the number of data points in the window and the adjustment factor to obtain a multiplication result; normalize the inverse of the multiplication result to obtain the stability of the data point;
[0027] Calculate the slope between one data point and the next data point, recorded as the instantaneous rate of change; obtain the standard deviation of the horizontal coordinates of the data points other than the data point in the window corresponding to the data point, recorded as the first horizontal coordinate standard deviation; multiply the instantaneous rate of change by the inverse of the first horizontal coordinate standard deviation and normalize them to obtain the turning point probability of the data point.
[0028] Preferably, the data points are screened to obtain key points according to the stability and turning point probability of the data points, including:
[0029] Set the screening threshold. If the stability or turning point probability of a data point in the sequence being analyzed is greater than the screening threshold, then the data point is a key point.
[0030] Preferably, the stationarity of each sequence in the reference sequence pair is calculated respectively, and the sequence with the greater stationarity is selected as the reference sequence, including:
[0031] A window is established with the first data point of a focused analysis sequence in the control sequence pair as the starting point, and the window is slid according to the set step size to obtain the mean of the standard deviation of the data points in each window during sliding, which is recorded as the stationarity of the focused analysis sequence; the focused analysis sequence with a larger stationarity in the control sequence pair is selected as the control sequence.
[0032] Preferably, the adaptive weight is obtained according to the confidence and credibility of the focused analysis sequence, including:
[0033] The confidence and credibility of a focused analysis sequence are multiplied as the numerator, and the confidence and credibility of each focused analysis sequence in the same period are multiplied and summed as the denominator. The ratio is the adaptive weight of the focused analysis sequence.
[0034] Preferably, obtaining the vehicle passing speed in the time period according to the adaptive weight and performing congestion identification includes:
[0035] The mean of the data in each focused analysis sequence in the same time period is weighted and summed based on the adaptive weights of each focused analysis sequence in the time period to obtain the vehicle passing speed of the time period; a speed threshold is set, and if the vehicle passing speed of the time period is lower than the speed threshold, the time period is marked as a congested time period.
[0036] The embodiments of the present invention have at least the following beneficial effects: the present invention obtains various speed sequences of a time period by collecting speed data of vehicles traveling on a road section within a time period, and the data sources of each speed sequence are different, thereby filtering out the focused analysis sequence, and subsequently only analyzing the focused analysis sequence to reduce the amount of calculation; further, analyzing the peak points and trough points in the focused analysis sequence to obtain the credibility of the focused analysis sequence; then obtaining the two focused analysis sequences with the greatest similarity as a reference sequence pair, and then using the sequence with the largest degree of stability as a reference sequence to obtain the confidence of each focused analysis sequence; then obtaining an adaptive weight by comprehensively considering the confidence and credibility of each focused analysis sequence, obtaining the vehicle passing speed of the time period according to the adaptive weight and performing congestion identification, thereby improving the accuracy of the obtained vehicle passing speed, and then accurately identifying the congestion situation of the traffic network. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0038] Figure 1 A method flow chart of a traffic network congestion identification method based on data fusion provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0039] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the specific implementation method, structure, features and effects of a traffic network congestion identification method based on data fusion proposed by the present invention in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures, or characteristics in one or more embodiments may be combined in any suitable form.
[0040] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0041] The following is a detailed description of a method for identifying traffic network congestion based on data fusion provided by the present invention in conjunction with the accompanying drawings.
[0042] Example:
[0043] The main application scenarios of the present invention are: the traditional traffic data fusion obtains the vehicle passing speed of each road through the weighted average method. The weight of each data source is usually given according to the empirical value, and the weight of each data source in each time period is fixed. There are limitations in the accuracy and rationality of the data, and it is difficult to accurately reflect the actual state of the vehicle passing speed in each time period, resulting in low accuracy in the subsequent congestion period identification. Therefore, it is necessary to correct the weight value when performing weighting, so as to obtain a more accurate vehicle passing speed.
[0044] See also Figure 1 , which shows a method flow chart of a method for identifying traffic network congestion based on data fusion provided by an embodiment of the present invention, the method comprising the following steps:
[0045] Step S1, collecting speed data of a vehicle traveling on a road section within a time period, obtaining each speed sequence of the time period, and obtaining the ordinate and abscissa of each data point.
[0046] Traffic flow detectors are devices used to monitor the flow of vehicles on the road (geomagnetic sensors, video surveillance, radar, laser sensors, ultrasonic sensors). Generally, in order to collect accurate data, at least two or more traffic flow detectors are installed on each road to collect the speed of each passing vehicle. Floating vehicles refer to the installation of GPS and other equipment on vehicles traveling on the road to collect real-time location and speed data of vehicles.
[0047] Traffic flow detectors are generally installed at fixed locations on each road to collect the speed of vehicles passing through the detectors within a specific time period. In the present invention, a preset number of traffic flow detectors are installed on a road section to collect speed data of each vehicle passing through the traffic flow detectors within a time period. The speed data of vehicles passing through each traffic flow detector within a time period form a speed sequence, which is the speed sequence corresponding to the traffic flow detector. The preset number of traffic flow detectors needs to be determined by the implementer based on actual conditions. The length of a time period is one hour, and the data collected every day can be divided into 24 speed sequences.
[0048] At the same time, floating vehicle data also needs to be collected. Floating vehicle data is all speed data uploaded at fixed intervals by vehicles passing through the section within a period of time. It also collects the vehicle's driving speed. Therefore, the average of all speed data uploaded at fixed intervals when a vehicle passes through the section within a period of time is obtained, which is recorded as the speed average. The speed average of all vehicles passing through the section within a period of time constitutes the speed sequence corresponding to the floating vehicle. The timestamp of each data point in the speed sequence corresponding to the floating vehicle is the moment when each vehicle enters the section.
[0049] Map each data point in the speed sequence corresponding to the traffic flow detector to a two-dimensional rectangular coordinate system, and the ordinate and abscissa of each data point are the speed value and the collection time respectively; map each data point in the speed sequence corresponding to the floating vehicle to a two-dimensional rectangular coordinate system, and the ordinate and abscissa of each data point are the mean speed value and the time when the vehicle enters the road section respectively.
[0050] In this way, all speed sequences corresponding to the road section within a period of time can be obtained for subsequent analysis.
[0051] It should be noted that when traffic flow detectors or floating vehicle data are collected, multiple data may be obtained at the same time. For example, the traffic flow detector collects the speeds of multiple vehicles at the same time, and multiple vehicles enter the road section at the same time. At this time, one timestamp may correspond to multiple collected data values. However, since the data points in the speed series are mapped to the coordinate system as a scatter plot, it will not affect the subsequent analysis.
[0052] Step S2, screening the speed sequence to obtain a focused analysis sequence; establishing a window with the trough point and the peak point in the focused analysis sequence as the center, calculating the trough credibility and the peak credibility respectively according to the data in the window, and obtaining a credible trough point and a credible peak point.
[0053] Each speed sequence records the speed of each vehicle within a certain period of time, and then takes its average as the output value of each speed sequence. Since each speed sequence is collected by sensors or on-board GPS and other devices, some data have low overall credibility and cannot reflect the average speed of the road. Therefore, it is necessary to obtain its credibility through the distribution characteristics of each speed sequence, and then obtain the confidence of each speed sequence based on the corresponding relationship between the data in each speed sequence, and combine the two to obtain the weight of each speed sequence.
[0054] For the speed sequence of each time period collected, we first need to determine whether the speed sequence of the time period needs to be calculated emphatically based on the amount of data. If the amount of data is large, the possibility of congestion is high and it needs to be calculated emphatically. Otherwise, the possibility of congestion is small and it does not need to be calculated emphatically.
[0055] Since the possible degree of congestion of the speed sequence needs to be determined based on the number of data in the speed sequence, the maximum speed limit of the road section needs to be obtained. Under unimpeded conditions, the speed of the vehicle is generally high and close to the maximum speed limit of the road section. The speed range is set according to the maximum speed limit of the road section, so the speed range is set to [0.7Vmax, 0.85Vmax], where Vmax is the maximum speed limit of the road section, and 0.7 and 0.85 are empirical values, which can be adjusted by the implementer according to the actual situation.
[0056] The number of data points in a speed sequence whose data values are within the vehicle speed range is obtained and recorded as the number of normal passing vehicles; a preset multiple of the number of data points in the speed sequence is used as the number threshold; if the number of normal passing vehicles corresponding to each speed sequence in a time period is greater than or equal to the number threshold, then each speed sequence in the time period is a conventional analysis sequence; if the number of normal passing vehicles corresponding to each speed sequence in a time period is not completely greater than or equal to the number threshold, then each speed sequence in the time period is a focused analysis sequence.
[0057] If the number of normal passing vehicles corresponding to each speed sequence in a time period is not completely greater than or equal to the quantity threshold, it means that there are many vehicles passing during the collection period and congestion may occur. Therefore, it is marked as a key calculation period, and each speed sequence in the period is a key analysis sequence. The credibility of the data in each key analysis sequence in the same period is continued to be calculated.
[0058] Since the vehicles passing through a road section are uncontrollable factors, and the driving conditions of vehicles will be affected by various factors and present different speeds, in some cases, vehicles will slow down due to some special circumstances, not congestion. Therefore, if there is a low speed in the data and the surrounding data points are distributed discretely, it means that the credibility of this group of data is low. When congestion occurs, the data generally shows low speed and dense data. When the traffic is unblocked, the data volume is small and the distribution is discrete. For low-speed data points, the more concentrated the distribution of the surrounding data is, the more credible the data is; for high-speed data points, the sparser the distribution of the surrounding data is, the more credible the data is.
[0059] For the focused analysis sequence, the peak points and trough points are obtained and marked. The marked peak points and trough points are matched to the original focused analysis sequence, where the peak points are high speed data points and the trough points are low speed data points.
[0060] For trough points, since trough data points belong to low-speed vehicle data points, the more data points there are around the trough points, the more reasonable it is, and the more uniform the density is, the more it conforms to the characteristics of low-speed congestion, that is, the collected data is more credible.
[0061] With a trough point as the center, a window of a preset time length is established to obtain the number of data points in the window; the horizontal coordinate difference between the latter data point and the previous data point in every two adjacent data points in the window is calculated, recorded as the time difference; the number of data points in the window is multiplied by the inverse of the standard deviation of the corresponding time difference in the window and normalized to obtain the trough credibility of the trough point;
[0062] For the peak points, the peak points belong to the data points of high-speed vehicles, so the fewer the number of data points around the peak points, the more reasonable it is, and the more discrete the density is, the more it conforms to the characteristics of high-speed traffic, that is, the collected data is more reliable.
[0063] With a peak point as the center, a window with a time scale of a preset time length n is established to obtain the number of data points in the window; the horizontal coordinate difference between the latter and the previous data point of each two adjacent data points in the window is calculated and recorded as the time difference; the reciprocal of the number of data points in the window is multiplied by the reciprocal of the standard deviation of the corresponding time difference in the window and normalized to obtain the peak credibility of the peak point. The preset time length n is 10 minutes, which can be adjusted by the implementer according to the actual situation.
[0064] The specific calculation formulas for the trough credibility of the trough point and the peak credibility of the peak point are:
[0065] ,
[0066] ,
[0067] Among them, βi represents the trough credibility of the i-th trough point; Gs is the number of data points contained in the trough point window; The time difference between each data point and the previous data point starting from the second data point in the window, that is, the mean of the horizontal axis difference; is the ath time difference in the window; Gs-1 is the number of time differences in the window; is the standard deviation of the time difference within the window; It means that the more data points there are in the window corresponding to the trough point, the smaller the standard deviation of the time difference is, and the greater the credibility of the data distribution around the trough point is.
[0068] γj represents the peak credibility of the j-th peak point; Gs is the number of data points contained in the peak point window; The time difference between each data point and the previous data point starting from the second data point in the window, that is, the mean of the horizontal axis difference; is the ath time difference in the window; is the number of time differences in the window; is the standard deviation of the time difference within the window; The smaller the number of data points in the window corresponding to the trough point, the smaller the standard deviation of the time difference, and the greater the credibility of the data distribution around the peak point. Thus, the peak credibility of each peak point and the trough credibility of each trough point in the sequence can be obtained.
[0069] The ultimate goal of this application is to obtain the passing speed that can characterize all vehicles passing through the road section within a period of time, and the passing speed of the vehicle requires all data to participate in the analysis as a whole. Therefore, after obtaining the credibility of the peak and trough data points, it is necessary to obtain the peak points and trough points with higher credibility in the overall data and compare and analyze them with the overall data in the sequence. Thus, a credible threshold is set, wherein the acquisition of the credible threshold can be calculated by historical data, and then determined by mathematical statistics to obtain the demarcation point, and then the credible threshold is obtained. If the peak credibility of the peak point is greater than the threshold, the peak point is a credible peak point, and if the trough credibility of the trough point is greater than the threshold, the trough point is a credible trough point.
[0070] Step S3, averaging the data values in the window corresponding to the credible trough point and the credible peak point to obtain a sampling mean; and calculating the credibility of the focused analysis sequence based on the sampling mean, the trough point and the peak point.
[0071] In step S2, the credible trough points and credible peak points are obtained, and further, the overall credibility of the focused analysis sequence needs to be analyzed. First, the sampling threshold needs to be calculated based on the credible trough points and credible peak points, and compared with the overall mean of the focused analysis sequence, and then the credibility of the sequence is obtained by combining the peak credibility and trough credibility of all peak points and trough points in the sequence.
[0072] The data values in the window corresponding to the credible trough point and the credible peak point are averaged to obtain the sampling mean.
[0073] Furthermore, the credibility of the focused analysis sequence is calculated based on the sampling mean, trough points and peak points. The absolute value of the difference between the sampling mean and the mean of the focused analysis sequence is inverted and multiplied with the adjustment factor to obtain the mean difference term; the peak credibility mean of all peak points of the focused analysis sequence and the trough credibility mean of all trough points are added to obtain the peak and trough distribution credibility term; the value difference term is multiplied by the peak and trough distribution credibility term to obtain the credibility of the focused analysis sequence.
[0074] The specific calculation formula is:
[0075] ,
[0076] Among them, δz is the credibility of the zth focused analysis sequence; CJ and ZJ are the sampling mean and the overall mean of the focused analysis sequence respectively; βi represents the trough credibility of the i-th trough point; represents the peak credibility of the j-th peak point; Fz and Gz are the number of peak points and trough points in the z-th focused analysis sequence respectively; k is the adjustment factor of the difference between the sampling mean and the overall mean, is the mean difference term, and its empirical value is 0.2. Distribute credible items for peaks and troughs, It means that the smaller the difference between the sampling mean and the overall mean in the zth focused analysis sequence, the greater the credibility of the peak and trough data point distribution, the higher the credibility of the zth focused analysis sequence. The credibility of other focused analysis sequences can be obtained in the same way.
[0077] Step S4, establish a window with a data point in the focused analysis sequence as the starting point, calculate the stationarity of the data point according to the ordinate and abscissa of each data point in the window and the number of data points, calculate the turning point probability of the data point according to the slope between the data point and the next data point; and screen the data points to obtain key points according to the stationarity and turning point probability of the data points.
[0078] Since the collected data are all the speeds of vehicles passing through the road section within a period of time, the distribution characteristics of each focused analysis sequence are similar. After obtaining the credibility of the focused analysis sequence, it is also necessary to obtain the key points in the focused analysis sequence, analyze the similarity of the key points of each focused analysis sequence, and then obtain the control data of all data based on the stability of the data in the sequence, that is, the most stable sequence among the two focused analysis sequences with the greatest similarity. The confidence of each focused analysis sequence is obtained from the control sequence, and then the respective weights are obtained based on the confidence and credibility of each focused analysis sequence.
[0079] For the above key points, the speed data is analyzed and obtained according to its characteristics. The speed data is related to whether the road is congested. In the corresponding scatter plot, it is shown that there are more data points at low speeds and greater volatility, and fewer data points at high speeds and less volatility. Therefore, the key points in each focused analysis sequence must not only obtain the turning points of the data but also the stable points. The two are combined to obtain the key points of the focused analysis sequence, and then analyzed. The turning point represents the part of each focused analysis sequence where the speed changes greatly, that is, the congested state turns to the unobstructed state, or the unobstructed state turns to the congested state; the stable point represents the part of each focused analysis sequence where the speed is relatively stable, that is, the data points in the congested state or the unobstructed state. The combination of the two can obtain the general trend of the speed of each focused analysis sequence, so as to better analyze the similarities between the focused analysis sequences.
[0080] When judging whether a data point is a stationary point or a turning point, it is necessary to establish a window with a data point in the sequence as the starting point, calculate the stationarity of the data point according to the ordinate and abscissa of each data point in the window and the number of data points, and calculate the turning point probability of the data point according to the slope between the data point and the next data point, where the length of the window is the preset time length.
[0081] For the stability of data points, since the distribution of data in the window is different between congested and unobstructed stable points, we need to consider not only the variance of the data points in the window, but also the number of data points in the window. Congested stable points have more data points in the window, which leads to more fluctuations, but the overall trend is stable; unobstructed stable points have fewer data points in the window and less fluctuations.
[0082] The standard deviation of the ordinate and the standard deviation of the abscissa of the data point in the window corresponding to a data point are obtained respectively, and recorded as the ordinate standard deviation and the abscissa standard deviation; the ordinate standard deviation, the abscissa standard deviation, the inverse of the number of data points in the window and the adjustment factor are multiplied to obtain the multiplication result; the inverse of the multiplication result is normalized to obtain the stationarity of the data point.
[0083] For the data points belonging to the turning points, first, the instantaneous change rate at the turning point is large, and secondly, except for the turning point, other data points gradually become sparse or dense. Therefore, when obtaining the turning point probability of each data point, first obtain the instantaneous change rate of the target data point, and then obtain the standard deviation of the horizontal coordinates of the data points in the window except the target data point. Combining the two, the turning point probability of each data point can be obtained.
[0084] The turning point probability of a data point is calculated according to the slope between a data point and the next data point. Specifically, the slope between a data point and the next data point is calculated, which is recorded as the instantaneous rate of change; the standard deviation of the horizontal coordinates of the data points other than the data point in the window corresponding to the data point is obtained, which is recorded as the first horizontal coordinate standard deviation; the instantaneous rate of change is multiplied by the inverse of the first horizontal coordinate standard deviation and normalized to obtain the turning point probability of the data point.
[0085] The specific calculation formula for the stability of the data points and the probability of turning points is:
[0086] ,
[0087] ,
[0088] in, and They represent the stability and turning point probability of the x-th data point respectively; norm represents the normalization operation; The standard deviation of the ordinate value of the window data point corresponding to the x-th data point, that is, the ordinate standard deviation; The standard deviation of the horizontal coordinate value of the window data point corresponding to the x-th data point, that is, the horizontal coordinate standard deviation; is the number of window data points corresponding to the x-th data point; k is the adjustment factor, which adjusts the impact of the number of data points on the stability, and the empirical value of k is 0.2; When there are many data points in the window, it is generally a congested section with dense data and a high possibility of fluctuation. Therefore, the variance needs to be appropriately adjusted by the inverse of the number of data points to reduce the variance and meet the stability standard. When there are fewer data points, the degree of fluctuation in the window is small, and the variance needs to be increased by the inverse of the number of data points to meet the stability standard. It means that the smaller the variance of the horizontal and vertical coordinates of the data in the window, the greater the stability.
[0089] It is the slope of the straight line between a data point and the next data point, that is, the instantaneous rate of change of the data point; is the standard deviation of the horizontal coordinates of the data points in the window corresponding to the data point except the data point, that is, the first horizontal coordinate standard deviation; It means that the greater the instantaneous change rate of the data point in the window, the smaller the standard deviation of the horizontal coordinates of the data points in the window except for the data point, and the greater the probability that the data point is a turning point. and Represent the values of the ordinates of the r+1th and rth data points respectively; and Represents the values of the horizontal coordinates of the r+1th and rth data points respectively.
[0090] Finally, the data points are screened according to the stability and turning point probability of the data points to obtain key points. The specific screening threshold is set, preferably 0.7, which can be determined by the implementer according to the actual situation. If the stability or turning point probability of a data point in the focused analysis sequence is greater than the screening threshold, the data point is a key point. In this way, the key points in each focused analysis sequence of the same period can be obtained for subsequent analysis.
[0091] Step S5, obtaining the similarity of key points between every two focused analysis sequences, taking the two focused analysis sequences with the greatest similarity as the reference sequence pair; calculating the stability of each sequence in the reference sequence pair, and selecting the sequence with the largest stability as the reference sequence.
[0092] After obtaining the key points in each focused analysis sequence in the same period, the key points of each focused analysis sequence are the stable part and the turning part in the data, representing the general trend of each focused analysis sequence, and each focused analysis sequence collected is obtained in the same period, so they are similar in general trend. The two focused analysis sequences with the highest similarity have the highest confidence, that is, the speed data in these two sequences are more accurate compared with the data in other sequences, but a single or a small amount of speed data is accidental and may not accurately represent the speed of vehicles passing through the road during the period, so it is also necessary to obtain the confidence of other focused analysis sequences. To obtain the confidence of other focused analysis sequences, it is first necessary to obtain a control sequence from the two focused analysis sequences with the highest similarity. The key points in the two acquired focused analysis sequences have high similarity, but the overall data stability is inconsistent, that is, the collected vehicle speed data may fluctuate due to different detector installation positions and autonomous deceleration and acceleration of some vehicles, affecting the overall vehicle passing speed accuracy. Therefore, it is also necessary to analyze the stability of the two focused analysis sequence data, use the focused analysis sequence with greater stability as the control sequence, and then obtain the confidence of other focused analysis sequence data.
[0093] After obtaining the key points of each emphasized analysis sequence, the DTW algorithm is used to calculate the similarity of the key points in the two emphasized analysis sequences as the similarity of the two emphasized analysis sequences. A similarity is obtained between every two emphasized analysis sequences, and the two emphasized analysis sequences with the largest similarity are obtained as the control sequence pair.
[0094] The similarity between the two focused analysis sequences in the control sequence pair is the largest, and then one sequence is selected as the control sequence to obtain the confidence of the other focused analysis sequences. The stability of each sequence in the control sequence pair is calculated separately, and the sequence with the largest stability is selected as the control sequence. Specifically:
[0095] The window is established with the first data point of a focused analysis sequence in the control sequence pair as the starting point, and the window is slid according to the set step size to obtain the mean of the standard deviation of the data points in each window during the sliding, which is recorded as the stability of the focused analysis sequence. Since the key points in the two sequences in the control sequence pair are highly similar, their overall trends are similar. The sequence with a smaller mean of the standard deviation sequence is the sequence with a higher degree of stability, and the focused analysis sequence in the control sequence pair with a higher degree of stability is used as the control sequence. The step size is set as Preferably, n is a preset time length of 10 minutes, which can be adjusted by the implementer according to the actual situation. When sliding, all the data in the sequence are traversed by the sliding window, and the sliding ends.
[0096] Therefore, the sequence with the highest stability in the reference sequence pair is selected as the reference sequence.
[0097] Step S6, calculating the similarity between each focused analysis sequence and the control sequence, recorded as confidence; obtaining an adaptive weight according to the confidence and credibility of the focused analysis sequence; obtaining the vehicle passing speed in the period according to the adaptive weight and performing congestion identification.
[0098] After obtaining the focused analysis sequence with the smallest standard deviation mean and the largest similarity in the control sequence pair, it is marked as the control sequence, and then the similarity between each focused analysis sequence and the control sequence in the same time period is obtained based on the control sequence. Here, since the data length may be different, DTW is used to obtain the similarity of the data. After normalization, the obtained similarity is used as the confidence of each focused analysis sequence, wherein the confidence of the control sequence is set to a preset value. Preferably, the preset value in the embodiment of the present invention is 1.
[0099] After obtaining the confidence of each focused analysis sequence in the same period, the confidence of each focused analysis sequence is combined with the credibility to obtain the adaptive weight of each focused analysis sequence.
[0100] Specifically, the confidence and credibility of a focused analysis sequence are multiplied as the numerator, the confidence and credibility of each focused analysis sequence in the same period are multiplied and summed as the denominator, and the ratio is the adaptive weight of the focused analysis sequence.
[0101] The traditional fusion of vehicle passing speed from multi-source data is to fuse the mean of each group of data under the weight of a fixed threshold without considering the characteristics of each group of data. Therefore, the obtained vehicle passing speed is not accurate.
[0102] Furthermore, the vehicle passing speed of the time period is obtained according to the adaptive weight and congestion identification is performed. Specifically, the mean of the data in each focused analysis sequence of the same time period is weighted and summed based on the adaptive weights of each focused analysis sequence of the time period to obtain the vehicle passing speed of the time period.
[0103] After obtaining the vehicle passing speed of a period through the above operation, the speed threshold is set according to the traditional threshold method, and the period below the speed threshold is marked as a congested period.
[0104] The time periods below the speed threshold are marked by the threshold method, and the frequency of traffic network data congestion on each road is counted. Professionals are reminded to dispatch technicians to investigate the causes of abnormal congestion on sections with high traffic network data congestion, and take corresponding measures to alleviate road pressure. The speed threshold is obtained by the implementer through statistical analysis of the vehicle speed when the specific section is congested. The implementer needs to set it according to the specific situation, such as counting the dividing line between vehicle speed when there is no congestion and when there is congestion, and using this as the speed threshold, or using other methods to obtain the threshold based on statistical data. This is a prior art and will not be elaborated on here.
[0105] The present invention collects data from each traffic flow detector and floating vehicle on each road to obtain a speed sequence, and then selects the time period that needs to be focused on. The speed sequence in the time period is recorded as a focused analysis sequence, and its credibility is obtained according to the distribution characteristics of each focused analysis sequence. Then, according to the corresponding relationship between each focused analysis sequence, the confidence of each focused analysis sequence is obtained, and the adaptive weight of each focused analysis sequence is obtained according to the credibility and confidence of each focused analysis sequence, and then a new weighted average speed is obtained, that is, the vehicle passing speed. Then, the threshold is set by the threshold method to mark the time period below the speed threshold, and the frequency of traffic network data congestion on each road is counted. Professionals are reminded to dispatch technicians to investigate the causes of abnormal congestion in sections with high incidence of traffic network data congestion, and take corresponding measures to alleviate road pressure.
[0106] It should be noted that the sequence of the above embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The above is a description of a specific embodiment of this specification. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0107] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.
[0108] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A traffic network congestion identification method based on data fusion, characterized in that: The method includes: Collect the speed data of vehicles traveling on a road section within a period of time, obtain the speed sequences of the period of time, and obtain the ordinate and abscissa of each data point; The velocity sequence is screened to obtain the focused analysis sequence; a window is established with the trough point and the peak point in the focused analysis sequence as the center, and the trough credibility and the peak credibility are calculated according to the data in the window, and the credible trough point and the credible peak point are obtained; The data values in the window corresponding to the credible trough point and the credible peak point are averaged to obtain the sampling mean; the credibility of the sequence is analyzed based on the sampling mean, trough point and peak point calculation; A window is established with a data point in the sequence being analyzed as the starting point. The stationarity of the data point is calculated based on the ordinate and abscissa of each data point in the window and the number of data points. The turning point probability of the data point is calculated based on the slope between the data point and the next data point. The data points are screened to obtain key points based on the stationarity and turning point probability of the data points. Calculate the similarity of key points between each two focused analysis sequences, and select the two focused analysis sequences with the largest similarity as the control sequence pair; calculate the stability of each sequence in the control sequence pair, and select the sequence with the largest stability as the control sequence; The similarity between each focused analysis sequence and the control sequence is calculated and recorded as confidence; the adaptive weight is obtained according to the confidence and credibility of the focused analysis sequence; the vehicle passing speed in the period is obtained according to the adaptive weight and congestion identification is performed.
2. A traffic network congestion identification method based on data fusion according to claim 1, characterized in that: The method of collecting speed data of a vehicle traveling on a road section within a period of time, obtaining each speed sequence of the period of time, and obtaining the ordinate and abscissa of each data point includes: Install a preset number of traffic flow detectors on a road section, use a traffic flow detector to collect vehicle speed data during the period, and form a speed sequence corresponding to the traffic flow detector; Get the average of all speed data uploaded at fixed intervals when a vehicle passes through the road section within a period of time, recorded as the speed average. The speed average of all vehicles passing through the road section within a period of time constitutes the speed sequence corresponding to the floating vehicle; the timestamp of each data point in the speed sequence corresponding to the floating vehicle is the time when each vehicle enters the road section; obtain each speed sequence corresponding to the period of time; Map each data in the speed sequence corresponding to the traffic flow detector to a two-dimensional rectangular coordinate system, where the ordinate and abscissa of each data point are the speed value and the acquisition time respectively; Each data point in the speed sequence corresponding to the floating vehicle is mapped to a two-dimensional rectangular coordinate system, and the ordinate and abscissa of each data point are the mean speed and the moment when the vehicle enters the road section, respectively.
3. The method for identifying traffic network congestion based on data fusion according to claim 1, characterized in that: The screening of the speed sequence to obtain the focused analysis sequence includes: The speed range is set according to the maximum speed limit of the road section, and the number of data values of data points in a speed sequence within the speed range is obtained, which is recorded as the number of normal passing vehicles; the number of data points in the speed sequence with a preset multiple is used as the number threshold, if the number of normal passing vehicles corresponding to each speed sequence in a time period is greater than or equal to the number threshold, then each speed sequence in this time period is a regular analysis sequence; if the number of normal passing vehicles corresponding to each speed sequence in a time period is not greater than or equal to the number threshold, then each speed sequence in this time period is a focused analysis sequence.
4. The method for identifying traffic network congestion based on data fusion according to claim 1, characterized in that: The step of calculating the trough credibility and the peak credibility respectively according to the data in the window and obtaining the credible trough point and the credible peak point includes: With a trough point as the center, a window of a preset time length is established to obtain the number of data points in the window; the horizontal coordinate difference between the latter data point and the previous data point in every two adjacent data points in the window is calculated, recorded as the time difference; the number of data points in the window is multiplied by the inverse of the standard deviation of the corresponding time difference in the window and normalized to obtain the trough credibility of the trough point; With a peak point as the center, a window of a preset time length is established to obtain the number of data points in the window; the horizontal coordinate difference between the latter data point and the previous data point in every two adjacent data points in the window is calculated, recorded as the time difference; the reciprocal of the number of data points in the window is multiplied by the reciprocal of the standard deviation of the corresponding time difference in the window and normalized to obtain the peak credibility of the peak point; A credible threshold is set. If the peak credibility of a peak point is greater than the threshold, the peak point is a credible peak point. If the trough credibility of a trough point is greater than the threshold, the trough point is a credible trough point.
5. The method for identifying traffic network congestion based on data fusion according to claim 1, characterized in that: The calculation based on the sampling mean, trough point and peak point focuses on analyzing the credibility of the sequence, including: The absolute value of the difference between the sampling mean and the mean of the focused analysis sequence is inverted and multiplied by the adjustment factor to obtain the mean difference term; the mean of the peak credibility of all peak points of the focused analysis sequence and the mean of the trough credibility of all trough points are added to obtain the peak and trough distribution credibility term; the value difference term is multiplied by the peak and trough distribution credibility term to obtain the credibility of the focused analysis sequence.
6. The method for identifying traffic network congestion based on data fusion according to claim 1, characterized in that: The step of establishing a window with a data point in the focused analysis sequence as the starting point, calculating the stationarity of the data point according to the ordinate and abscissa of each data point in the window and the number of data points, and calculating the turning point probability of the data point according to the slope between the data point and the next data point, comprises: Respectively obtain the standard deviation of the ordinate and the standard deviation of the abscissa of the data point in the window corresponding to a data point, and record them as the ordinate standard deviation and the abscissa standard deviation; multiply the ordinate standard deviation, the abscissa standard deviation, the inverse of the number of data points in the window and the adjustment factor to obtain a multiplication result; normalize the inverse of the multiplication result to obtain the stability of the data point; Calculate the slope between one data point and the next data point, recorded as the instantaneous rate of change; obtain the standard deviation of the horizontal coordinates of the data points other than the data point in the window corresponding to the data point, recorded as the first horizontal coordinate standard deviation; multiply the instantaneous rate of change by the inverse of the first horizontal coordinate standard deviation and normalize them to obtain the turning point probability of the data point.
7. The method for identifying traffic network congestion based on data fusion according to claim 1, characterized in that: The step of screening the data points to obtain key points according to the stability and turning point probability of the data points includes: Set the screening threshold. If the stability or turning point probability of a data point in the sequence being analyzed is greater than the screening threshold, then the data point is a key point.
8. The method for identifying traffic network congestion based on data fusion according to claim 1, characterized in that: The step of calculating the stationarity of each sequence in the reference sequence pair and selecting the sequence with the largest stationarity as the reference sequence comprises: A window is established with the first data point of a focused analysis sequence in the control sequence pair as the starting point, and the window is slid according to the set step size to obtain the mean of the standard deviation of the data points in each window during sliding, which is recorded as the stationarity of the focused analysis sequence; the focused analysis sequence with a larger stationarity in the control sequence pair is selected as the control sequence.
9. The method for identifying traffic network congestion based on data fusion according to claim 1, characterized in that: The step of obtaining the adaptive weight based on the confidence and credibility of the focused analysis sequence includes: The confidence and credibility of a focused analysis sequence are multiplied as the numerator, and the confidence and credibility of each focused analysis sequence in the same period are multiplied and summed as the denominator. The ratio is the adaptive weight of the focused analysis sequence.
10. The method for identifying traffic network congestion based on data fusion according to claim 1, characterized in that: The obtaining of the vehicle passing speed in the time period according to the adaptive weight and performing congestion identification includes: The mean of the data in each focused analysis sequence in the same time period is weighted and summed based on the adaptive weights of each focused analysis sequence in the time period to obtain the vehicle passing speed of the time period; a speed threshold is set, and if the vehicle passing speed of the time period is lower than the speed threshold, the time period is marked as a congested time period.
Citation Information
Patent Citations
Expression method of road traffic information credibility space characteristics based on sensor network
CN103413428A
Method and device for identifying chromatin open region based on sequencing data
CN111724860A