A method for evaluating ecological adaptability of transportation network

By classifying and collecting proportions of data in the ecological adaptability evaluation of the transportation road network, combining the impact index, correlation index and data quality index, the problem of imbalance in data collection proportions is solved, and the accuracy of data analysis and scientific decision-making are improved.

CN119669943BActive Publication Date: 2025-06-06TIANJIN RES INST FOR WATER TRANSPORT ENG M O T +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510180086.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-06
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

The existing technology has the problem of imbalance in the proportion of data acquisition in the ecological adaptability evaluation of the transportation road network, which leads to a decrease in the accuracy and representativeness of data analysis, affecting the scientificity and effectiveness of the final decision.

Method used

By classifying the data to be collected, and designing the acquisition ratio for each type of data, the comprehensive score of each type of data is calculated based on the combined analysis of the impact index, correlation index, comprehensive data quality index and weight, so as to determine its acquisition ratio and ensure the balance and representativeness of data acquisition.

Benefits of technology

It improves the scientificity and accuracy of data collection, avoids analysis deviations, improves the scientificity and accuracy of the final evaluation results, and thus improves the effectiveness of decision-making execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669943B_ABST
    Figure CN119669943B_ABST
Patent Text Reader

Abstract

The invention discloses a method for evaluating the ecological adaptability of a traffic network, and relates to the technical field of ecological adaptability evaluation. The invention aims to ensure the balance and representativeness of data in a collection process, by classifying the data to be collected, and then designing a collection ratio for each type of data, and because the collection ratio is calculated based on a series of data-driven methods, specifically in the calculation process of obtaining the collection ratio, a comprehensive score for each type of data is obtained by combining and analyzing an influence index, a correlation index, a comprehensive data quality index and a weight, and then a collection ratio is formed according to the ratio of the comprehensive score to the total score. Therefore, the method has high scientificity and accuracy, and good executability, thereby helping to avoid analysis deviations caused by imbalanced data collection ratios, and further being able to effectively improve the scientificity and accuracy of the final evaluation results to the greatest extent, so as to effectively improve the effect of final decision execution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ecological adaptability evaluation, and in particular to a method for evaluating the ecological adaptability of a transportation network. Background Art

[0002] The evaluation of the ecological adaptability of a transportation network refers to the process of evaluating the performance of a road transportation network in an ecological environment. This evaluation aims to ensure that the development of transportation infrastructure not only meets transportation needs, but also minimizes the negative impact on the natural environment and promotes sustainable development. In order to achieve an accurate evaluation of the ecological adaptability of a transportation network, big data analysis is an important method. This method utilizes a large amount of data collected from various channels, and uses complex algorithms and analytical tools to reveal hidden patterns, unknown correlations, and other useful information to support decision-making. The following is a complete process for evaluating the ecological adaptability of a transportation network through big data analysis:

[0003] First, data collection is performed, including mobile device data, social media data, and sensor data. Mobile device data includes the use of location data from smartphones and other mobile devices to track the movement patterns of people. This data can show when and where people travel, and the mode of transportation they choose. Social media data includes analyzing posts, pictures, and videos on social media to understand people's views and experiences of traffic conditions. Sensor data includes sensors installed on transportation infrastructure such as roads and bridges, which can collect information about traffic flow, speed, weather conditions, etc. Next, data cleaning and preprocessing are performed, including cleaning data to remove errors and duplicates; processing missing values; standardizing data formats from different sources for subsequent integration and analysis; then data integration is performed, including integrating all relevant data sets onto a unified platform, such as a data warehouse or data lake, to facilitate cross-domain correlation analysis; establishing connections between data, such as linking traffic flow data with social media sentiment analysis results; finally, generating an evaluation report after performing data analysis, including descriptive analysis, which summarizes the basic characteristics of the current traffic situation; diagnostic analysis, which identifies the causes of traffic congestion, accident-prone locations, and other problems, and then writing an evaluation report based on the analysis results to convey key findings and recommendations to decision makers;

[0004] Although big data analysis provides a powerful tool for evaluating the ecological adaptability of transportation networks, there are still some defects in practical applications, especially the problem of imbalanced data collection ratio:

[0005] Specifically, in big data analysis, we often encounter situations where the amount of data collected for certain types is too large, while the amount of data collected for other types is too small. This imbalance will lead to reduced accuracy and representativeness in subsequent data analysis. When the data sample is not sufficient to represent the overall situation, the analysis results will tend to be overly biased towards certain specific conditions and ignore other important variables, thereby affecting the scientificity and accuracy of the final results. This will lead to evaluations based on one-sided data that easily ignore the needs of certain important groups, resulting in the final decision being prone to poor results due to lack of scientificity.

[0006] Therefore, the prior art urgently needs a technical solution for evaluating the ecological adaptability of a transportation network. Summary of the invention

[0007] In order to solve the above technical problems, the present invention provides a method for evaluating the ecological adaptability of a transportation network, which specifically includes the following steps:

[0008] Step S1, classify the data to be collected according to their nature, perform analysis on each type of data, and obtain the collection ratio of each type of data based on the analysis results;

[0009] Step S1a, by analyzing the historical data, determining the contribution of each type of data to the final analysis result, and obtaining the influence index of each type of data;

[0010] Step S1a1, collecting usage of each type of data in the historical data, wherein the usage includes the usage frequency, application scope and coverage of each type of data;

[0011] Step S1a2: based on the usage frequency, application scope and coverage of each type of data, obtain the contribution of each type of data to the final analysis result;

[0012] Among them, the calculation formula for the contribution of each type of data to the final analysis result is:

[0013] ;

[0014] in, Represents the contribution of the i-th type of data to the final analysis results; Represents the usage frequency of the i-th category of data; Represents the application scope of the i-th category of data; Represents the coverage of the i-th category data; , and Represent the weight coefficients of usage frequency, application scope and coverage respectively, and ;

[0015] Step S1a3, performing a conversion operation on the contribution of each type of data to the final analysis result to obtain an influence index of each type of data;

[0016] The calculation formula for the impact index of each type of data is:

[0017] ;

[0018] in, Represents the impact index of the i-th category of data; Represents the contribution of the i-th type of data to the final analysis results; represents a monotonically increasing function;

[0019] Step S1b, analyzing the correlation between each type of data and the final analysis result by statistical methods to obtain a correlation index for each type of data;

[0020] Step S1b1, obtaining the time series value of each type of data and the time series value of the final analysis result;

[0021] Step S1b2, using a sliding window method to segment the time series values ​​of each type of data and the final analysis result, to obtain data sets of several sub-time periods of each type of data and data sets of several sub-time periods of the final analysis result, wherein the data sets include several data points;

[0022] Step S1b3, using each data point in the data set of each sub-time period of each type of data and each data point in the data set of each sub-time period of the final analysis result, applying the Pearson correlation coefficient formula to analyze the sub-correlation between each type of data and the final analysis result, and obtaining the sub-correlation coefficient of each type of data in each sub-time period;

[0023] Among them, the calculation formula for obtaining the sub-correlation coefficient of each type of data in each sub-time period is:

[0024] ;

[0025] in, represents the sub-correlation coefficient of the i-th category of data in the j-th sub-time period; represents the kth data point in the data set of the i-th category of data in the j-th sub-time period; represents the mean of all data points of the data set of the i-th category in the j-th sub-time period; The kth data point in the data set representing the final analysis result in the jth sub-time period; Represents the mean of all data points in the data set in the jth sub-period of the analysis result; represents the number of data points in the dataset for the jth sub-period;

[0026] Step S1b4, combining the sub-correlation coefficients of each type of data in each sub-time period and taking the average, to obtain the correlation index of each type of data;

[0027] Among them, the calculation formula for the correlation index of each type of data is:

[0028] ;

[0029] in, Represents the correlation index of the i-th category of data; represents the sub-correlation coefficient of the i-th category of data in the j-th sub-time period; represents the number of sub-periods;

[0030] Step S1c, determining the data quality of each type of data by analyzing the historical data, and obtaining a data quality index for each type of data;

[0031] Step S1c1, calling each type of data collected for the N-1th, Nth and N+1th times in the historical data;

[0032] Step S1c2, performing accuracy evaluation on each type of data collected each time, and obtaining an accuracy score for each type of data collected each time;

[0033] Step S1c3, performing integrity assessment on each type of data collected each time, and obtaining an integrity score for each type of data collected each time;

[0034] Step S1c4, performing timeliness evaluation on each type of data collected each time, and obtaining a timeliness score for each type of data collected each time;

[0035] Step S1c5, based on the accuracy score, completeness score and timeliness score, obtaining a data quality index for each type of data collected each time;

[0036] The calculation formula for the data quality index of each type of data collected each time is:

[0037] ;

[0038] in, represents the data quality index of the i-th type of data collected at the p-th time; Represents the accuracy score of the i-th category of data collected at the p-th time; represents the completeness score of the i-th type of data collected at the p-th time; Represents the timeliness score of the i-th type of data collected at the p-th time; , and Represent the weight coefficients of accuracy score, completeness score and timeliness score respectively, and ; ;

[0039] Step S1c6: Calculate the comprehensive data quality index of each type of data according to the data quality index of each type of data, specifically:

[0040] ;

[0041] in, Represents the comprehensive data quality index of the i-th category of data; represents the data quality index of the i-th type of data collected at the N-1th time, represents the data quality index of the i-th type of data collected for the Nth time, Represents the data quality index of the i-th type of data collected for the N+1th time.

[0042] Step S1d, obtaining a comprehensive score for each type of data based on the influence index, relevance index and comprehensive data quality index of each type of data;

[0043] Step S1d1, based on the influence index, relevance index and comprehensive data quality index of each type of data, obtain a comprehensive index for each type of data;

[0044] Among them, the calculation formula for the comprehensive index of each type of data is:

[0045] ;

[0046] in, Represents the comprehensive index of the i-th category of data; Represents the impact index of the i-th category of data; Represents the correlation index of the i-th category of data; Represents the comprehensive data quality index of the i-th category of data;

[0047] Step S1d2, obtaining the weight of each type of data according to the comprehensive index of each type of data;

[0048] Among them, the calculation formula for the weight of each type of data is:

[0049] ;

[0050] in, Represents the weight of the i-th category data; represents the comprehensive index of the i-th category of data; x represents the number of data categories;

[0051] Step S1d3, obtaining a comprehensive score for each type of data based on the influence index, relevance index, comprehensive data quality index and weight of each type of data;

[0052] The calculation formula for obtaining the comprehensive score of each type of data is:

[0053] ;

[0054] in, Represents the comprehensive score of the i-th category data; Represents the impact index of the i-th category of data; Represents the correlation index of the i-th category of data; Represents the comprehensive data quality index of the i-th category of data; Represents the weight of the i-th category data;

[0055] Step S1e, obtaining the collection ratio of each type of data according to the comprehensive score of each type of data;

[0056] Step S1e1, summing up the comprehensive scores of each category of data to obtain the total scores of all categories of data;

[0057] Step S1e2, obtaining the collection ratio of each type of data according to the ratio of the comprehensive score of each type of data to the total score;

[0058] Step S2, performing data collection according to the collection ratio of each type of data, and performing preprocessing on each type of collected data;

[0059] Step S3, integrating each type of pre-processed data into a unified data platform, and performing descriptive analysis and diagnostic analysis on the integrated data respectively;

[0060] Step S4: Generate an evaluation report based on the analysis results of the descriptive analysis and the diagnostic analysis.

[0061] The embodiments of the present invention have the following technical effects:

[0062] The present invention aims to ensure the balance and representativeness of data in the collection process. It classifies the data to be collected and designs a collection ratio for each type of data. Since the collection ratio is calculated based on a series of data-driven methods, specifically in the calculation process of obtaining the collection ratio, the comprehensive score of each type of data is obtained by combining and analyzing the influence index, correlation index, comprehensive data quality index and weight, and then the collection ratio is formed according to the ratio of the comprehensive score to the total score. Therefore, it has high scientificity and accuracy, and good executability, which helps to avoid analysis deviations caused by imbalance in data collection ratios, and can effectively improve the scientificity and accuracy of the final evaluation results to the greatest extent, so as to effectively improve the effect of the final decision execution. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0064] Figure 1 It is a flow chart of a method for evaluating ecological adaptability of a transportation network provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the scope of protection of the present invention.

[0066] Embodiment 1: Figure 1 As shown, the present invention provides a method for evaluating the ecological adaptability of a transportation network, comprising the following steps:

[0067] Step S1: classify the data to be collected according to their properties, perform analysis on each type of data, and obtain the collection ratio of each type of data based on the analysis results.

[0068] Step S1a: By analyzing the historical data, the contribution of each type of data to the final analysis result is determined, and the influence index of each type of data is obtained.

[0069] Step S1a1: Collect usage of each type of data in historical data, wherein the usage includes usage frequency, application scope and coverage of each type of data.

[0070] First, export historical data from the database of the traffic management system, including all types of data records. At the same time, ensure that the data format is consistent, usually in CSV or JSON format, including timestamp, data type, data value and other fields. The definition of frequency of use refers to the number of times a certain type of data is used within a certain time range. Then extract the data type of each record from the historical data, and count the number of records of each data type within the specified time range. The specific steps are: use Python's pandas library to read the historical data file, then group the data by data type, and finally count the number of records of each data type. The statistical results are the frequency of use of each type of data; the definition of application scope refers to the use of a certain type of data in different application scenarios, such as traffic flow prediction, accident analysis, weather forecast and other scenarios. Then extract the application scenario of each record from the historical data, and count the number of times each data type is used in different application scenarios. The specific steps are the same as above. Continue to use the pandas library to read the historical data file, then double group the data by data type and application scenario, and finally count the number of records of each data type in different application scenarios, which is the application scope of each type of data; the definition of coverage refers to the integrity and continuity of a certain type of data within a specified time range. For example, the collection frequency and time interval of a certain type of data in a day, then extract the timestamp of each record from the historical data, and calculate the collection frequency and time interval of each data type within the specified time range. The specific steps are the same as above and will not be repeated here.

[0071] Step S1a2: Based on the usage frequency, application scope and coverage of each type of data, the contribution of each type of data to the final analysis result is obtained.

[0072] Among them, the calculation formula for the contribution of each type of data to the final analysis result is:

[0073] ;

[0074] in, Represents the contribution of the i-th type of data to the final analysis results; Represents the usage frequency of the i-th category of data; Represents the application scope of the i-th category of data; Represents the coverage of the i-th category data; , and Represent the weight coefficients of usage frequency, application scope and coverage respectively, and .

[0075] Step S1a3, performing a conversion operation on the contribution of each type of data to the final analysis result to obtain an influence index of each type of data;

[0076] The calculation formula for the impact index of each type of data is:

[0077] ;

[0078] in, Represents the impact index of the i-th category of data; Represents the contribution of the i-th type of data to the final analysis results; Represents a monotonically increasing function, which is used to map contribution to impact index, ensuring that the greater the contribution, the greater the impact index.

[0079] Step S1b, analyzing the correlation between each type of data and the final analysis result by statistical methods to obtain a correlation index for each type of data;

[0080] Step S1b1, obtaining the time series value of each type of data and the time series value of the final analysis result;

[0081] Export historical data from the database of the traffic management system, including all types of data records, and ensure that the data format is consistent, usually in CSV or JSON format, including timestamp, data type, data value and other fields; the definition of the time series value of each type of data refers to the observation value of a certain type of data at different time points, then extract the timestamp and data value of each record from the historical data, and organize the timestamp and data value of each data type into a time series. The specific steps include: exporting historical data files from the database of the traffic management system, ensuring that the data files contain the following fields: timestamp, data type and data value, then grouping the data according to the data type field, grouping the same type of data into one group, and finally for each type of data, Extract its timestamp and corresponding data value, and organize these timestamps and data values ​​into a time series, so as to form a time series value for each type of data; the definition of the time series value of the analysis result refers to the observation value of the analysis result at different time points, and the specific implementation steps include: exporting the analysis result data file from the database of the traffic management system, and ensuring that the data file contains the following fields: timestamp, analysis result type and analysis result value; then grouping the data according to the analysis result type field, grouping the analysis results of the same type into one group, and finally for each analysis result type, extract its timestamp and corresponding analysis result value, and organize these timestamps and analysis result values ​​into a time series, so as to form a time series value for each analysis result.

[0082] Step S1b2, using a sliding window method to segment the time series values ​​of each type of data and the final analysis result, to obtain data sets of several sub-time periods of each type of data and data sets of several sub-time periods of the final analysis result, wherein the data sets include several data points;

[0083] Step S1b3, using each data point in the data set of each sub-time period of each type of data and each data point in the data set of each sub-time period of the final analysis result, applying the Pearson correlation coefficient formula to analyze the sub-correlation between each type of data and the final analysis result, and obtaining the sub-correlation coefficient of each type of data in each sub-time period;

[0084] Among them, the calculation formula for obtaining the sub-correlation coefficient of each type of data in each sub-time period is:

[0085] ;

[0086] in, represents the sub-correlation coefficient of the i-th category of data in the j-th sub-time period; represents the kth data point in the data set of the i-th category of data in the j-th sub-time period; represents the mean of all data points of the data set of the i-th category in the j-th sub-time period; The kth data point in the data set representing the final analysis result in the jth sub-time period; Represents the mean of all data points in the data set in the jth sub-period of the analysis result; Represents the number of data points in the dataset for the jth sub-period.

[0087] Step S1b4, combining the sub-correlation coefficients of each type of data in each sub-time period and taking the average, to obtain the correlation index of each type of data;

[0088] Among them, the calculation formula for the correlation index of each type of data is:

[0089] ;

[0090] in, Represents the correlation index of the i-th category of data; represents the sub-correlation coefficient of the i-th category of data in the j-th sub-time period; Represents the number of sub-periods.

[0091] Step S1c, determining the data quality of each type of data by analyzing the historical data, and obtaining a data quality index for each type of data;

[0092] Step S1c1, calling each type of data collected for the N-1th, Nth and N+1th times in the historical data;

[0093] When calling each type of data collected for the N-1th, Nth and N+1th times in historical data, the first thing to understand is that when each type of data is collected for the Nth time, there must be a corresponding time point. If the corresponding time point for this collection is three o'clock in the afternoon of a certain month and year, and according to N-1 and N+1, it can be known that the data collection frequency is once an hour, then the corresponding N-1 and N+1 times are each type of data collected one hour before and after three o'clock in the afternoon of a certain month and year; in the specific calling process, first export the historical data file from the database of the traffic management system, and ensure that the data file contains the following fields: timestamp, data type and data value, then group the data according to the data type, and group the same type of data into one group. For each type of data, determine the time point of the N-1th, Nth and N+1th collection according to the timestamp, and execute the calling and collection work after the determination is completed.

[0094] Step S1c2, performing accuracy evaluation on each type of data collected each time, and obtaining an accuracy score for each type of data collected each time;

[0095] Accuracy refers to the degree of consistency between the data collected each time and the reference data. When evaluating, for each type of data, the data value collected each time is mainly compared with the reference data value, and error indicators such as absolute error, relative error, mean square error, etc. are used to quantify the accuracy of the data, thereby obtaining the accuracy score of each type of data.

[0096] Step S1c3, performing integrity assessment on each type of data collected each time, and obtaining an integrity score for each type of data collected each time;

[0097] Completeness refers to whether there are missing or incomplete records in the data collected each time. During the evaluation, for each type of data, we mainly check whether the data collected each time contains all the expected data points, and calculate the proportion of missing data or the data completeness rate, so as to obtain the completeness score of each type of data;

[0098] Step S1c4, performing timeliness evaluation on each type of data collected each time, and obtaining a timeliness score for each type of data collected each time;

[0099] Timeliness refers to whether the data collected each time is collected in a timely manner within the specified time range. During the evaluation, for each type of data, we mainly check whether the timestamp of each collected data is within the specified time range, and calculate the delay time or timeout ratio of data collection, so as to obtain the timeliness score of each type of data;

[0100] Step S1c5, based on the accuracy score, completeness score and timeliness score, obtaining a data quality index for each type of data collected each time;

[0101] The calculation formula for the data quality index of each type of data collected each time is:

[0102] ;

[0103] in, represents the data quality index of the i-th type of data collected at the p-th time; Represents the accuracy score of the i-th category of data collected at the p-th time; represents the completeness score of the i-th type of data collected at the p-th time; Represents the timeliness score of the i-th type of data collected at the p-th time; , and Represent the weight coefficients of accuracy score, completeness score and timeliness score respectively, and ; ;

[0104] Step S1c6: Calculate the comprehensive data quality index of each type of data according to the data quality index of each type of data, specifically:

[0105] ;

[0106] in, Represents the comprehensive data quality index of the i-th category of data; represents the data quality index of the i-th type of data collected at the N-1th time, represents the data quality index of the i-th type of data collected for the Nth time, Represents the data quality index of the i-th type of data collected for the N+1th time.

[0107] Step S1d, obtaining a comprehensive score for each type of data based on the influence index, relevance index and comprehensive data quality index of each type of data;

[0108] Step S1d1, based on the influence index, relevance index and comprehensive data quality index of each type of data, obtain a comprehensive index for each type of data;

[0109] Among them, the calculation formula for the comprehensive index of each type of data is:

[0110] ;

[0111] in, Represents the comprehensive index of the i-th category of data; Represents the impact index of the i-th category of data; Represents the correlation index of the i-th category of data; Represents the comprehensive data quality index of the i-th category of data;

[0112] It is worth noting that designing the comprehensive data quality index as an inverse form can amplify the importance of high-quality data. High-quality data, that is, data with a smaller comprehensive data quality index, will occupy a larger proportion in the calculation, thereby ensuring that more attention is given in subsequent analysis, which helps to ensure the reliability of the final analysis results; and using the ratio of the impact index to the correlation index can balance the importance of the data and its correlation with the analysis results. If a certain type of data has a high correlation but has little impact on the overall analysis, then its comprehensive index will not be too high. On the contrary, if a certain type of data is crucial to the overall analysis, even if the correlation is not particularly high, a higher comprehensive index will be obtained. This can avoid the deviation caused by a single dimension being too high and make the comprehensive index more reasonable.

[0113] Step S1d2, obtaining the weight of each type of data according to the comprehensive index of each type of data;

[0114] Among them, the calculation formula for the weight of each type of data is:

[0115] ;

[0116] in, Represents the weight of the i-th category data; represents the comprehensive index of the i-th category of data; x represents the number of data categories;

[0117] Step S1d3, obtaining a comprehensive score for each type of data based on the influence index, relevance index, comprehensive data quality index and weight of each type of data;

[0118] The calculation formula for obtaining the comprehensive score of each type of data is:

[0119] ;

[0120] in, Represents the comprehensive score of the i-th category data; Represents the impact index of the i-th category of data; Represents the correlation index of the i-th category of data; Represents the comprehensive data quality index of the i-th category of data; Represents the weight of the i-th category data;

[0121] Step S1e, obtaining the collection ratio of each type of data according to the comprehensive score of each type of data;

[0122] Step S1e1, summing up the comprehensive scores of each category of data to obtain the total scores of all categories of data;

[0123] Step S1e2, obtaining the collection ratio of each type of data according to the ratio of the comprehensive score of each type of data to the total score;

[0124] Step S2, performing data collection according to the collection ratio of each type of data, and performing preprocessing on each type of collected data;

[0125] Step S3, integrating each type of pre-processed data into a unified data platform, and performing descriptive analysis and diagnostic analysis on the integrated data respectively;

[0126] Descriptive analysis is to count and describe the basic characteristics of data, including data distribution, central tendency, degree of dispersion, etc. The specific analysis method includes calculating the basic statistics of each type of data after integration, such as mean, median, standard deviation, minimum value, maximum value, etc., and generating statistical reports based on these basic statistics to show the basic statistical characteristics of each type of data; then, it is necessary to draw data distribution diagrams based on each type of integrated data, including histograms and box plots, in order to show the distribution characteristics of the data. Specifically, the method of drawing a histogram is to divide the data into several intervals, and count the number of data in each interval, and then draw a bar chart, in order to show the frequency distribution of the data; the method of drawing a box plot is to draw a box, the upper and lower boundaries of the box are the first quartile and the third quartile, respectively, the horizontal line in the middle is the median, and the whiskers outside the box extend to the minimum and maximum values ​​to exclude outliers, and outliers need to be marked separately. The purpose of drawing a box plot is to show the five-number summary of the data, such as minimum value, first quartile, median, third quartile, maximum value and outliers;

[0127] Diagnostic analysis is an in-depth analysis of data to identify anomalies, trends, associations, etc. in the data, to help discover problems and potential causes; analysis methods include identifying outliers and outliers and performing trend analysis to identify long-term trends and short-term fluctuations in the data. Specifically, with regard to identifying outliers and outliers, statistical methods such as Z-score, IQR, etc. should be used to identify outliers and outliers in each type of data, and an outlier report should be generated based on this, listing all outliers and outliers, and providing possible cause analysis; with regard to trend analysis, trend analysis should be performed on each type of data to identify long-term trends and short-term fluctuations in the data, and a trend chart should be generated based on this to show the trend changes in the data, while identifying the key nodes and events of the trend changes.

[0128] Step S4: Generate an evaluation report based on the analysis results of the descriptive analysis and the diagnostic analysis.

[0129] Finally, the results of descriptive analysis and diagnostic analysis are integrated into a comprehensive analysis report, which should include basic statistical characteristics of the data, data distribution graphs, outlier reports and trend graphs, etc. Then, based on the integrated comprehensive analysis report, further write a detailed analysis report and explain the analysis results of each part. At the same time, based on the analysis results, an interpretation of the data characteristics and potential problems should be provided, and suggestions for improvement should be put forward.

[0130] It should be noted that the terms used in the present invention are only for describing specific embodiments, rather than limiting the scope of the present application. As shown in the present specification, unless the context clearly indicates an exception, the words "one", "a", "a kind of" and / or "the" do not specifically refer to the singular, but may also include the plural. The terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of more restrictions, the elements defined by the sentence "include one..." do not exclude the presence of other identical elements in the process, method or device including the elements.

[0131] It should also be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. Unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", etc. should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be a connection between the two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0132] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A method for evaluating the ecological adaptability of a transportation network, characterized in that: The following steps are involved: Step S1, classify the data to be collected according to their nature, perform analysis on each type of data, and obtain the collection ratio of each type of data based on the analysis results; Specifically include: Step S1a, by analyzing the historical data, determining the contribution of each type of data to the final analysis result, and obtaining the influence index of each type of data; Step S1b, analyzing the correlation between each type of data and the final analysis result by statistical methods to obtain a correlation index for each type of data; Step S1c, determining the data quality of each type of data by analyzing the historical data, and obtaining a data quality index for each type of data; Step S1d, obtaining a comprehensive score for each type of data based on the influence index, relevance index and comprehensive data quality index of each type of data; Specifically include: Step S1d1: Based on the influence index, relevance index and comprehensive data quality index of each type of data, a comprehensive index of each type of data is obtained: ; in, Represents the comprehensive index of the i-th category of data; Represents the impact index of the i-th category of data; Represents the correlation index of the i-th category of data; Represents the comprehensive data quality index of the i-th category of data; Step S1d2: According to the comprehensive index of each type of data, the weight of each type of data is obtained: ; In the formula, Represents the weight of the i-th category data; represents the comprehensive index of the i-th category of data; x represents the number of data categories; Step S1d3: Obtain a comprehensive score for each type of data based on the impact index, relevance index, comprehensive data quality index and weight of each type of data: ; Represents the comprehensive score of the i-th category data; Represents the impact index of the i-th category of data; Represents the correlation index of the i-th category of data; Represents the comprehensive data quality index of the i-th category of data; Represents the weight of the i-th category data; Step S1e, obtaining the collection ratio of each type of data according to the comprehensive score of each type of data; Step S2, performing data collection according to the collection ratio of each type of data, and performing preprocessing on each type of collected data; Step S3, integrating each type of pre-processed data into a unified data platform, and performing descriptive analysis and diagnostic analysis on the integrated data respectively; Step S4: Generate an evaluation report based on the analysis results of the descriptive analysis and the diagnostic analysis.

2. A method for evaluating ecological adaptability of a transportation network according to claim 1, characterized in that: The above method determines the contribution of each type of data to the final analysis result by analyzing the historical data, and obtains the influence index of each type of data, including: Step S1a1, collecting usage of each type of data in the historical data, wherein the usage includes the usage frequency, application scope and coverage of each type of data; Step S1a2: Based on the usage frequency, application scope and coverage of each type of data, the contribution of each type of data to the final analysis result is obtained: ; in, Represents the contribution of the i-th type of data to the final analysis results; Represents the usage frequency of the i-th category of data; Represents the application scope of the i-th category of data; Represents the coverage of the i-th category data; , and Represent the weight coefficients of usage frequency, application scope and coverage respectively, and ; Step S1a3, performing a conversion operation on the contribution of each type of data to the final analysis result to obtain an influence index of each type of data; The calculation formula for the impact index of each type of data is: ; In the formula, Represents the impact index of the i-th category of data; Represents the contribution of the i-th type of data to the final analysis results; Represents a monotonically increasing function.

3. A method for evaluating the ecological adaptability of a transportation network according to claim 1, characterized in that: The correlation between each type of data and the final analysis result is analyzed by statistical methods to obtain the correlation index of each type of data, including: Step S1b1, obtaining the time series value of each type of data and the time series value of the final analysis result; Step S1b2, using a sliding window method to segment the time series values ​​of each type of data and the analysis results, to obtain data sets of several sub-time periods for each type of data and data sets of several sub-time periods for the final analysis results, wherein the data sets include several data points; Step S1b3, using each data point in the data set of each sub-time period of each type of data and each data point in the data set of each sub-time period of the analysis result, apply the Pearson correlation coefficient formula to analyze the sub-correlation between each type of data and the analysis result, and obtain the sub-correlation coefficient of each type of data in each sub-time period: ; in, represents the sub-correlation coefficient of the i-th category of data in the j-th sub-time period; represents the kth data point in the data set of the i-th category of data in the j-th sub-time period; represents the mean of all data points of the data set of the i-th category in the j-th sub-time period; represents the kth data point in the data set of the jth sub-time period in the analysis result; Represents the mean of all data points in the data set in the jth sub-period of the analysis result; represents the number of data points in the dataset for the jth sub-period; Step S1b4: synthesize the sub-correlation coefficients of each type of data in each sub-time period and take the average to obtain the correlation index of each type of data: ; In the formula, Represents the correlation index of the i-th category of data; represents the sub-correlation coefficient of the i-th category of data in the j-th sub-time period; Represents the number of sub-time periods.

4. A method for evaluating the ecological adaptability of a transportation network according to claim 1, characterized in that: The data quality of each type of data is determined by analyzing the historical data, and the data quality index of each type of data is obtained, including: Step S1c1, calling each type of data collected for the N-1th, Nth and N+1th times in the historical data; Step S1c2, performing accuracy evaluation on each type of data collected each time, and obtaining an accuracy score for each type of data collected each time; Step S1c3, performing integrity assessment on each type of data collected each time, and obtaining an integrity score for each type of data collected each time; Step S1c4, performing timeliness evaluation on each type of data collected each time, and obtaining a timeliness score for each type of data collected each time; Step S1c5: Based on the accuracy score, completeness score and timeliness score, obtain the data quality index of each type of data collected each time: ; in, represents the data quality index of the i-th type of data collected at the p-th time; Represents the accuracy score of the i-th category of data collected at the p-th time; represents the completeness score of the i-th type of data collected at the p-th time; Represents the timeliness score of the i-th type of data collected at the p-th time; , and Represent the weight coefficients of accuracy score, completeness score and timeliness score respectively, and ; ; Step S1c6: Calculate the comprehensive data quality index of each type of data according to the data quality index of each type of data, specifically: ; in, Represents the comprehensive data quality index of the i-th category data; Represents the data quality index of the i-th type of data collected for the N-1th time, represents the data quality index of the i-th type of data collected for the Nth time, Represents the data quality index of the i-th type of data collected for the N+1th time.

5. A method for evaluating the ecological adaptability of a transportation network according to claim 1, characterized in that: The collection ratio of each type of data is obtained based on the comprehensive score of each type of data, including: Step S1e1, summing up the comprehensive scores of each category of data to obtain the total scores of all categories of data; Step S1e2: Obtain the collection ratio of each type of data based on the ratio of the comprehensive score of each type of data to the total score.

Citation Information

Patent Citations

  • Optimized method for analyzing maturity of regional development

    CN105868906A

  • Evaluation method based on boiler economic operation

    CN113205429A