Spatiotemporal data evaluation methods, equipment and computer-readable storage media
By employing a multi-level, progressive spatiotemporal data evaluation method, the evaluation granularity is refined layer by layer, solving the problem of balancing accuracy and efficiency in the evaluation of massive spatiotemporal data, and achieving efficient and accurate data quality evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
- Filing Date
- 2023-05-26
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies struggle to balance accuracy and efficiency when assessing the quality of massive amounts of spatiotemporal data.
A multi-level, progressive quality assessment approach is adopted, including macro, meso, and micro assessments. The assessment granularity is refined layer by layer. After determining whether the requirements are met through the macro quality assessment results, further meso and micro assessments are conducted, and finally, a dataset that meets all quality requirements is output.
It improves the efficiency of spatiotemporal data evaluation, reduces the invalid evaluation time and resource consumption of data that does not meet the upper-level quality requirements, and ensures the accuracy of evaluation results.
Smart Images

Figure CN116719812B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data technology, and in particular to a spatiotemporal data evaluation method, device, and computer-readable storage medium. Background Technology
[0002] Spatiotemporal data refers to data that simultaneously possesses temporal and spatial dimensions, such as vehicle and pedestrian trajectory data, and checkpoint capture data. It is characterized by being multi-source, massive, and rapidly updated. Currently, spatiotemporal data is applied in many fields, such as traffic management, disease monitoring, environmental monitoring, and public health. Emerging computing technologies are developed and applied to analyze massive, high-dimensional spatiotemporal data and uncover valuable information within it. In the context of big data, the value of data depends on its quality, which in turn affects people's analysis and decision-making. The impact of low-quality data has always been a significant problem faced by various fields, research institutions, and even enterprises.
[0003] Currently, although there are various methods for assessing data quality, they often fail to balance accuracy and efficiency when applied to assessing the quality of massive amounts of spatiotemporal data. Summary of the Invention
[0004] The main objective of this invention is to provide a spatiotemporal data evaluation method, device, and computer-readable storage medium, which aims to improve the efficiency of quality evaluation while ensuring the accuracy of the evaluation through a multi-level, progressive quality evaluation approach.
[0005] To achieve the above objectives, the present invention provides a spatiotemporal data evaluation method, the method comprising the following steps:
[0006] Obtain the spatiotemporal dataset to be evaluated, wherein the spatiotemporal dataset to be evaluated includes multiple spatiotemporal data to be evaluated, and each spatiotemporal data to be evaluated has attribute values of multiple dimensions;
[0007] The dataset quality evaluation indicators of the spatiotemporal dataset to be evaluated are used to obtain macroscopic quality evaluation results;
[0008] If the dataset quality assessment index of the spatiotemporal dataset to be assessed meets the macro quality requirements based on the macro quality assessment results, then the dimensional quality assessment index of the spatiotemporal dataset to be assessed is evaluated to obtain the meso quality assessment results.
[0009] If the dimensional quality assessment index of the spatiotemporal dataset to be evaluated is determined to meet the meso-level quality requirements based on the meso-level quality assessment results, then the data quality assessment index of the spatiotemporal dataset to be evaluated is assessed to obtain the micro-level quality assessment results.
[0010] If the data quality assessment indicators of the spatiotemporal dataset to be evaluated meet the micro-quality requirements based on the micro-quality assessment results, then the dataset quality assessment result that has passed the assessment is output.
[0011] To achieve the above objectives, the present invention also provides a spatiotemporal data evaluation device, the device comprising:
[0012] The acquisition module is used to acquire the spatiotemporal dataset to be evaluated, wherein the spatiotemporal dataset to be evaluated includes multiple spatiotemporal data to be evaluated, and each spatiotemporal data to be evaluated has attribute values in multiple dimensions;
[0013] The first evaluation module is used to evaluate the dataset quality evaluation indicators of the spatiotemporal dataset to be evaluated, and obtain macro-quality evaluation results.
[0014] The second evaluation module is used to evaluate the dimensional quality evaluation index of the spatiotemporal dataset to be evaluated if the dataset quality evaluation index of the dataset to be evaluated meets the macro quality requirements based on the macro quality evaluation result, and to obtain the meso quality evaluation result.
[0015] The third evaluation module is used to evaluate the data quality evaluation indicators of the spatiotemporal dataset to be evaluated if the dimensional quality evaluation indicators of the dataset to be evaluated meet the meso-quality requirements based on the meso-quality evaluation results, and to obtain the micro-quality evaluation results.
[0016] The output module is used to output the dataset quality assessment result if the data quality assessment index of the spatiotemporal dataset to be evaluated meets the micro-quality requirements based on the micro-quality assessment result.
[0017] To achieve the above objectives, the present invention also provides an electronic device, the electronic device comprising: a memory, a processor, and a spatiotemporal data evaluation program stored in the memory and executable on the processor, wherein the spatiotemporal data evaluation program, when executed by the processor, implements the steps of the spatiotemporal data evaluation method as described above.
[0018] Furthermore, to achieve the above objectives, the present invention also proposes a computer-readable storage medium storing a spatiotemporal data evaluation program, which, when executed by a processor, implements the steps of the spatiotemporal data evaluation method as described above.
[0019] In this embodiment of the invention, a spatiotemporal dataset to be evaluated is obtained. This dataset includes multiple pieces of spatiotemporal data, each with attribute values across multiple dimensions. The dataset quality evaluation indicators of the dataset are then evaluated to obtain a macro-quality evaluation result, thus achieving a macro-quality evaluation of the dataset. Furthermore, if the dataset quality evaluation indicators of the dataset meet the macro-quality requirements based on the macro-quality evaluation result, the dimensional quality evaluation indicators of the dataset are evaluated to obtain a meso-quality evaluation result. This allows for further evaluation of the dimensional quality of the dataset after the dataset quality evaluation indicators meet the macro-quality requirements. The evaluation metrics are assessed, and then, if the dimensional quality evaluation metrics of the spatiotemporal dataset to be evaluated meet the meso-level quality requirements based on the meso-level quality evaluation results, the data quality evaluation metrics of the spatiotemporal dataset to be evaluated are evaluated to obtain micro-level quality evaluation results. This achieves a multi-level progressive quality evaluation of the spatiotemporal dataset to be evaluated, from macro to meso to micro, and from dataset to dimension to data, after the dimensional quality evaluation metrics of the spatiotemporal dataset to be evaluated meet the meso-level quality requirements. Furthermore, if the data quality evaluation metrics of the spatiotemporal dataset to be evaluated meet the micro-level quality requirements based on the micro-level quality evaluation results, the dataset quality evaluation result that has passed the evaluation is output. On the one hand, compared to comprehensive quality assessment of spatiotemporal data, this invention employs a progressive, layer-by-layer assessment approach, moving from macro to micro levels, with the assessment granularity increasing from dataset to dimension to data. As the assessment granularity refines, the computational load increases layer by layer. Only spatiotemporal datasets that meet the upper-level quality requirements can be assessed at the lower level. This effectively reduces the time and resources required to assess spatiotemporal datasets that do not meet the upper-level quality requirements at the lower level, thus improving the efficiency of quality assessment. On the other hand, compared to methods that improve efficiency by reducing the number of assessment dimensions, the spatiotemporal datasets output by this invention must comprehensively meet macro, meso, and micro quality requirements. This ensures the accuracy of the quality assessment and overcomes the technical limitation that often makes it difficult to balance accuracy and efficiency when assessing massive amounts of spatiotemporal data. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating an embodiment of the spatiotemporal data evaluation method of the present invention;
[0021] Figure 2 This is a schematic diagram illustrating an example of the spatiotemporal data evaluation method of the present invention.
[0022] Figure 3This is a flowchart illustrating an example of the spatiotemporal data evaluation method of the present invention.
[0023] Figure 4 This is a flowchart illustrating an embodiment of step S20 of the spatiotemporal data evaluation method of the present invention;
[0024] Figure 5 This is a flowchart illustrating an embodiment of steps S60 to S70 of the spatiotemporal data evaluation method of the present invention.
[0025] Figure 6 This is a flowchart illustrating the steps for detecting noise points in the spatiotemporal data evaluation method of the present invention.
[0026] Figure 7 This is a flowchart illustrating an example of data processing for abnormal situations in the spatiotemporal data evaluation method of the present invention.
[0027] Figure 8 This is a flowchart illustrating another embodiment of step S60 of the spatiotemporal data evaluation method of the present invention;
[0028] Figure 9 This is a flowchart illustrating another example of data processing for abnormal situations in the spatiotemporal data evaluation method of the present invention.
[0029] Figure 10 This is a schematic diagram of the hardware operating environment involved in the embodiments of the present invention.
[0030] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0031] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. It should be noted that all data acquisition and processing in this specification are performed with the knowledge and authorization of the relevant users.
[0032] Reference Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the spatiotemporal data evaluation method of the present invention.
[0033] This invention provides an embodiment of a spatiotemporal data evaluation method. It should be noted that although the flowchart shows a logical order, in some cases, the steps shown or described may be executed in a different order. In this embodiment, the executing entity of the spatiotemporal data evaluation method can be a smartphone, personal computer, server, or other device; no limitation is imposed in this embodiment. For ease of description, the executing entity is omitted in this embodiment. In this embodiment, the spatiotemporal data evaluation method includes the following steps:
[0034] Step S10: Obtain the spatiotemporal dataset to be evaluated, wherein the spatiotemporal dataset to be evaluated includes multiple spatiotemporal data to be evaluated, and each spatiotemporal data to be evaluated has attribute values of multiple dimensions.
[0035] In this embodiment, spatiotemporal data refers to data that simultaneously possesses time and space dimensions, such as vehicle trajectory data and checkpoint capture data. It has the comprehensive characteristics of being multi-source, massive, and rapidly updated. The spatiotemporal dataset to be evaluated refers to a collection of multiple spatiotemporal datasets to be evaluated, collected by the acquisition device. Each piece of spatiotemporal dataset to be evaluated has attribute values in multiple dimensions, including at least time-dimensional attribute values and space-dimensional attribute values. It may also include attribute values in the dimensions of identification codes, speed, and data type. The acquisition device can be a GPS (Global Positioning System) sensor, a camera, etc.
[0036] In one feasible implementation, the spatiotemporal dataset to be evaluated, which is collected and uploaded by the acquisition device, can be obtained from the acquisition device, or the spatiotemporal dataset to be evaluated, which is collected by the acquisition device and stored in the database, can be obtained from the database. In this embodiment, no limitation is imposed.
[0037] Step S20: Evaluate the dataset quality evaluation indicators of the spatiotemporal dataset to be evaluated to obtain macroscopic quality evaluation results;
[0038] In this embodiment, macro-quality assessment refers to the process of calculating, statistically analyzing, and / or evaluating all the spatiotemporal data to be assessed in the spatiotemporal dataset to obtain the dataset quality assessment index value. The dataset quality assessment index can be used to characterize the overall quality of all data in the dataset, including data volume, timeliness, and data types. Data volume refers to the total number of all spatiotemporal data to be assessed in the dataset; timeliness refers to the time range of all the spatiotemporal data to be assessed in the dataset; and data types refer to the types of data covered by all the spatiotemporal data to be assessed in the dataset. The specific selection of the dataset quality assessment index can be determined according to the actual use of the spatiotemporal dataset to be assessed, and is not limited in this embodiment. After obtaining the spatiotemporal dataset to be assessed, the data volume, sampling time range, and data types of the spatiotemporal data to be assessed in the dataset can be statistically analyzed.
[0039] In one feasible implementation, at least one dataset quality assessment indicator to be evaluated for the spatiotemporal dataset to be evaluated can be determined in advance based on actual needs or user selection. Then, statistical analysis can be performed on each piece of spatiotemporal data to be evaluated in the spatiotemporal dataset to be evaluated to determine the dataset quality assessment indicator values of each dataset quality assessment indicator of the spatiotemporal dataset to be evaluated. The dataset quality assessment indicator values can be used as macro-quality assessment results. Alternatively, the dataset quality assessment indicators can be used to perform dataset quality scoring or dataset quality classification analysis on the spatiotemporal dataset to be evaluated. The obtained dataset quality score or dataset quality classification results can be used as macro-quality assessment results. The dataset quality classification results can include meeting macro-quality requirements, not meeting macro-quality requirements, meeting macro-quality requirements but requiring data processing, etc. The macro-quality assessment results can also include specific information on macro-quality assessment. For example, the data volume of the spatiotemporal dataset to be evaluated is evaluated. The data volume of the spatiotemporal dataset to be evaluated is 40,000, and the data volume required for actual use is 50,000. The evaluation result may be that it does not meet the macro quality requirements. The ratio of 40,000 to 50,000 can be used as the score, and the score can be used as the evaluation result. The evaluation result may also include specific information that the data volume of the spatiotemporal dataset to be evaluated is 40,000.
[0040] Step S30: If the dataset quality assessment index of the spatiotemporal dataset to be assessed meets the macro quality requirements based on the macro quality assessment result, then the dimensional quality assessment index of the spatiotemporal dataset to be assessed is evaluated to obtain the meso quality assessment result.
[0041] In this embodiment, meso-level quality assessment refers to the process of calculating, statistically analyzing, and / or evaluating all attribute values of at least one dimension in the spatiotemporal dataset to be assessed to obtain the dimensional quality assessment index value of the spatiotemporal dataset to be assessed. The dimensional quality assessment index can be used to characterize the overall quality of all attribute values of a certain dimension in the dataset, including dimensional stability, dimensional completeness, and dimensional accuracy. Dimensional stability characterizes the fluctuation range of all attribute values of a certain dimension of the spatiotemporal data. Spatiotemporal data needs to have a certain degree of stability within a certain dimensional range. For example, vehicles traveling on urban roads have similar speeds, while vehicles traveling on highways have similar speeds. Stability can be characterized by calculating standard deviation, variance, etc. The smaller the standard deviation or variance, the higher the stability, and the smaller the fluctuation range of the spatiotemporal data in a specific dimension. For example, stability can be at least one of spatiotemporal logical stability (e.g., sampling frame rate fluctuations, trajectory smoothness, etc.) and numerical stability (e.g., fluctuations in traffic values at the same time within the same period, etc.). Dimensional completeness characterizes the missing value of all attribute values of a certain dimension of the spatiotemporal data. The data needs to cover the scope required for the actual use of spatiotemporal data in a certain dimension. The completeness can be characterized by calculating the proportion of missing spatiotemporal data. The smaller the proportion of missing spatiotemporal data, the higher the completeness and the less missing spatiotemporal data. For example, the completeness can be at least one of temporal completeness (e.g., whether there is missing spatiotemporal data within a certain time range) and spatial completeness (e.g., whether the spatiotemporal data covers every district of a certain city). The accuracy is used to characterize the logical correctness of all attribute values of a certain dimension of spatiotemporal data. Spatiotemporal data needs to conform to certain logic within a certain dimension. For example, a car cannot be in Hainan one minute ago and in Beijing one minute later. It can be characterized by calculating the proportion of logically incorrect spatiotemporal data. The smaller the proportion of logically incorrect spatiotemporal data, the higher the accuracy and the more logically correct spatiotemporal data. For example, the accuracy can be at least one of format accuracy (e.g., whether the time format is correct, whether the latitude and longitude format is correct), spatiotemporal logical accuracy (e.g., trajectory noise points, speed anomalies), and contextual logical accuracy (e.g., device data conflicts).
[0042] In one feasible implementation, at least one dimension to be evaluated in the spatiotemporal dataset to be evaluated and at least one dimension quality evaluation index corresponding to each dimension can be determined in advance according to actual needs or user selection. Based on the macro-quality evaluation results, it is determined whether the spatiotemporal dataset to be evaluated meets the macro-quality requirements. If the dataset quality evaluation index of the spatiotemporal dataset to be evaluated meets the macro-quality requirements, statistical analysis can be performed on all attribute values corresponding to each dimension to be evaluated in the spatiotemporal dataset to be evaluated to determine the dimension quality evaluation index values of each dimension to be evaluated. The dimension quality evaluation index values can be used as the meso-level quality evaluation results. Alternatively, dimension quality scoring or dimension quality classification analysis can be performed on each dimension to be evaluated and / or the spatiotemporal dataset to be evaluated based on the dimension quality evaluation index values. The obtained dimension quality scoring or dimension quality classification results can be used as the meso-level quality evaluation results. The dimension quality classification results can include meeting the meso-level quality requirements, not meeting the meso-level quality requirements, meeting the meso-level quality requirements but requiring data processing, and can also include specific information about the meso-level quality evaluation. For example, the dimensional stability score is 70, the dimensional integrity score is 50, and the dimensional accuracy score is 75. The dimensional stability score of 70, the dimensional integrity score of 50, and the dimensional accuracy score of 75 can be used as the meso-level quality assessment result, or the average score of 65 or the total score of 195 can be used as the meso-level quality assessment result.
[0043] If the macro-quality assessment results determine that the spatiotemporal dataset to be evaluated does not meet the macro-quality requirements, the process can return to the step of obtaining the spatiotemporal dataset to be evaluated, and the dataset can be obtained again. A prompt message can also be output to remind the user that the currently obtained spatiotemporal dataset does not meet the macro-quality requirements and can be obtained again. Regardless of the meso- or micro-quality assessment results, spatiotemporal data that does not meet the macro-quality requirements cannot meet the practical needs of spatiotemporal data. For example, if the data volume is too small, even if each data point is of high quality, it may lose statistical significance. Similarly, data from ten years ago, even if each data point is of high quality, may not reflect the current situation. Furthermore, if only trajectory data is available without urban road network data, even if each data point is of high quality, it is impossible to determine the area where peak traffic flow is concentrated. Therefore, spatiotemporal data that does not meet the macro-quality requirements does not need to undergo further spatiotemporal data assessment steps, thus effectively reducing the time and resources spent on evaluating invalid data and micro-quality assessments, and improving the efficiency of quality assessment.
[0044] In a specific implementation, if the macro-quality assessment result is a dataset quality classification result, then the method for determining whether the spatiotemporal dataset to be evaluated meets the macro-quality requirements based on the macro-quality assessment result can be to directly determine whether the spatiotemporal dataset to be evaluated meets the macro-quality requirements based on the dataset quality classification result; if the macro-quality assessment result is a dataset quality score, then the method for determining whether the spatiotemporal dataset to be evaluated meets the macro-quality requirements based on the macro-quality assessment result can be to compare the dataset quality score with a preset dataset quality score threshold, and determine whether the spatiotemporal dataset to be evaluated meets the macro-quality requirements based on the comparison result. This embodiment does not impose any restrictions.
[0045] In a specific implementation, the method of determining whether the spatiotemporal dataset to be evaluated meets the macro quality requirements based on the macro quality assessment result can be to output the macro quality assessment result to the user terminal, obtain feedback information from the user based on the macro quality assessment result, and determine whether the macro quality assessment result meets the macro quality requirements based on the feedback information; the method of determining whether the spatiotemporal dataset to be evaluated meets the macro quality requirements based on the macro quality assessment result can also be to pre-set macro quality requirements and automatically determine whether the macro quality assessment result meets the macro quality requirements, and no limitation is made in this embodiment.
[0046] Optionally, the step of evaluating the dimensional quality assessment index of the spatiotemporal dataset to be evaluated to obtain the meso-level quality assessment result includes:
[0047] Step S31: Divide the spatiotemporal dataset to be evaluated into multiple first data subsets based on a preset first dimension;
[0048] In one feasible implementation, the spatiotemporal dataset to be evaluated is arranged according to a preset first dimension, and the arranged spatiotemporal dataset to be evaluated is divided into multiple first data subsets. The preset first dimension can be time, space, length, data volume, etc. The method of dividing the arranged spatiotemporal dataset to be evaluated into multiple first data subsets can be determined based on user operation, big data, statistical analysis results, etc., and is not limited in this embodiment. In a specific implementation, the spatiotemporal dataset to be evaluated, arranged in order, can be output on the interface, and the user's first segmentation operation based on the user interface can be detected. Based on the first segmentation operation, the arranged spatiotemporal dataset to be evaluated can be divided into multiple first data subsets. Alternatively, a first segmentation point can be preset according to the actual use and data volume of the spatiotemporal dataset to be evaluated, and the arranged spatiotemporal dataset to be evaluated can be divided into multiple first data subsets according to the preset first segmentation point. Alternatively, segmentation can be performed by combining user operation and a preset first segmentation point; this is not limited in this embodiment.
[0049] Step S32: Based on a preset first sampling ratio, at least one target first data subset is extracted from each of the first data subsets, and the dimensional quality assessment index of each target first data subset is evaluated to obtain the mesoscopic quality assessment result.
[0050] In one feasible implementation, based on a preset first sampling ratio, at least one target first data subset is extracted from each of the first data subsets, and the dimensional quality assessment index of each target first data subset is evaluated to obtain the initial meso-level quality assessment result corresponding to each target first data subset. The initial meso-level quality assessment results can be directly determined as the final meso-level quality assessment result, or the initial meso-level quality assessment results can be further combined, calculated, and analyzed to obtain the final meso-level quality assessment result. The preset first sampling ratio can be determined based on big data, actual needs, test results, etc., and is not limited in this embodiment.
[0051] In this embodiment, the evaluation is conducted by sampling, which can effectively improve the efficiency of meso-level quality assessment.
[0052] Step S40: If the dimensional quality assessment index of the spatiotemporal dataset to be evaluated meets the meso-quality requirements based on the meso-quality assessment result, then the data quality assessment index of the spatiotemporal dataset to be evaluated is evaluated to obtain the micro-quality assessment result.
[0053] In this embodiment, micro-quality assessment refers to the process of calculating, statistically analyzing, and / or evaluating the attribute values of at least one piece of spatiotemporal data in the spatiotemporal dataset to obtain data quality assessment index values for the dataset. These data quality assessment indexes characterize the quality of a single piece of data in the dataset, including data uniqueness and data consistency. Uniqueness characterizes whether data items, combinations of data items, data labels, etc., are repeated in the spatiotemporal data. This can be characterized by calculating the proportion of repeated spatiotemporal data; the smaller the proportion of repeated spatiotemporal data, the higher the uniqueness, and the fewer the repeated spatiotemporal data. For example, the uniqueness can be at least one of the following: uniqueness of data items, uniqueness of data item combinations, and uniqueness of data tags (e.g., a spatiotemporal data point will not belong to both the trajectory point of car A and the trajectory point of car B); the consistency is used to characterize whether the attribute values of a spatiotemporal data point are logically consistent, which can be characterized by calculating the proportion of spatiotemporal data points that are not logically consistent. The smaller the proportion of logically incorrect spatiotemporal data points, the higher the consistency, and the more logically consistent spatiotemporal data points there are. For example, if the same spatiotemporal data point has both the attribute value that the license plate color is green and the license plate number is 7 digits, then it is not logically consistent.
[0054] In one feasible implementation, at least one data quality assessment indicator for the spatiotemporal dataset to be evaluated can be determined in advance based on actual needs or user selection. The meso-level quality assessment results determine whether the spatiotemporal dataset meets the meso-level quality requirements. If the meso-level quality assessment results indicate that the dimensional quality assessment indicators of the spatiotemporal dataset meet the meso-level quality requirements, statistical analysis can be performed on all attribute values corresponding to one, multiple, or all of the spatiotemporal data to be evaluated in the dataset. The data quality assessment indicator values for each data point can be determined. These data quality assessment indicator values can be used as micro-level quality assessment results. Alternatively, data quality scoring or classification analyses can be performed on each spatiotemporal data and / or the dataset to be evaluated based on the data quality assessment indicator values. The resulting data quality scores or classification results can be used as micro-level quality assessment results. The data quality classification results can include meeting micro-level quality requirements, not meeting micro-level quality requirements, meeting micro-level quality requirements but requiring data processing, and can also include specific information related to the micro-level quality assessment. For example, if the data uniqueness score is 95 and the data consistency score is 97, the data uniqueness score of 95 and the data consistency score of 97 can be used as the micro-quality assessment result, or the average score of 96 or the total score of 192 can be used as the micro-quality assessment result.
[0055] If the meso-level quality assessment results determine that the spatiotemporal dataset to be evaluated does not meet the meso-level quality requirements, the process can return to the step of obtaining the spatiotemporal dataset to be evaluated, and the dataset can be obtained again. A prompt message can also be output to remind the user that the currently obtained spatiotemporal dataset does not meet the meso-level quality requirements and that it can be obtained again. Regardless of the micro-level quality assessment results, spatiotemporal data that does not meet the meso-level quality requirements cannot meet the actual needs of spatiotemporal data applications. For example, the data may have excessive fluctuations, too many missing data points, or too many noisy data points. Even if the consistency and uniqueness of each data point are high, the overall accuracy of the data will be low, potentially leading to excessive final errors. Therefore, spatiotemporal data that does not meet the meso-level quality requirements does not need to undergo further spatiotemporal data evaluation steps, thus effectively reducing the time and resources spent on evaluating invalid data and improving the efficiency of quality assessment.
[0056] For example, if the meso-quality assessment result is a dimensional quality classification result, then the method for determining whether the spatiotemporal dataset to be evaluated meets the meso-quality requirements based on the meso-quality assessment result can be to directly determine whether the spatiotemporal dataset to be evaluated meets the meso-quality requirements based on the dimensional quality classification result; if the meso-quality assessment result is a meso-quality score, then the method for determining whether the spatiotemporal dataset to be evaluated meets the meso-quality requirements based on the meso-quality assessment result can be to compare the meso-quality score with a preset meso-quality score threshold, and determine whether the spatiotemporal dataset to be evaluated meets the meso-quality requirements based on the comparison result. This embodiment does not impose any restrictions.
[0057] In a specific implementation, the method of determining whether the spatiotemporal dataset to be evaluated meets the meso-quality requirements based on the meso-quality assessment result can be to output the meso-quality assessment result to the user terminal, obtain feedback information from the user based on the meso-quality assessment result, and determine whether the meso-quality assessment result meets the meso-quality requirements based on the feedback information; the method of determining whether the spatiotemporal dataset to be evaluated meets the meso-quality requirements based on the meso-quality assessment result can also be to pre-set the meso-quality requirements and automatically determine whether the meso-quality assessment result meets the meso-quality requirements, and no limitation is made in this embodiment.
[0058] Optionally, the step of evaluating the data quality assessment indicators of the spatiotemporal dataset to be evaluated to obtain the micro-quality assessment results includes:
[0059] Step S41: Divide the spatiotemporal dataset to be evaluated into multiple second data subsets based on a preset second dimension;
[0060] In one feasible implementation, the spatiotemporal dataset to be evaluated is arranged according to a preset second dimension, and the arranged spatiotemporal dataset to be evaluated is divided into multiple second data subsets. The preset second dimension can be time, space, length, data volume, etc. The preset second dimension can be the same as or different from the preset first dimension, and can be determined according to actual needs. In this embodiment, no limitation is imposed. The method of dividing the arranged spatiotemporal dataset to be evaluated into multiple second data subsets can be determined according to user operations, big data, statistical analysis results, etc., and is not limited in this embodiment.
[0061] In a specific implementation, the spatiotemporal dataset to be evaluated, arranged in order, can be output in the user interface. A second segmentation operation by the user based on the user interface can be detected, and the arranged spatiotemporal dataset to be evaluated can be segmented into multiple second data subsets based on the second segmentation operation. Alternatively, a second segmentation point can be set in advance according to the actual use of the spatiotemporal dataset to be evaluated, and the arranged spatiotemporal dataset to be evaluated can be segmented into multiple second data subsets according to the preset second segmentation point. Alternatively, segmentation can be performed by combining user operation and preset second segmentation point, which is not limited in this embodiment.
[0062] Step S42: Based on a preset second sampling ratio, at least one target spatiotemporal data to be evaluated is extracted from each of the second data subsets, and the data quality evaluation indicators of each target spatiotemporal data to be evaluated are evaluated to obtain micro-quality evaluation results.
[0063] In one feasible implementation, based on a preset second sampling ratio, at least one target spatiotemporal data point to be evaluated is extracted from each of the second data subsets. The data quality evaluation indicators of each target spatiotemporal data point are evaluated to obtain the initial micro-quality evaluation results corresponding to each target spatiotemporal data point. The initial micro-quality evaluation results can be directly determined as the final micro-quality evaluation results, or the initial micro-quality evaluation results can be further combined, calculated, and analyzed to obtain the final micro-quality evaluation results. The preset second sampling ratio can be determined based on big data, actual needs, test results, etc. The preset second sampling ratio can be the same as or different from the first sampling ratio, and is not limited in this embodiment.
[0064] In this embodiment, the evaluation is carried out by sampling, which can effectively improve the efficiency of microscopic quality assessment.
[0065] Step S50: If the data quality assessment index of the spatiotemporal dataset to be evaluated meets the micro-quality requirements based on the micro-quality assessment result, then the dataset quality assessment result that has passed the assessment is output.
[0066] In one feasible implementation, the micro-quality assessment results determine whether the spatiotemporal dataset to be assessed meets the micro-quality requirements. If the micro-quality assessment results determine that the data quality assessment indicators of the spatiotemporal dataset to be assessed meet the micro-quality requirements, the dataset quality assessment results that have passed the assessment can be output. After outputting the dataset quality assessment results that have passed the assessment, the spatiotemporal dataset to be assessed can be used for modeling, data analysis, etc.
[0067] If the micro-quality assessment results determine that the spatiotemporal dataset to be evaluated does not meet the micro-quality requirements, the process can return to the step of obtaining the spatiotemporal dataset to be evaluated and re-obtain it. Alternatively, a prompt message can be output to remind the user that the currently obtained spatiotemporal dataset to be evaluated does not meet the meso-quality requirements and that the dataset can be re-obtained.
[0068] For example, if the micro-quality assessment result is a data quality classification result, then the method of determining whether the spatiotemporal dataset to be evaluated meets the micro-quality requirements based on the micro-quality assessment result can be to directly determine whether the spatiotemporal dataset to be evaluated meets the micro-quality requirements based on the data quality classification result; if the micro-quality assessment result is a micro-quality score, then the method of determining whether the spatiotemporal dataset to be evaluated meets the micro-quality requirements based on the micro-quality assessment result can be to compare the micro-quality score with a preset micro-quality score threshold, and determine whether the spatiotemporal dataset to be evaluated meets the micro-quality requirements based on the comparison result. This embodiment does not impose any restrictions.
[0069] In a specific implementation, the method of determining whether the spatiotemporal dataset to be evaluated meets the micro-quality requirements based on the micro-quality assessment result can be to output the micro-quality assessment result to the user terminal, obtain feedback information from the user based on the micro-quality assessment result, and determine whether the micro-quality assessment result meets the micro-quality requirements based on the feedback information; the method of determining whether the spatiotemporal dataset to be evaluated meets the micro-quality requirements based on the micro-quality assessment result can also be to pre-set the micro-quality requirements and automatically determine whether the micro-quality assessment result meets the micro-quality requirements, and no limitation is made in this embodiment.
[0070] In a specific implementation, the spatiotemporal data evaluation method further includes: if it is determined that the data quality evaluation indicators of the spatiotemporal dataset to be evaluated do not meet the macro-quality requirements, meso-quality requirements, and / or micro-quality requirements, then at least one unprocessed anomaly in the spatiotemporal dataset to be evaluated can be detected, and data processing can be performed on the spatiotemporal dataset to be evaluated according to each unprocessed anomaly to obtain an optimized spatiotemporal dataset, thereby improving the quality of the spatiotemporal dataset to be evaluated. For example, refer to... Figure 2 and Figure 3 First, the spatiotemporal data to be evaluated is collected using acquisition equipment and stored in a database. The spatiotemporal data evaluation equipment or device acquires the spatiotemporal dataset to be evaluated stored in the database. First, a macro-quality assessment is performed on the dataset, calculating usability indicators. If the dataset quality assessment indicators meet the macro-quality requirements, the data is deemed usable. Then, a meso-quality assessment is performed. If the dataset quality assessment indicators meet the meso-quality requirements, the data is deemed usable. Finally, a micro-quality assessment is performed. If the dataset quality assessment indicators meet the micro-quality requirements, the data is deemed usable. Data preprocessing can be performed on the dataset, or not, to complete the data preprocessing. The preprocessed spatiotemporal dataset can then be used for modeling. If it is determined that the dataset quality assessment indicators of the spatiotemporal dataset to be evaluated do not meet the macro-quality requirements, meso-quality requirements, and / or micro-quality requirements, the data is deemed unusable, and the current data preprocessing can be terminated, and a new spatiotemporal dataset to be evaluated can be obtained. After determining that the data is unusable, data monitoring can be performed on the spatiotemporal dataset to be evaluated to detect whether there are any unprocessed anomalies that can be improved in the dataset. If anomalies are detected, the dataset to be evaluated is processed to obtain an optimized spatiotemporal dataset, which is then used to replace the dataset to be evaluated for subsequent quality assessment. If no unprocessed anomalies are detected, the process returns to the step of obtaining the spatiotemporal dataset to be evaluated stored in the database, and a new dataset to be evaluated is obtained.
[0071] In this embodiment, by acquiring a spatiotemporal dataset to be evaluated, wherein the dataset includes multiple spatiotemporal data points, each with attribute values in multiple dimensions, the dataset quality evaluation indicators of the dataset are evaluated to obtain a macro-quality evaluation result, thus achieving a macro-quality evaluation of the dataset. Furthermore, if the dataset quality evaluation indicators of the dataset meet the macro-quality requirements based on the macro-quality evaluation results, the dimensional quality evaluation indicators of the dataset are then evaluated to obtain a meso-quality evaluation result. This allows for further evaluation of the dimensional quality of the dataset after the dataset quality evaluation indicators meet the macro-quality requirements. The evaluation metrics are assessed, and then, if the dimensional quality evaluation metrics of the spatiotemporal dataset to be evaluated meet the meso-level quality requirements based on the meso-level quality evaluation results, the data quality evaluation metrics of the spatiotemporal dataset to be evaluated are evaluated to obtain micro-level quality evaluation results. This achieves a multi-level progressive quality evaluation of the spatiotemporal dataset to be evaluated, from macro to meso to micro, and from dataset to dimension to data, after the dimensional quality evaluation metrics of the spatiotemporal dataset to be evaluated meet the meso-level quality requirements based on the micro-level quality evaluation results. The quality evaluation is divided into three levels: macro-level quality evaluation, meso-level quality evaluation, and micro-level quality evaluation. From macro to micro, the evaluation granularity decreases layer by layer, while the amount of data required for quality evaluation increases layer by layer, as does the algorithm complexity and processing performance. On the one hand, compared to methods that perform comprehensive macro-, meso-, and micro-quality assessments of spatiotemporal data before determining whether they meet quality requirements, this invention employs a progressive, layer-by-layer assessment, moving from macro to meso to micro, and from dataset to dimension to data. This eliminates spatiotemporal datasets that do not meet the higher-level quality requirements with relatively little time and resources. Only those datasets that meet the higher-level requirements can undergo lower-level quality assessment. Therefore, this effectively reduces the time and resources required for lower-level quality assessments of datasets that do not meet the higher-level requirements, thus improving the efficiency of quality assessment. On the other hand, compared to methods that improve efficiency by reducing the number of quality assessment dimensions, this invention ensures that the final output spatiotemporal dataset fully meets macro, meso, and micro-level quality requirements. This guarantees the accuracy of the quality assessment and overcomes the technical challenge of balancing accuracy and efficiency when assessing massive amounts of spatiotemporal data.
[0072] Furthermore, based on the above embodiments, another embodiment of the spatiotemporal data evaluation method of the present invention is proposed, referring to... Figure 3 , Figure 3 This is a flowchart illustrating an embodiment of steps S60 to S70 of the spatiotemporal data evaluation method of the present invention. In this embodiment, after step S50, the method further includes:
[0073] Step S60: If it is determined that data processing is required for the spatiotemporal dataset to be evaluated based on the meso-quality assessment results and the micro-quality assessment results, then at least one unprocessed anomaly exists in the spatiotemporal dataset to be evaluated, and data processing is performed on the spatiotemporal dataset to be evaluated based on each unprocessed anomaly to obtain an optimized spatiotemporal dataset.
[0074] In one feasible implementation, the micro-quality assessment results determine whether the spatiotemporal dataset to be evaluated meets the micro-quality requirements. If the micro-quality assessment results determine that the data quality assessment indicators of the spatiotemporal dataset to be evaluated meet the micro-quality requirements, then the meso-quality assessment results and the micro-quality assessment results determine whether to process the spatiotemporal dataset to be evaluated. If it is determined that the spatiotemporal dataset to be evaluated should be processed, then unprocessed anomalies in the spatiotemporal dataset to be evaluated are detected, and at least one unprocessed anomaly is identified in the spatiotemporal dataset to be evaluated. Based on the preset mapping relationship between the anomalies and the data processing algorithms, the target data processing algorithm corresponding to each unprocessed anomaly is determined. The spatiotemporal dataset to be evaluated is processed according to each target data processing algorithm to obtain an optimized spatiotemporal dataset.
[0075] In this embodiment, before determining whether to process the spatiotemporal dataset to be evaluated, a macro-quality assessment, a meso-quality assessment, and a micro-quality assessment have been performed on the spatiotemporal dataset to be evaluated. If the macro-quality assessment, the meso-quality assessment, and the micro-quality assessment are all full assessments of the spatiotemporal dataset to be evaluated, then all unprocessed anomalies in the spatiotemporal dataset to be evaluated can be directly determined based on the macro-quality assessment results, the meso-quality assessment results, and the micro-quality assessment results. If the meso-quality assessment and / or the micro-quality assessment are sampling assessments of the spatiotemporal dataset to be evaluated, then each spatiotemporal dataset to be evaluated can be detected sequentially in a certain order to determine all unprocessed anomalies in the spatiotemporal dataset to be evaluated.
[0076] In one feasible implementation, if it is determined that no data processing is performed on the spatiotemporal dataset to be evaluated, a qualified dataset quality assessment result can be output. After outputting the qualified dataset quality assessment result, the spatiotemporal dataset to be evaluated can be used for modeling, data analysis, etc. The output spatiotemporal dataset to be evaluated is used for modeling, data analysis, etc. The results of modeling and data analysis are unknown. The less data processing is performed on the spatiotemporal dataset to be evaluated, the higher the data authenticity, and the closer the subsequent modeling and data analysis results are to the true results.
[0077] In a specific implementation, the method for determining whether to process the spatiotemporal dataset to be evaluated based on the meso-quality assessment results and the micro-quality assessment results can be as follows: outputting the meso-quality assessment results and the micro-quality assessment results in the user interface, detecting the user's data processing operation based on the user interface, and determining whether to process the spatiotemporal dataset to be evaluated based on the data processing operation.
[0078] In a specific implementation, the method of determining whether to process the spatiotemporal dataset to be evaluated based on the meso-level quality assessment results and the micro-level quality assessment results can also be to detect whether there are any manageable quality problems among the quality problems in the meso-level quality assessment results and the micro-level quality assessment results. The manageable quality problems can be preset according to actual needs; for example, they may include noise points, missing data, etc. If manageable quality problems are detected among the quality problems in the meso-level quality assessment results and the micro-level quality assessment results, then based on the detailed information of the manageable quality problems, the processing method is determined. The process determines whether the quality issues meet preset data processing conditions. These preset conditions can be determined based on threshold data processing algorithms and actual needs. For example, when performing data completion processing using a data completion algorithm, it may not be suitable to complete large-scale data gaps. Therefore, a completion threshold can be set. If the range of missing data corresponding to a processable quality issue does not exceed the completion threshold, then the preset data processing conditions are met. If no processable quality issues are detected in the meso-level quality assessment results and the micro-level quality assessment results, then it is determined that no data processing will be performed on the spatiotemporal dataset to be evaluated. Furthermore, if it is determined that the processable quality issues meet the preset data processing conditions, then it is determined that data processing will be performed on the spatiotemporal dataset to be evaluated; if it is determined that the processable quality issues do not meet the preset data processing conditions, then it is determined that no data processing will be performed on the spatiotemporal dataset to be evaluated.
[0079] Step S70: Output the optimized spatiotemporal dataset.
[0080] In one feasible implementation, the optimized spatiotemporal dataset is output, and the output optimized spatiotemporal dataset can be used for modeling, data analysis, etc.
[0081] For example, the step of outputting the optimized spatiotemporal dataset may further include: performing a quality re-evaluation on the optimized spatiotemporal dataset to obtain a quality re-evaluation result, and outputting the optimized spatiotemporal dataset and the quality re-evaluation result. The quality re-evaluation may include macro-level quality evaluation, meso-level quality evaluation, and micro-level quality evaluation. Generally, the quality re-evaluation result is better than the quality evaluation result; therefore, the probability of the quality re-evaluation failing to meet quality requirements is low. Thus, the quality re-evaluation process does not need to be conducted in a hierarchical manner, but rather a comprehensive re-evaluation of the processed optimized spatiotemporal dataset to allow users to understand the specific quality status of the processed optimized spatiotemporal dataset.
[0082] In this embodiment, the amount of spatiotemporal data is large. Even if the spatiotemporal dataset to be evaluated meets the macroscopic, mesoscopic, and microscopic quality requirements, it is usually not all of high quality. That is, the spatiotemporal dataset to be evaluated that meets the microscopic quality requirements may still contain low-quality spatiotemporal data. In this regard, this embodiment improves the quality of the output spatiotemporal data by processing these low-quality spatiotemporal data, so as to help users solve problems in a timely manner after discovering them.
[0083] Furthermore, based on the above embodiments, another embodiment of the spatiotemporal data evaluation method of the present invention is proposed, referring to... Figure 4 , Figure 4 This is a flowchart illustrating an embodiment of step S60 of the spatiotemporal data evaluation method of the present invention. In this embodiment, the abnormal situations to be processed include abnormal situations to be deleted, abnormal situations to be completed, and abnormal situations to be adjusted. Step S60 includes:
[0084] Step A10: If at least one abnormal situation to be deleted is detected in the spatiotemporal dataset to be evaluated, the spatiotemporal dataset to be evaluated is processed according to a preset data deletion algorithm to obtain a first intermediate dataset.
[0085] In one feasible implementation, the anomalies to be deleted in the spatiotemporal dataset to be evaluated are detected. If at least one anomaly to be deleted is detected in the spatiotemporal dataset to be evaluated, the dataset is processed according to a preset data deletion algorithm to obtain a first intermediate dataset. The anomalies to be deleted refer to anomalies that need to be resolved by deleting data, such as noise points or duplicate data. They can be detected according to a preset anomaly detection algorithm, such as a noise point detection algorithm or a duplicate data detection algorithm. The data deletion algorithm refers to an algorithm for deleting data, such as an algorithm for deleting noise points.
[0086] If no deletion processing exception is detected in the spatiotemporal dataset to be evaluated, the spatiotemporal dataset to be evaluated can be directly determined as the first intermediate dataset for subsequent detection of exceptions to be completed.
[0087] In a specific implementation, refer to Figure 6 The spatiotemporal dataset to be evaluated includes driving trajectory data collected by vehicle GPS sensors. The steps for detecting noise points include: calculating the speed interval, distance interval, sampling interval, etc. between any two adjacent trajectory points of each vehicle based on the driving trajectory data; then using box plots or normal distribution algorithms to find trajectory points with abnormal speeds and / or abnormal distances; comparing the parameter values of these trajectory points with abnormal conditions with the corresponding preset parameter thresholds; and filtering out noise points whose values exceed the preset thresholds.
[0088] Step A20: If at least one abnormal situation to be completed is detected in the first intermediate dataset, the first intermediate dataset is processed according to the preset data completion algorithm to obtain the second intermediate dataset.
[0089] In one feasible implementation, the imputation anomalies in the first intermediate dataset are detected. If at least one imputation anomaly is detected in the first intermediate dataset, the first intermediate dataset is processed according to a preset data imputation algorithm to obtain a second intermediate dataset. The imputation anomalies refer to anomalies that need to be resolved by adding new data, such as missing data. They can be detected according to a preset imputation anomaly detection algorithm, such as a missing data detection algorithm. The data imputation algorithm refers to an algorithm for imputing data, such as an algorithm for imputing missing data.
[0090] If no anomalies to be completed are detected in the first intermediate dataset, the first intermediate dataset can be directly identified as the second intermediate dataset for subsequent anomaly detection.
[0091] Step A30: If at least one abnormal value processing condition is detected in the second intermediate dataset, the second intermediate dataset is processed according to the preset data adjustment algorithm to obtain the optimized spatiotemporal dataset.
[0092] In one feasible implementation, anomalies to be optimized in the second intermediate dataset are detected. If at least one anomaly to be adjusted is detected in the second intermediate dataset, the second intermediate dataset is processed according to a preset data adjustment algorithm to obtain an optimized spatiotemporal dataset. The anomaly to be adjusted refers to anomalies that need to be resolved by modifying the values of the data, such as numerical fluctuations or numerical anomalies. It can be detected according to a preset anomaly detection algorithm, such as an anomaly value detection algorithm. The data adjustment algorithm refers to an algorithm that adjusts the values of the data, such as a smoothing algorithm.
[0093] If no abnormality in the processing of values to be adjusted is detected in the second intermediate dataset, the second intermediate dataset can be directly identified as the optimized spatiotemporal dataset.
[0094] For example, the spatiotemporal dataset to be evaluated includes driving trajectory data collected by vehicle GPS sensors. Based on the driving trajectory data, trajectory parameters such as speed, distance, sampling interval, and trajectory length for each trajectory point of each vehicle can be calculated. Abnormal and duplicate data can be identified by comparing these trajectory parameters. Noise points are identified by comparing the speed of each trajectory point with a preset speed threshold and the distance between each trajectory point and its adjacent trajectory points with a preset distance threshold. Abnormal, duplicate, and noise points are then deleted. Furthermore, the existence of missing data is determined by comparing the time interval between each trajectory point and its adjacent trajectory points with a preset time threshold. If missing data exists, it is first filled in, followed by frame rate and trajectory point coordinate smoothing. If no missing data exists, frame rate and trajectory point coordinate smoothing can be performed directly.
[0095] For example, refer to Figure 7The spatiotemporal dataset to be evaluated includes driving trajectory data collected by vehicle GPS sensors. Based on this data, trajectory parameters such as speed, distance, sampling interval, and trajectory length for each trajectory point of each vehicle can be calculated. Global thresholds for each trajectory parameter can then be determined using conventional statistical analysis methods. Furthermore, the size of a sliding window is determined based on the trajectory length and frame rate, and this window is used to divide the spatiotemporal dataset into multiple data subsets. Then, diagnostic processing is performed for issues such as time anomalies and duplicate records. Finally, the trajectory points within the sliding window are traversed, and the speed interval, distance interval, and time interval between adjacent trajectory points are calculated. The threshold values for each diagnostic processing method are confirmed by combining the full and partial data conditions. The partial data condition refers to the data condition of the data subsets after the sliding window segmentation, while the full data condition refers to the global data condition of the entire dataset. For data with speed intervals greater than the speed interval threshold and distance intervals greater than the distance interval threshold, noise point diagnosis processing is performed first, followed by comparison of interval times. For data with speed intervals greater than the speed interval threshold and distance intervals greater than the distance interval threshold, the time interval is directly compared. For data with speed intervals greater than the speed interval threshold, distance intervals greater than the distance interval threshold, and time intervals greater than the time interval threshold, data missing point diagnosis processing is performed first, followed by frame rate and coordinate point smoothing processing. For data that meets any one of the following conditions—speed interval less than or equal to the speed interval threshold, distance interval less than or equal to the distance interval threshold, and time interval less than or equal to the time interval threshold—frame rate and coordinate point smoothing processing is performed directly. After performing frame rate and coordinate point smoothing processing on any trajectory point, the process can return to traversing the trajectory points within the sliding window until all trajectory points within the sliding window have been traversed.
[0096] In this embodiment, currently, quality assessment results are typically output so that users can choose the data to be processed and the processing method after knowing the results. For example, if a user sees that the gap between two data points is too large through the user interface, they can input an operation command to fill in the gap between these two data points, thereby achieving data processing. However, since deleting data leads to a reduction in data, which may result in data loss, it increases user operations and may even lead to uncontrollable or excessively biased data processing results. Therefore, this embodiment automatically processes data deletion first, then processes data completion, and then determines whether data optimization is needed after data completion. This optimizes the data processing order and automates the data processing process. On the one hand, it can effectively avoid the re-emergence of previously resolved anomalies due to subsequent data processing, reducing unnecessary time and resources wasted on repeated detection and processing, and improving the smoothness and efficiency of data processing. On the other hand, it can also reduce user operations and provide convenience for users.
[0097] Furthermore, based on the above embodiments, another embodiment of the spatiotemporal data evaluation method of the present invention is proposed, referring to... Figure 8 , Figure 8 This is a flowchart illustrating an embodiment of step S60 in the spatiotemporal data evaluation method of the present invention. In this embodiment, step S60 includes:
[0098] Step B10: Sort the spatiotemporal data to be evaluated in the spatiotemporal dataset to be evaluated based on the preset target dimension;
[0099] In one feasible implementation, the target dimension and sorting method for sorting the spatiotemporal data in the dataset to be evaluated can be determined in advance according to actual needs or user selection. The spatiotemporal data in the dataset to be evaluated are sorted based on the attribute values of the preset target dimension. For example, the spatiotemporal data to be evaluated can be sorted in order from early to late according to the time attribute values, or the spatiotemporal data to be evaluated can be sorted in order from small to large according to the spatial attribute values (coordinate values).
[0100] Step B20: According to the preset statistical classification algorithm, the sorted spatiotemporal dataset to be evaluated is divided into multiple windows;
[0101] In one feasible implementation, at least one segmentation point is determined according to a preset statistical classification algorithm, and the spatiotemporal dataset to be evaluated is segmented from each segmentation point to obtain multiple windows. The spatiotemporal data to be evaluated in each window should conform to the same data pattern. The statistical classification algorithm refers to the method of classifying the data by performing statistical analysis on the data and then classifying the data according to the results of the statistical analysis. The specific method of statistical classification is similar to the prior art and will not be elaborated here.
[0102] For example, for vehicle trajectory data, the vehicle trajectory data can be arranged in chronological order, and the speed change and sampling interval change of the vehicle trajectory data can be determined by statistical analysis. Based on the speed change and sampling interval change, the segmentation time point is determined, and the vehicle trajectory data is segmented from each segmentation time point to obtain multiple windows, so as to divide the vehicle trajectory data whose speed and sampling interval fluctuate within a certain range around the mean into the same window.
[0103] In one feasible implementation, a data volume threshold can be preset for each window. If a window to be segmented is detected to have a data volume exceeding the threshold, it can be further segmented to avoid reducing data processing efficiency due to excessive data in a single window.
[0104] Step B30: Perform anomaly detection on the spatiotemporal data to be evaluated in each window, and determine at least one target window with at least one anomaly to be processed from each window.
[0105] In one feasible implementation, anomaly detection is performed on the spatiotemporal data to be evaluated in each window, and the window in which at least one anomaly is detected is determined as the target window. At least one target window is determined from each window.
[0106] Optionally, the step of detecting unprocessed anomalies in the spatiotemporal data to be evaluated in each window, and determining at least one target window containing at least one unprocessed anomaly from each window, includes:
[0107] Step B31: Traverse the spatiotemporal data to be evaluated in each window and calculate the anomaly detection parameters corresponding to each spatiotemporal data to be evaluated in each window.
[0108] In one feasible implementation, the spatiotemporal data to be evaluated in each window are traversed, and the anomaly detection parameters corresponding to each spatiotemporal data to be evaluated in each window are calculated. The anomaly detection parameters are parameters used to identify anomalies. They can be attribute values of the spatiotemporal data to be evaluated, or other parameters calculated based on the attribute values of the spatiotemporal data to be evaluated. For example, they can be speed, trajectory length, sampling interval time, etc. The specific parameters can be determined according to actual needs.
[0109] Step B32: Perform statistical analysis on the anomaly detection parameters corresponding to each window to determine the threshold values for the anomaly detection parameters corresponding to each window.
[0110] In one feasible implementation, statistical analysis is performed on the anomaly detection parameters corresponding to each window, and the threshold values for the anomaly detection parameters corresponding to each window are determined based on the results of the statistical analysis. The statistical analysis includes normal distribution analysis, box plot analysis, mean analysis, mode analysis, etc. The method of determining the threshold values based on the results of the statistical analysis is similar to existing technologies and will not be elaborated further here. Compared to determining the anomaly detection parameter threshold values based on all the spatiotemporal data to be evaluated in the spatiotemporal dataset, the anomaly detection parameter threshold values determined by statistical analysis of the spatiotemporal data to be evaluated within each window are more consistent with the data patterns within the window range, thus enabling more accurate detection of anomalies.
[0111] Step B33: Compare the anomaly detection parameters and anomaly detection parameter thresholds for each window respectively;
[0112] In one feasible implementation, for each window, the anomaly detection parameters and the anomaly detection parameter threshold can be compared separately, and the spatiotemporal data to be processed that exceed the anomaly detection parameter threshold can be determined from the spatiotemporal data to be evaluated in each window.
[0113] Step B34: If anomaly detection parameters exceeding the anomaly detection parameter threshold are detected in each of the spatiotemporal data to be evaluated, then the window corresponding to the spatiotemporal data to be processed is determined as the target window.
[0114] In one feasible implementation, if at least one piece of spatiotemporal data to be processed is detected from each of the spatiotemporal data to be evaluated, such that the anomaly detection parameter exceeds the anomaly detection parameter threshold, then the window corresponding to each piece of spatiotemporal data to be processed is determined as the target window.
[0115] Step B40: Perform data processing on the spatiotemporal data to be evaluated in each target window according to the corresponding unprocessed anomaly situation to obtain an optimized spatiotemporal dataset.
[0116] In one feasible implementation, data processing is performed on the spatiotemporal data to be evaluated in each of the target windows. The data processing process for the data to be evaluated in each target window includes determining the target data processing algorithm corresponding to each anomaly to be processed in the window, and processing the spatiotemporal data to be evaluated in the window according to each target data processing algorithm to obtain an optimized spatiotemporal dataset.
[0117] In specific implementations, anomaly detection and data processing can be performed in parallel on the spatiotemporal data to be evaluated in each window to improve data processing efficiency. For example, big data components, such as Spark, can be used to perform anomaly detection and data processing in parallel on the spatiotemporal data to be evaluated in each window.
[0118] Optionally, the preset target dimension includes a time dimension, the anomalies to be processed include data missing anomalies, and the step of processing the spatiotemporal data to be evaluated in each target window according to the anomalies to be processed corresponding to each target window to obtain an optimized spatiotemporal dataset includes:
[0119] Step B41: If it is determined that the change in the spatial attribute value of the spatiotemporal data to be evaluated corresponding to the data missing anomaly in the target window exceeds the preset dwell range, then calculate the average sampling time interval and average moving speed of the target window.
[0120] In this embodiment, the preset target dimension includes the time dimension, and the anomaly to be processed includes the data missing anomaly. The data missing anomaly refers to the situation where there may be a lack of spatiotemporal data between two adjacent spatiotemporal data to be evaluated in time sequence. For example, by comparing the time interval between each trajectory point and the adjacent trajectory point with a preset time threshold, if the time interval exceeds the preset time threshold, it can be determined that a data missing anomaly has been detected.
[0121] There are usually two reasons for data missing anomalies. One is that no data is collected for a period of time due to equipment failure or other reasons. The other is that a stationary situation occurs, such as when a vehicle stops at a certain location. In the case of stationary situations, the sampling time interval will be extended, and the spatial attribute values of the data collected each time may be the same or very close due to sampling errors.
[0122] In one feasible implementation, for a target window where a data missing anomaly is detected, it is further determined whether the change in the spatial attribute value of the spatiotemporal data to be evaluated corresponding to the data missing anomaly in the target window exceeds a preset dwell range. If it is determined that the change in the spatial attribute value of the spatiotemporal data to be evaluated corresponding to the data missing anomaly in the target window exceeds the preset dwell range, it indicates that the data missing is due to the lack of data collection and needs to be supplemented. Therefore, the average sampling time interval and average movement speed of the target window are calculated based on all the spatiotemporal datasets to be evaluated in the target window. If it is determined that the change in the spatial attribute value of the spatiotemporal data to be evaluated corresponding to the data missing anomaly in the target window does not exceed the preset dwell range, it indicates that the data missing is due to dwelling. In this case, no supplementation is required, but the data pattern before and after dwelling may change. Therefore, the target window can be directly divided into two windows from the dwelling point.
[0123] Step B42: Determine the completion time threshold based on the average moving speed and the completion distance threshold corresponding to the preset completion algorithm;
[0124] In one feasible implementation, a completion distance threshold corresponding to a preset completion algorithm is obtained, and the ratio of the completion distance threshold to the average moving speed is determined as the completion time threshold.
[0125] Step B43: If it is determined that the time interval corresponding to the spatiotemporal data to be evaluated corresponding to the data missing anomaly does not exceed the completion time threshold, then the spatiotemporal data to be evaluated in the target window is processed according to the preset completion algorithm to obtain an optimized spatiotemporal dataset.
[0126] In one feasible implementation, the time interval between the spatiotemporal data to be evaluated corresponding to the data missing anomaly and its adjacent spatiotemporal data to be evaluated is calculated. The time interval is compared with the completion time threshold. If the time interval does not exceed the completion time threshold, the spatiotemporal data to be evaluated corresponding to the data missing anomaly in the target window is processed according to a preset completion algorithm to obtain an optimized spatiotemporal dataset. If the time interval exceeds the completion time threshold, it indicates that the data missing anomaly has lasted for a long time. In this case, it is not advisable to process the data using the completion algorithm. Furthermore, since the data missing anomaly has lasted for a long time, the data before and after the missing data usually do not have the same data pattern. Therefore, the two endpoints of the time interval can be used as dividing points, and the target window can be divided into two windows according to the dividing points.
[0127] For example, refer to Figure 9 The spatiotemporal dataset to be evaluated includes driving trajectory data collected by vehicle GPS sensors. The steps for data processing in case of missing data include: calculating the average sampling time interval T1 and average moving speed V1 of the target window; determining whether the trajectory points fluctuate irregularly within a certain range; if the trajectory points fluctuate irregularly within a certain range, determining the dwell point based on the fluctuation of the trajectory points, and directly dividing the target window into two windows from the dwell point; if the trajectory points do not fluctuate irregularly within a certain range, calculating the time threshold Th1 based on the preset complete distance threshold and average vehicle speed; if T1 < Th1, calling the completion algorithm to complete the trajectory points according to the interval time; if T1 ≥ Th1, dividing the target window into two windows based on the time of the trajectory points.
[0128] In this embodiment, since spatiotemporal data is dynamic, the data patterns may differ at different stages of the entire spatiotemporal data's trajectory. For example, vehicle speeds vary with time and road segment. The speed of vehicles traveling on the same road segment is lower during rush hour and higher at night. Similarly, the speed of vehicles traveling in the same time period is lower in urban areas and higher on highways. Detecting a speed of 100 km / h in an urban area might be an anomaly, while detecting a speed of 100 km / h on a highway is not. Therefore, this embodiment divides the spatiotemporal dataset to be evaluated into multiple windows, and performs anomaly detection and data processing on the spatiotemporal data to be evaluated in each window. This allows the detected anomalies and corresponding data processing to adapt to the changes in spatiotemporal data, thereby improving the accuracy of data processing.
[0129] Furthermore, embodiments of the present invention also propose a spatiotemporal data evaluation device, the device comprising:
[0130] The acquisition module is used to acquire the spatiotemporal dataset to be evaluated, wherein the spatiotemporal dataset to be evaluated includes multiple spatiotemporal data to be evaluated, and each spatiotemporal data to be evaluated has attribute values in multiple dimensions;
[0131] The first evaluation module is used to evaluate the dataset quality evaluation indicators of the spatiotemporal dataset to be evaluated, and obtain macro-quality evaluation results.
[0132] The second evaluation module is used to evaluate the dimensional quality evaluation index of the spatiotemporal dataset to be evaluated if the dataset quality evaluation index of the dataset to be evaluated meets the macro quality requirements based on the macro quality evaluation result, and to obtain the meso quality evaluation result.
[0133] The third evaluation module is used to evaluate the data quality evaluation indicators of the spatiotemporal dataset to be evaluated if the dimensional quality evaluation indicators of the dataset to be evaluated meet the meso-quality requirements based on the meso-quality evaluation results, and to obtain the micro-quality evaluation results.
[0134] The output module is used to output the dataset quality assessment result if the data quality assessment index of the spatiotemporal dataset to be evaluated meets the micro-quality requirements based on the micro-quality assessment result.
[0135] Optionally, the second evaluation module is further used for:
[0136] The spatiotemporal dataset to be evaluated is divided into multiple first data subsets based on a preset first dimension;
[0137] Based on a preset first sampling ratio, at least one target first data subset is extracted from each of the first data subsets;
[0138] The dimensional quality assessment indicators of each of the first data subsets of the targets are evaluated to obtain the meso-level quality assessment results.
[0139] Optionally, the third evaluation module is further configured to:
[0140] The spatiotemporal dataset to be evaluated is divided into multiple second data subsets based on a preset second dimension;
[0141] Based on a preset second sampling ratio, at least one target spatiotemporal data point to be evaluated is extracted from each of the second data subsets.
[0142] The data quality assessment indicators of the spatiotemporal data to be evaluated for each target are evaluated to obtain the micro-quality assessment results.
[0143] Optionally, the spatiotemporal data evaluation device further includes a data processing module, which is used for:
[0144] If it is determined from the meso-quality assessment results and the micro-quality assessment results that data processing is required on the spatiotemporal dataset to be assessed, then at least one unprocessed anomaly in the spatiotemporal dataset to be assessed is detected, and data processing is performed on the spatiotemporal dataset to be assessed according to each unprocessed anomaly to obtain an optimized spatiotemporal dataset.
[0145] Output the optimized spatiotemporal dataset.
[0146] Optionally, the data processing module is used for:
[0147] If at least one abnormal situation to be deleted is detected in the spatiotemporal dataset to be evaluated, the spatiotemporal dataset to be evaluated is processed according to a preset data deletion algorithm to obtain a first intermediate dataset.
[0148] If at least one abnormal situation requiring completion is detected in the first intermediate dataset, the first intermediate dataset is processed according to a preset data completion algorithm to obtain the second intermediate dataset.
[0149] If at least one abnormal value processing condition is detected in the second intermediate dataset, the second intermediate dataset is processed according to the preset data adjustment algorithm to obtain an optimized spatiotemporal dataset.
[0150] Optionally, the data processing module is used for:
[0151] Based on a preset target dimension, the spatiotemporal data to be evaluated in the spatiotemporal dataset to be evaluated are sorted.
[0152] Based on the preset statistical classification algorithm, the sorted spatiotemporal dataset to be evaluated is divided into multiple windows;
[0153] Each window is used to detect unprocessed anomalies in the spatiotemporal data to be evaluated, and at least one target window with at least one unprocessed anomaly is identified from each window.
[0154] The spatiotemporal data to be evaluated in each of the target windows are processed according to the corresponding anomalies to be processed, so as to obtain an optimized spatiotemporal dataset.
[0155] Optionally, the data processing module is used for:
[0156] Iterate through the spatiotemporal data to be evaluated in each window and calculate the anomaly detection parameters corresponding to each spatiotemporal data to be evaluated in each window.
[0157] Statistical analysis is performed on the anomaly detection parameters corresponding to each window to determine the threshold values of the anomaly detection parameters corresponding to each window.
[0158] The anomaly detection parameters and anomaly detection parameter thresholds corresponding to each window are compared separately.
[0159] If anomaly detection parameters exceeding the anomaly detection parameter threshold are detected in the spatiotemporal data to be evaluated, then the window corresponding to the spatiotemporal data to be processed is determined as the target window.
[0160] Optionally, the data processing module is used for:
[0161] If it is determined that the change in the spatial attribute value of the spatiotemporal data to be evaluated corresponding to the data missing anomaly in the target window exceeds the preset dwell range, then the average sampling time interval and average moving speed of the target window are calculated.
[0162] The completion time threshold is determined based on the average moving speed and the completion distance threshold corresponding to the preset completion algorithm;
[0163] If it is determined that the time interval corresponding to the spatiotemporal data to be evaluated in the data missing anomaly does not exceed the completion time threshold, then the spatiotemporal data to be evaluated in the target window is processed according to the preset completion algorithm to obtain an optimized spatiotemporal dataset.
[0164] Furthermore, embodiments of the present invention also propose an electronic device, such as... Figure 10 As shown, Figure 10 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of the present invention. It should be noted that the electronic device in the embodiments of the present invention can be a smartphone, a personal computer, a server, or other devices, and no specific limitations are imposed here.
[0165] like Figure 10 As shown, the electronic device may include: a processor 1001, such as a CPU; a network interface 1004; a user interface 1003; a memory 1005; and a communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0166] Those skilled in the art will understand that Figure 5The device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0167] like Figure 5 As shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a spatiotemporal data evaluation program. The operating system is a program that manages and controls the device's hardware and software resources, supporting the operation of the spatiotemporal data evaluation program and other software or programs. Figure 5 In the device shown, the user interface 1003 is mainly used for data communication with the client; the network interface 1004 is mainly used for establishing a communication connection with the server; and the processor 1001 can be used to call the spatiotemporal data evaluation program stored in the memory 1005 and perform the following operations:
[0168] Obtain the spatiotemporal dataset to be evaluated, wherein the spatiotemporal dataset to be evaluated includes multiple spatiotemporal data to be evaluated, and each spatiotemporal data to be evaluated has attribute values of multiple dimensions;
[0169] The dataset quality evaluation indicators of the spatiotemporal dataset to be evaluated are used to obtain macroscopic quality evaluation results;
[0170] If the dataset quality assessment index of the spatiotemporal dataset to be assessed meets the macro quality requirements based on the macro quality assessment results, then the dimensional quality assessment index of the spatiotemporal dataset to be assessed is evaluated to obtain the meso quality assessment results.
[0171] If the dimensional quality assessment index of the spatiotemporal dataset to be evaluated is determined to meet the meso-level quality requirements based on the meso-level quality assessment results, then the data quality assessment index of the spatiotemporal dataset to be evaluated is assessed to obtain the micro-level quality assessment results.
[0172] If the data quality assessment indicators of the spatiotemporal dataset to be evaluated meet the micro-quality requirements based on the micro-quality assessment results, then the dataset quality assessment result that has passed the assessment is output.
[0173] Furthermore, the processor 1001 can also be used to call the spatiotemporal data evaluation program stored in the memory 1005 to perform the following operations:
[0174] The spatiotemporal dataset to be evaluated is divided into multiple first data subsets based on a preset first dimension;
[0175] Based on a preset first sampling ratio, at least one target first data subset is extracted from each of the first data subsets;
[0176] The dimensional quality assessment indicators of each of the first data subsets of the targets are evaluated to obtain the meso-level quality assessment results.
[0177] Furthermore, the processor 1001 can also be used to call the spatiotemporal data evaluation program stored in the memory 1005 to perform the following operations:
[0178] The spatiotemporal dataset to be evaluated is divided into multiple second data subsets based on a preset second dimension;
[0179] Based on a preset second sampling ratio, at least one target spatiotemporal data point to be evaluated is extracted from each of the second data subsets.
[0180] The data quality assessment indicators of the spatiotemporal data to be evaluated for each target are evaluated to obtain the micro-quality assessment results.
[0181] Furthermore, after the step of outputting the dataset quality assessment result indicating that the data quality assessment index of the spatiotemporal dataset to be evaluated meets the micro-quality requirements based on the micro-quality assessment result, the processor 1001 can also be used to call the spatiotemporal data assessment program stored in the memory 1005 to perform the following operations:
[0182] If it is determined from the meso-quality assessment results and the micro-quality assessment results that data processing is required on the spatiotemporal dataset to be assessed, then at least one unprocessed anomaly in the spatiotemporal dataset to be assessed is detected, and data processing is performed on the spatiotemporal dataset to be assessed according to each unprocessed anomaly to obtain an optimized spatiotemporal dataset.
[0183] Output the optimized spatiotemporal dataset.
[0184] Furthermore, the processor 1001 can also be used to call the spatiotemporal data evaluation program stored in the memory 1005 to perform the following operations:
[0185] If at least one abnormal situation to be deleted is detected in the spatiotemporal dataset to be evaluated, the spatiotemporal dataset to be evaluated is processed according to a preset data deletion algorithm to obtain a first intermediate dataset.
[0186] If at least one abnormal situation requiring completion is detected in the first intermediate dataset, the first intermediate dataset is processed according to a preset data completion algorithm to obtain the second intermediate dataset.
[0187] If at least one abnormal value processing condition is detected in the second intermediate dataset, the second intermediate dataset is processed according to the preset data adjustment algorithm to obtain an optimized spatiotemporal dataset.
[0188] Furthermore, the processor 1001 can also be used to call the spatiotemporal data evaluation program stored in the memory 1005 to perform the following operations:
[0189] Based on a preset target dimension, the spatiotemporal data to be evaluated in the spatiotemporal dataset to be evaluated are sorted.
[0190] Based on the preset statistical classification algorithm, the sorted spatiotemporal dataset to be evaluated is divided into multiple windows;
[0191] Each window is used to detect unprocessed anomalies in the spatiotemporal data to be evaluated, and at least one target window with at least one unprocessed anomaly is identified from each window.
[0192] The spatiotemporal data to be evaluated in each of the target windows are processed according to the corresponding anomalies to be processed, so as to obtain an optimized spatiotemporal dataset.
[0193] Furthermore, the processor 1001 can also be used to call the spatiotemporal data evaluation program stored in the memory 1005 to perform the following operations:
[0194] Iterate through the spatiotemporal data to be evaluated in each window and calculate the anomaly detection parameters corresponding to each spatiotemporal data to be evaluated in each window.
[0195] Statistical analysis is performed on the anomaly detection parameters corresponding to each window to determine the threshold values of the anomaly detection parameters corresponding to each window.
[0196] The anomaly detection parameters and anomaly detection parameter thresholds corresponding to each window are compared separately.
[0197] If anomaly detection parameters exceeding the anomaly detection parameter threshold are detected in the spatiotemporal data to be evaluated, then the window corresponding to the spatiotemporal data to be processed is determined as the target window.
[0198] Furthermore, the processor 1001 can also be used to call the spatiotemporal data evaluation program stored in the memory 1005 to perform the following operations:
[0199] If it is determined that the change in the spatial attribute value of the spatiotemporal data to be evaluated corresponding to the data missing anomaly in the target window exceeds the preset dwell range, then the average sampling time interval and average moving speed of the target window are calculated.
[0200] The completion time threshold is determined based on the average moving speed and the completion distance threshold corresponding to the preset completion algorithm;
[0201] If it is determined that the time interval corresponding to the spatiotemporal data to be evaluated in the data missing anomaly does not exceed the completion time threshold, then the spatiotemporal data to be evaluated in the target window is processed according to the preset completion algorithm to obtain an optimized spatiotemporal dataset.
[0202] Furthermore, embodiments of the present invention also propose a computer-readable storage medium storing a spatiotemporal data evaluation program, which, when executed by a processor, implements the steps of the spatiotemporal data evaluation method described below.
[0203] The embodiments of the electronic device and computer-readable storage medium of the present invention can all refer to the embodiments of the spatiotemporal data evaluation method of the present invention, and will not be repeated here.
[0204] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0205] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0206] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0207] The above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of protection of the present invention.
Claims
1. A spatiotemporal data evaluation method, characterized in that, include: Obtain the spatiotemporal dataset to be evaluated, wherein the spatiotemporal dataset to be evaluated includes multiple spatiotemporal data to be evaluated, each spatiotemporal data to be evaluated has attribute values of multiple dimensions, and the spatiotemporal data to be evaluated is vehicle driving trajectory data; The dataset quality evaluation indicators of the spatiotemporal dataset to be evaluated are evaluated to obtain macro-quality evaluation results. The dataset quality evaluation indicators are used to characterize the overall quality of all data in the spatiotemporal dataset to be evaluated, including data volume, timeliness and data type. If the dataset quality assessment index of the spatiotemporal dataset to be evaluated meets the macro quality requirements based on the macro quality assessment result, then the dimensional quality assessment index of the spatiotemporal dataset to be evaluated is evaluated to obtain the meso quality assessment result. The dimensional quality assessment index is used to characterize the overall quality of all attribute values of each dimension in the spatiotemporal dataset to be evaluated, including dimensional stability, dimensional completeness and dimensional accuracy. If the dimensional quality assessment index of the spatiotemporal dataset to be evaluated is determined to meet the meso-level quality requirements based on the meso-level quality assessment results, then the data quality assessment index of the spatiotemporal dataset to be evaluated is evaluated to obtain the micro-level quality assessment results. The data quality assessment index is used to characterize the quality of each data item in the spatiotemporal dataset to be evaluated, including data uniqueness and data consistency. If the data quality assessment indicators of the spatiotemporal dataset to be evaluated meet the micro-quality requirements based on the micro-quality assessment results, then the dataset quality assessment result that has passed the assessment is output. If it is determined from the meso-level quality assessment results and the micro-level quality assessment results that data processing is required on the spatiotemporal dataset to be assessed, then the spatiotemporal data to be assessed in the spatiotemporal dataset to be assessed is sorted based on a preset target dimension, the preset target dimension including the time dimension. Based on the preset statistical classification algorithm, the sorted spatiotemporal dataset to be evaluated is divided into multiple windows; Each window's spatiotemporal data to be evaluated is subjected to anomaly detection, and at least one target window with at least one anomaly to be processed is identified from each window; the anomalies to be processed include data missing anomalies. If it is determined that the change in the spatial attribute value of the spatiotemporal data to be evaluated corresponding to the data missing anomaly in the target window exceeds the preset dwell range, then the average sampling time interval and average moving speed of the target window are calculated. The completion time threshold is determined based on the average moving speed and the completion distance threshold corresponding to the preset completion algorithm; If it is determined that the time interval corresponding to the spatiotemporal data to be evaluated corresponding to the data missing anomaly does not exceed the completion time threshold, then the spatiotemporal data to be evaluated in the target window is processed according to the preset completion algorithm to obtain an optimized spatiotemporal dataset. If the time interval exceeds the completion time threshold, the two endpoints of the time interval are used as dividing points, and the target window is divided into two windows according to the dividing points.
2. The spatio-temporal data evaluation method of claim 1, wherein, The step of evaluating the dimensional quality assessment index of the spatiotemporal dataset to be evaluated to obtain the mesoscopic quality assessment result includes: The spatiotemporal dataset to be evaluated is divided into multiple first data subsets based on a preset first dimension; Based on a preset first sampling ratio, at least one target first data subset is extracted from each of the first data subsets; The dimensional quality assessment indicators of each of the first data subsets of the targets are evaluated to obtain the meso-level quality assessment results.
3. The spatiotemporal data evaluation method as described in claim 1, characterized in that, The step of evaluating the data quality assessment indicators of the spatiotemporal dataset to be evaluated and obtaining the micro-quality assessment results includes: The spatiotemporal dataset to be evaluated is divided into multiple second data subsets based on a preset second dimension; Based on a preset second sampling ratio, at least one target spatiotemporal data point to be evaluated is extracted from each of the second data subsets. The data quality assessment indicators of the spatiotemporal data to be evaluated for each target are evaluated to obtain the micro-quality assessment results.
4. The spatiotemporal data evaluation method as described in claim 1, characterized in that, After the step of outputting the dataset quality assessment result that has passed the assessment if the data quality assessment index of the spatiotemporal dataset to be assessed meets the micro-quality requirements based on the micro-quality assessment result, the method further includes: If it is determined from the meso-quality assessment results and the micro-quality assessment results that data processing is required on the spatiotemporal dataset to be assessed, then at least one unprocessed anomaly in the spatiotemporal dataset to be assessed is detected, and data processing is performed on the spatiotemporal dataset to be assessed according to each unprocessed anomaly to obtain an optimized spatiotemporal dataset. Output the optimized spatiotemporal dataset.
5. The spatiotemporal data evaluation method as described in claim 4, characterized in that, The unprocessed anomalies include unprocessed deletion anomalies, unprocessed completion anomalies, and unprocessed value adjustment anomalies. The step of detecting at least one unprocessed anomaly in the spatiotemporal dataset to be evaluated, and processing the spatiotemporal dataset to be evaluated according to each unprocessed anomaly to obtain an optimized spatiotemporal dataset includes: If at least one abnormal situation to be deleted is detected in the spatiotemporal dataset to be evaluated, the spatiotemporal dataset to be evaluated is processed according to a preset data deletion algorithm to obtain a first intermediate dataset. If at least one abnormal situation requiring completion is detected in the first intermediate dataset, the first intermediate dataset is processed according to a preset data completion algorithm to obtain the second intermediate dataset. If at least one abnormal value processing condition is detected in the second intermediate dataset, the second intermediate dataset is processed according to the preset data adjustment algorithm to obtain an optimized spatiotemporal dataset.
6. The spatiotemporal data evaluation method as described in claim 1, characterized in that, The step of detecting unprocessed anomalies in the spatiotemporal data to be evaluated in each window, and determining at least one target window containing at least one unprocessed anomaly from each window, includes: Iterate through the spatiotemporal data to be evaluated in each window and calculate the anomaly detection parameters corresponding to each spatiotemporal data to be evaluated in each window. Statistical analysis is performed on the anomaly detection parameters corresponding to each window to determine the threshold values of the anomaly detection parameters corresponding to each window. The anomaly detection parameters and anomaly detection parameter thresholds corresponding to each window are compared separately. If anomaly detection parameters exceeding the anomaly detection parameter threshold are detected in the spatiotemporal data to be evaluated, then the window corresponding to the spatiotemporal data to be processed is determined as the target window.
7. A spatiotemporal data evaluation device, characterized in that, include: The acquisition module is used to acquire the spatiotemporal dataset to be evaluated, wherein the spatiotemporal dataset to be evaluated includes multiple spatiotemporal data to be evaluated, each spatiotemporal data to be evaluated has attribute values of multiple dimensions, and the spatiotemporal data to be evaluated is vehicle driving trajectory data; The first evaluation module is used to evaluate the dataset quality evaluation indicators of the spatiotemporal dataset to be evaluated and obtain macro-quality evaluation results. The dataset quality evaluation indicators are used to characterize the overall quality of all data in the spatiotemporal dataset to be evaluated, including data volume, timeliness and data type. The second evaluation module is used to evaluate the dimensional quality evaluation index of the spatiotemporal dataset to be evaluated if the dataset quality evaluation index of the dataset to be evaluated meets the macro quality requirements based on the macro quality evaluation result, and to obtain the meso quality evaluation result. The dimensional quality evaluation index is used to characterize the overall quality of all attribute values of each dimension in the spatiotemporal dataset to be evaluated, including dimensional stability, dimensional completeness and dimensional accuracy. The third evaluation module is used to evaluate the data quality evaluation indicators of the spatiotemporal dataset to be evaluated if the dimensional quality evaluation indicators of the dataset to be evaluated meet the meso-quality requirements based on the meso-quality evaluation results, and to obtain micro-quality evaluation results. The data quality evaluation indicators are used to characterize the quality of each data in the spatiotemporal dataset to be evaluated, including data uniqueness and data consistency. The output module is used to output the dataset quality assessment result that has passed the assessment if the data quality assessment index of the spatiotemporal dataset to be assessed meets the micro-quality requirements based on the micro-quality assessment result. The data processing module is used for: If it is determined from the meso-level quality assessment results and the micro-level quality assessment results that data processing is required on the spatiotemporal dataset to be assessed, then the spatiotemporal data to be assessed in the spatiotemporal dataset to be assessed is sorted based on a preset target dimension, the preset target dimension including the time dimension. Based on the preset statistical classification algorithm, the sorted spatiotemporal dataset to be evaluated is divided into multiple windows; Each window's spatiotemporal data to be evaluated is subjected to anomaly detection, and at least one target window with at least one anomaly to be processed is identified from each window; the anomalies to be processed include data missing anomalies. If it is determined that the change in the spatial attribute value of the spatiotemporal data to be evaluated corresponding to the data missing anomaly in the target window exceeds the preset dwell range, then the average sampling time interval and average moving speed of the target window are calculated. The completion time threshold is determined based on the average moving speed and the completion distance threshold corresponding to the preset completion algorithm; If it is determined that the time interval corresponding to the spatiotemporal data to be evaluated corresponding to the data missing anomaly does not exceed the completion time threshold, then the spatiotemporal data to be evaluated in the target window is processed according to the preset completion algorithm to obtain an optimized spatiotemporal dataset. If the time interval exceeds the completion time threshold, the two endpoints of the time interval are used as dividing points, and the target window is divided into two windows according to the dividing points.
8. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a spatiotemporal data evaluation program stored in the memory and executable on the processor, wherein the spatiotemporal data evaluation program, when executed by the processor, implements the steps of the spatiotemporal data evaluation method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a spatiotemporal data evaluation program, which, when executed by a processor, implements the steps of the spatiotemporal data evaluation method as described in any one of claims 1 to 6.