Positioning data cleaning method and system based on data analysis

WO2026200433A1PCT designated stage Publication Date: 2026-10-01GUANGZHOU SEEWORLD TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/081116
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-03
Publication Date
2026-10-01

Smart Images

  • Figure CN2026081116_01102026_PF_FP_ABST
    Figure CN2026081116_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data processing, and in particular to a positioning data cleaning method and system based on data analysis. The method comprises: acquiring position data of a vehicle at various moments, the position data comprising longitude and latitude values; using a saliency detection CA algorithm to perform detection on the position data at the various moments, and obtaining saliency values at the various moments, comprising: determining longitude and latitude special values and longitude and latitude reliability at a target moment; acquiring the reliability of equally dividing, on the basis of the length of each preset scale, the position data into a plurality of data segments at each scale; and determining reference data segments and preliminary saliency values, and further obtaining the saliency values; and, on the basis of the magnitudes of the saliency values, performing interpolation processing, so as to complete positioning data cleaning based on data analysis. By means of calculating indicators such as the longitude and latitude special values and reliability, the present invention provides a more reliable basis for calculating the preliminary saliency values, effectively reducing interference, and improving the accuracy of invalid data determination.
Need to check novelty before this filing date? Find Prior Art

Description

A Location Data Cleaning Method and System Based on Data Analysis Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a location data cleaning method and system based on data analysis. Background Technology

[0002] With the development of big data technology, especially the widespread adoption of mobile internet and the Internet of Things (IoT), location data has become a core data type in many applications. Location data, as a type of positioning data, may contain noise, errors, and missing information due to various reasons. These problems may stem from unstable device hardware, sensor errors, or data loss during transmission. If these data errors are not cleaned, they will lead to distortion in downstream analysis, decision-making, and prediction, and may even affect important applications such as public safety and traffic management. Cleaning location data can improve its accuracy and completeness, thus providing more reliable information for relevant decisions. One existing method for detecting invalid location data utilizes the saliency detection (CA) algorithm, which has advantages such as strong real-time performance and robustness. It can quickly and accurately identify and filter invalid data in large-scale location data processing, significantly improving the reliability and accuracy of the positioning system.

[0003] Patent application CN114296063A discloses a cooperative positioning and ranging method based on CSMA / CA, including the following steps: S1: The transmitting station detects the state of the wireless channel and uses a backoff algorithm to choose whether to transmit the entire data frame; S2: The transmitting station uses a backoff timer to repeat and transmits a data frame to wait for confirmation; S3: When the backoff timer of a transmitting station counts down to 0, it uses an RTS / CTS data frame to detect whether the receiving station has received the data frame; S4: The transmitting station repeats step S1 and continues to execute the next frame.

[0004] However, the aforementioned patent application does not address the cleaning of location data and does not solve the problem that when using the saliency detection (CA) algorithm to analyze whether location data at a given moment is invalid, the determination is made by calculating the magnitude of its saliency value. Due to the existence of potentially invalid data, when calculating the initial saliency value of a data point at one scale, if invalid data exists in other data segments at the same scale, the data authenticity of the data segment to which the data point belongs is reduced, thereby affecting the accuracy of the initial saliency value of the data point at one scale and the saliency value at multiple scales, leading to errors in the determination of invalid data. Summary of the Invention

[0005] To address the issue that the saliency detection CA algorithm determines whether location data at each time point is invalid by calculating its saliency value, but the presence of potentially invalid data reduces the accuracy of the data segment, thus affecting the accuracy of the initial saliency value of the data point at one scale and the saliency value at multiple scales, leading to errors in the determination of invalid data, this invention provides a location data cleaning method and system based on data analysis.

[0006] In a first aspect, the present invention provides a location data cleaning method based on data analysis, employing the following technical solution:

[0007] A location data cleaning method based on data analysis includes: acquiring vehicle location data at various times, wherein the location data includes latitude and longitude values; and using saliency detection... The algorithm detects location data at each time point and obtains the saliency value for each time point; it then performs interpolation based on the magnitude of the saliency value to complete the location data cleaning based on data analysis; finally, it utilizes the saliency... During the detection process, the algorithm divides the location data into multiple data segments at each scale according to the preset length of each scale. It calculates the preliminary significance value of the target time at each scale and determines the significance value of the target time based on these preliminary significance values. This includes: determining the latitude and longitude special values ​​of the target time at each scale based on the latitude and longitude values ​​of the target time and several previous times, as well as the extreme values ​​of the latitude and longitude values ​​at each time; determining the latitude and longitude accuracy of the target time based on these special values ​​and the latitude and longitude values ​​of the target time and several previous times; using the sum of the latitude and longitude accuracy values ​​as the location accuracy of the target time; for data segments that do not contain the target time, using the sum of the location accuracy values ​​of all times within the data segment as the accuracy of the data segment; determining a reference data segment based on the magnitude of the accuracy; recording the average of the sum of the latitude and longitude special values ​​of all times within each data segment as the feature value of each data segment; and determining the preliminary significance value of the target time at the target scale based on the feature values ​​of the data segment to which the target time belongs and the reference data segment.

[0008] The beneficial effects are as follows: By combining the saliency detection CA algorithm with multi-scale analysis, not only can the accuracy of the data be improved, but invalid data can also be effectively filtered out through saliency value judgment and interpolation at each scale, reducing the possibility of misjudgment; the saliency detection process uses comparative analysis of the target time with the location data of several previous times, reducing the impact of invalid data on the overall authenticity of the data segment, thereby avoiding positioning errors caused by invalid data; by analyzing the location data at multiple scales, potential effective information in the data can be captured at different scales, and inaccurate or abnormal values ​​can be effectively isolated, enhancing the method's adaptability to data changes; by calculating the special values ​​of latitude and longitude of the target time and its adjacent times, and combining them with changes in their authenticity, the characteristic values ​​of the data segment are further determined, ultimately improving the comprehensive evaluation ability of time and data segment; judging the saliency value based on the characteristic values ​​of the data segment and the target time can more accurately filter out reliable data, making the final positioning result more accurate, thereby improving the stability and reliability of the positioning system in practical applications.

[0009] Furthermore, the specific values ​​of longitude and latitude respectively satisfy:

[0010] ,

[0011] In the formula, and They are time points The special values ​​of longitude and latitude, and They are time points The longitude and latitude values, and They are time points The average of the longitude and latitude values ​​at several previous moments. and These are the maximum values ​​of longitude and latitude at each time point. and These are the minimum longitude and latitude values ​​at each time point, respectively. For hyperparameters, For the standard normalized function, It is a minimum value function. It is the absolute value symbol.

[0012] The beneficial effects are as follows: by calculating the specific values ​​of latitude and longitude, the changes in location data can be reflected more accurately, thereby improving the accuracy of data cleaning; by considering the mean of historical data, the interference of instantaneous errors can be reduced, and the representativeness of the target time data can be improved; normalization and standardization operations eliminate the errors caused by differences in data range, ensuring that different data can be compared fairly; by calculating the difference between the maximum and minimum values, extreme data can be effectively identified and processed, reducing their negative impact on the results.

[0013] Furthermore, the accuracy of the latitude and longitude coordinates respectively satisfies:

[0014] ,

[0015] In the formula, and They are time points The accuracy of longitude and latitude and They are time points The special values ​​of longitude and latitude, and They are time points Compared with the variance of longitude and latitude values ​​at previous times, and They are time points The variance of longitude and the variance of latitude values ​​at previous times. and They are time points The difference in longitude and latitude compared to the previous moment.

[0016] The beneficial effects are as follows: by calculating the accuracy of latitude and longitude, the authenticity of positioning data at each moment can be accurately assessed, ensuring that the data is more reliable; based on the comparison of latitude and longitude variance, the credibility of the data can be dynamically adjusted, data with large fluctuations can be identified, and data processing can be optimized; by introducing differences between moments, the trend of location changes can be captured, which helps to identify and adjust abrupt changes in data and avoid their impact on the results; and by combining special values, variance, and differences for comprehensive evaluation, the depth and accuracy of data cleaning can be improved.

[0017] Furthermore, determining the reference data segment based on the magnitude of the realism includes: in response to the realism being greater than a preset reference threshold, identifying the data segment as a reference data segment.

[0018] Furthermore, the preliminary significance value satisfies:

[0019] In the formula, For a moment In scale The initial significance value, For a moment In scale The number of reference data segments below, For a moment In scale The feature values ​​of the data segment to which it belongs. For a moment In scale The next Feature values ​​of a reference data segment It is the absolute value symbol.

[0020] The beneficial effects are as follows: by calculating the preliminary significance value, it is possible to effectively assess the significant changes in data at different scales at the target time, and help identify key changes or anomalies in the data; by comparing with multiple reference data segments, it improves the accuracy and robustness of significance assessment and avoids the bias of a single reference data; by averaging the differences, it effectively reduces the impact of a single outlier on the significance calculation, and ensures that the assessment results are more stable and reliable.

[0021] Furthermore, the significance value satisfies:

[0022] In the formula, For a moment The significance value, For the number of scales, For a moment In scale The initial significance value.

[0023] The beneficial effects are as follows: by averaging the preliminary significance values ​​at multiple scales, the significance of the target time can be evaluated more comprehensively, avoiding the bias of a single scale; by smoothing the significance values ​​at each scale, the impact of data noise is reduced, ensuring more accurate results; and by normalization and averaging operations, the sensitivity to subtle changes is improved, enhancing the detection capability.

[0024] Furthermore, the interpolation process based on the magnitude of the significance value includes: in response to the significance value being greater than a preset invalid threshold, identifying the location data at the target time as invalid data; otherwise, identifying the location data at the target time as real data; for all invalid data, selecting multiple real data immediately preceding the invalid data, and assigning the average of the latitude and longitude values ​​of the multiple real data to the invalid data to complete the interpolation process.

[0025] Secondly, the present invention provides a positioning data cleaning system based on data analysis, which adopts the following technical solution:

[0026] A location data cleaning system based on data analysis includes a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the aforementioned location data cleaning method based on data analysis is implemented.

[0027] By adopting the above technical solution, a computer program is generated from the above-mentioned data analysis-based location data cleaning method and stored in the memory so that it can be loaded and executed by the processor. In this way, a terminal device can be made based on the memory and the processor for convenient use.

[0028] The present invention has the following technical effects:

[0029] Traditional saliency detection algorithms (CA) are susceptible to interference from potentially invalid data, leading to reduced data segment authenticity and affecting the accuracy of saliency values, resulting in incorrect invalid data identification. This invention refines the assessment of data reliability by calculating special values ​​of latitude and longitude and authenticity indicators. Simultaneously, it retains only data segments with high authenticity as reference data segments to avoid error propagation. This ensures that the preliminary saliency values ​​calculated at each scale of the location data at the target time are more accurate, providing a more reliable basis for the calculation of preliminary saliency values, effectively reducing interference and improving the accuracy of invalid data identification. Interpolation processing based on more accurate saliency values ​​allows for clearer identification of invalid data, making interpolation more reasonable, and the cleaned data after filling in invalid data more closely matches the actual vehicle trajectory. Multi-scale analysis is used when calculating preliminary saliency values, dividing the location data into multiple data segments at different scales. This allows for comprehensive and accurate evaluation of the saliency at the target time from both local and global perspectives, providing rich information for invalid data identification. This invention effectively removes invalid data, improves the quality of positioning data, provides a reliable basis for applications such as vehicle positioning and trajectory analysis, enhances system practicality, and helps related industries improve operational and management efficiency. Attached Figure Description

[0030] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent from the following detailed description taken in conjunction with the accompanying drawings. In the drawings, several embodiments of the invention are illustrated by way of example and not limitation, and the same or corresponding reference numerals denote the same or corresponding parts.

[0031] Figure 1 is a flowchart of a location data cleaning method based on data analysis according to an embodiment of the present invention.

[0032] Figure 2 is a flowchart of step S2 in a location data cleaning method based on data analysis according to an embodiment of the present invention. Detailed Implementation

[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0034] It should be understood that when the terms "first," "second," etc., are used in the claims, specification, and drawings of this invention, they are only used to distinguish different objects and not to describe a specific order. The terms "comprising" and "including" used in the specification and claims of this invention indicate the presence of the described features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof.

[0035] This invention discloses a location data cleaning method based on data analysis, referring to Figure 1, including steps S1-S3:

[0036] S1: Obtain the vehicle's location data at each time point, including latitude and longitude values.

[0037] GPS is used to collect the vehicle's location data at various times, i.e., latitude and longitude coordinate data. Since the longitude and latitude data corresponding to each time need to be analyzed separately, it is necessary to use an analog-to-digital converter to convert the longitude and latitude data to digital to obtain the longitude and latitude values ​​at each time.

[0038] The implementers can set the data collection frequency according to the specific implementation situation, for example, 1 second / time.

[0039] S2: Using significance testing The algorithm detects the location data at each time point and obtains the significance value at each time point.

[0040] It should be noted that the core objective of this invention is to use the saliency detection CA algorithm to detect invalid data in location data, and then perform interpolation processing on all invalid data. During the process, when calculating the saliency value of location data at a scale at each time point, only reference data segments with high accuracy are retained. This is mainly based on the analysis of the numerical characteristics of location data at each time point, thereby achieving the cleaning of location data.

[0041] Utilizing significance During the detection process, the algorithm divides the location data into multiple data segments at each scale according to the preset length of each scale (for example, the first scale has 21 data points, the upper limit of the expanded scale is 5, and 5 data points are selected on both sides of the previous scale for each expansion). (The data segment to which the target time belongs is centered on the target time. If there are not enough data points on either side when the target time is the center point, the data points on the other side are used to make up the difference). The algorithm calculates the preliminary significance value of the target time at each scale and determines the significance value of the target time based on the preliminary significance value. Referring to Figure 2, step S2 includes steps S201-S206, as follows:

[0042] S201: At the target scale in each scale, determine the special values ​​of longitude and latitude at the target time based on the longitude and latitude values ​​of the target time and several previous times, as well as the extreme values ​​of longitude and latitude values ​​at each time.

[0043] The implementers can set the number of previous time points according to the specific implementation situation, for example, 50. If there are less than 50 location data points before the sampling time, the location data collected at the beginning of the first minute will be used as a reference for analysis and will not be used for the actual significance value calculation.

[0044] It should be noted that this step will analyze the latitude and longitude values ​​at each moment to obtain the unique latitude and longitude values ​​for each moment. Quantifying this indicator is for the subsequent analysis of the accuracy of the location data. Since invalid data in location data is usually quite special and extreme in terms of numerical performance, when analyzing this indicator, the smaller the difference between the latitude and longitude values ​​corresponding to each moment and the maximum / minimum values ​​of latitude and longitude in the vehicle's historical location data (at each moment), the more special the latitude and longitude data of the target moment is, and the greater its unique value. However, since there may be extreme longitudes that a vehicle has recently traveled to in a period of time that it has never been to before, it is necessary to further analyze the difference between the latitude and longitude values ​​of the target moment and the latitude and longitude values ​​of its nearest moments. The greater the difference, the more special the latitude and longitude values ​​of the target moment are.

[0045] Specifically, the specific values ​​of longitude and latitude satisfy the following:

[0046] ,

[0047] ;

[0048] In the formula, and They are time points The special values ​​of longitude and latitude, and They are time points The longitude and latitude values, and They are time points The average of the longitude and latitude values ​​at several previous moments. and These are the maximum values ​​of longitude and latitude at each time point. and These are the minimum longitude and latitude values ​​at each time point, respectively. For hyperparameters, For the standard normalized function, It is a minimum value function. It is the absolute value symbol.

[0049] Implementers can set hyperparameters according to specific implementation conditions, for example, 0.000001. The existence of hyperparameters is to prevent... When the value of is 0, the calculation result is meaningless.

[0050] in, Quantified time The difference between the longitude value and the extreme longitude values ​​in the vehicle's historical location data (at various times) is the value of the vehicle's longitude. The smaller this value, the smaller the difference, which indicates the time. The more unique the longitude value, the greater its uniqueness. Indicates the time The relationship between a longitude value and its longitude values ​​at several previous times; the larger the value, the longer the time. The more unique the longitude value is among the longitude values ​​of several previous times, the more indicative it is of the time. The smaller the difference between the longitude value and the extreme value among the longitude values ​​in the vehicle's historical location data, the greater the reliability. The longer the longitude, the greater the value; the same applies to the latitude, which will not be elaborated here.

[0051] S202: Determine the accuracy of the latitude and longitude at the target time based on the specific values ​​of latitude and longitude, as well as the latitude and longitude values ​​of the target time and several previous times.

[0052] It should be noted that this step analyzes the changes in location data at each moment compared to the location data at several previous moments, and optimizes the special values ​​of the location data at each moment to obtain the accuracy of the location data at each moment. Specifically, if the latitude and longitude values ​​of the target moment contribute more to the variance of the latitude and longitude values ​​of the previous moments, it indicates that the data is more likely to be invalid due to interference, and the accuracy is lower. However, during vehicle movement, there are common behaviors such as turning and U-turns. If such behaviors occur at the target moment, they will increase the contribution of the latitude and longitude values ​​of the target moment to the variance of the latitude and longitude values ​​of the previous moments. Therefore, in this step, it is necessary to continue to analyze the difference between the location data of the target moment and the location data of the immediately preceding moment. If the difference between the latitude and longitude values ​​of the target moment and the location data of the immediately preceding moment is greater, it indicates that the location data of the target moment is more likely to be invalid due to interference, and the accuracy is lower.

[0053] Specifically, the accuracy of the latitude and longitude coordinates respectively satisfies:

[0054] ,

[0055] ;

[0056] In the formula, and They are time points The accuracy of longitude and latitude and They are time points The special values ​​of longitude and latitude, and They are time points Compared with the variance of longitude and latitude values ​​at previous times, and They are time points The variance of longitude and the variance of latitude values ​​at previous times. and They are time points The difference in longitude and latitude compared to the previous moment.

[0057] in, The larger the value, the more significant the time. The larger the longitude value, the greater the possibility that it is interfered with and becomes invalid data, and the lower its authenticity will be. Indicates the time The contribution of a given longitude value to the variance formed by its longitude values ​​at previous times. A larger value indicates a greater contribution and provides more information about the time elapsed. The greater the likelihood that a longitude value is interfered with and becomes invalid data, the lower the accuracy of the corresponding longitude value. The larger the value, the more accurate it is to indicate the time period. The greater the contribution of a longitude value to the variance of its longitude values ​​at previous times, the greater its credibility. This indicates that the longitude value at the target time is more likely to be invalid data due to interference, and its longitude accuracy is lower. The same applies to latitude accuracy, which will not be elaborated here.

[0058] S203: The sum of the latitude and longitude accuracy is taken as the location accuracy of the target time; for data segments that do not contain the target time, the sum of the location accuracy of each time within the data segment is taken as the accuracy of the data segment.

[0059] It should be noted that after obtaining the location accuracy at the target time, this step will calculate the accuracy of all data segments when calculating the significance value of the location data at a certain scale based on this indicator.

[0060] S204: Determine the reference data segment based on the magnitude of the realism.

[0061] It should be noted that after obtaining the authenticity of the data segments, data segments with lower authenticity need to be removed, and only data segments with higher authenticity should be retained as reference data segments.

[0062] Specifically, determining the reference data segment based on the magnitude of the realism includes:

[0063] If the authenticity is greater than a preset reference threshold, it indicates that the data segment is less likely to be disturbed, the data segment is more authentic, and the data segment is identified as a reference data segment.

[0064] Implementers can set a reference threshold based on the specific implementation situation, for example, 0.6.

[0065] S205: The average of the sum of the longitude and latitude special values ​​at each time point within each data segment is recorded as the characteristic value of each data segment; based on the characteristic values ​​of the data segment to which the target time belongs and the reference data segment, the preliminary significance value of the target time at the target scale is determined.

[0066] It should be noted that after the data segment selection was completed, a reference data segment with high accuracy was obtained. Subsequently, based on the reference data segment, the preliminary significance value of the location data at the target time was calculated at one scale.

[0067] Specifically, the preliminary significance value satisfies:

[0068] ;

[0069] In the formula, For a moment In scale The initial significance value, For a moment In scale The number of reference data segments below, For a moment In scale The feature values ​​of the data segment to which it belongs. For a moment In scale The next Feature values ​​of a reference data segment It is the absolute value symbol.

[0070] S206: Determine the significance value at the target time based on the preliminary significance value.

[0071] It should be noted that, through the above analysis steps, the preliminary significance value of the target time location data at one scale is obtained. Now, the optimized significance detection CA algorithm is used to detect the target time location data, and the mean of the preliminary significance values ​​of the target time location data at all scales is calculated as its significance value.

[0072] Specifically, the significance value satisfies:

[0073] ;

[0074] In the formula, For a moment The significance value, For the number of scales, For a moment In scale The initial significance value.

[0075] S3: Perform interpolation based on the magnitude of the significance value to complete the location data cleaning based on data analysis.

[0076] Specifically, the interpolation process based on the magnitude of the significance value includes:

[0077] If the significance value is greater than a preset invalid threshold, the location data at the target time is determined to be invalid data; otherwise, the location data at the target time is considered real data.

[0078] For all invalid data, select the multiple real data immediately preceding the invalid data, and assign the average of the latitude and longitude values ​​of the multiple real data to the invalid data to complete the interpolation process.

[0079] Implementers can set an invalid threshold and the number of adjacent real data points based on the specific implementation situation. For example, the invalid threshold can be 0.75 and the number of real data points can be 5.

[0080] This invention also discloses a location data cleaning system based on data analysis, including a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement a location data cleaning method based on data analysis according to the present invention.

[0081] The system also includes other components well known to those skilled in the art, such as communication buses and communication interfaces, the settings and functions of which are known in the art and will not be described in detail here.

[0082] In this invention, the aforementioned memory can be any tangible medium containing or storing a program that can be used or combined with an instruction execution system, apparatus, or device. For example, a computer-readable storage medium can be any suitable magnetic or magneto-optical storage medium, or any other medium that can be used to store desired information and can be accessed by an application, module, or both. Any such computer storage medium can be part of a device or accessible to or connected to a device.

[0083] While this specification has shown and described numerous embodiments of the invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of this invention.

[0084] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A location data cleaning method based on data analysis, characterized in that, include: Acquire the vehicle's location data at various times, including latitude and longitude values; Using significance detection The algorithm detects the location data at each time point and obtains the significance value at each time point; interpolation is performed based on the magnitude of the significance value to complete the location data cleaning based on data analysis. In utilizing significance During the detection process, the algorithm divides the location data into multiple data segments at each preset scale according to the length of each scale, calculates the preliminary significance value of the target time at each scale, and determines the significance value of the target time based on the preliminary significance value, including: At the target scale within each scale, the latitude and longitude values ​​of the target time are determined based on the latitude and longitude values ​​of the target time and several previous times, as well as the extreme values ​​among the latitude and longitude values ​​of each time. The latitude and longitude accuracy of the target time is determined based on these special values ​​and the latitude and longitude values ​​of the target time and several previous times. The sum of these latitude and longitude accuracy values ​​is taken as the location accuracy of the target time. For data segments that do not contain the target time, the sum of the location accuracy values ​​of each time within the data segment is taken as the accuracy of the data segment. Reference data segments are determined based on the magnitude of these accuracy values. The average of the sum of the special longitude and special latitude values ​​of each time within each data segment is recorded as the characteristic value of each data segment. Based on the characteristic values ​​of the data segment to which the target time belongs and the reference data segment, the preliminary significance value of the target time at the target scale is determined.

2. The location data cleaning method based on data analysis according to claim 1, characterized in that, The specific values ​​of longitude and latitude respectively satisfy: , ; In the formula, and They are time points The special values ​​of longitude and latitude, and They are time points The longitude and latitude values, and They are time points The average of the longitude and latitude values ​​at several previous moments. and These are the maximum values ​​of longitude and latitude at each time point. and These are the minimum longitude and latitude values ​​at each time point, respectively. For hyperparameters, For the standard normalized function, It is a minimum value function. It is the absolute value symbol.

3. The location data cleaning method based on data analysis according to claim 1, characterized in that, The accuracy of the latitude and longitude coordinates respectively satisfies: , ; In the formula, and They are time points The accuracy of longitude and latitude. and They are time points The special values ​​of longitude and latitude, and They are time points Compared with the variance of longitude and latitude values ​​at previous times, and They are time points The variance of longitude and the variance of latitude values ​​at previous times. and They are time points The difference in longitude and latitude compared to the previous moment.

4. The location data cleaning method based on data analysis according to claim 1, characterized in that, The step of determining the reference data segment based on the degree of realism includes: In response to the fact that the accuracy is greater than a preset reference threshold, the data segment is identified as a reference data segment.

5. The location data cleaning method based on data analysis according to claim 1, characterized in that, The preliminary significance value satisfies: ; In the formula, For a moment In scale The initial significance value, For a moment In scale The number of reference data segments below, For a moment In scale The feature values ​​of the data segment to which it belongs. For a moment In scale The next Feature values ​​of a reference data segment It is the absolute value symbol.

6. The location data cleaning method based on data analysis according to claim 1, characterized in that, The significance value satisfies: ; In the formula, For a moment The significance value, For the number of scales, For a moment In scale The initial significance value.

7. The location data cleaning method based on data analysis according to claim 1, characterized in that, The interpolation process based on the significance value includes: If the significance value is greater than a preset invalid threshold, the location data at the target time is determined to be invalid data; otherwise, the location data at the target time is considered real data. For all invalid data, select the multiple real data immediately preceding the invalid data, and assign the average of the latitude and longitude values ​​of the multiple real data to the invalid data to complete the interpolation process.

8. A location data cleaning system based on data analysis, characterized in that, include: A processor and a memory, the memory storing computer program instructions that, when executed by the processor, implement a location data cleaning method based on data analysis according to any one of claims 1-7.