A positioning data cleaning method and system based on data analysis

By calculating the special values ​​and true values ​​of longitude and latitude, refine the data evaluation, and adopting multi-scale analysis and interpolation processing, the error problem of the significance detection CA algorithm when judging the invalid position data is solved, and the accuracy of the positioning data and the stability of the system are improved.

CN119884627BActive Publication Date: 2025-06-03GUANGZHOU SEEWORLD TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510378484.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-03
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

When the significance detection CA algorithm determines whether the position data is invalid data, it is disturbed by potentially invalid data, resulting in a decrease in the authenticity of the data segment, affecting the accuracy of the significance value, and thus leading to errors in the judgment of invalid data.

Method used

By calculating indicators such as longitude and latitude special values ​​and realism, we will refine the evaluation of data reliability, and only data segments with higher authenticity are retained as reference data segments, and multi-scale analysis and interpolation processing are used to ensure that the significance value of the position data at the target moment is more accurate at each scale.

Benefits of technology

It effectively reduces the error in judging invalid data, improves the accuracy and reliability of positioning data, enhances the stability and adaptability of the system, and ensures the accuracy of positioning results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119884627B_ABST
    Figure CN119884627B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data processing, and in particular, to a positioning data cleaning method and system based on data analysis. The method includes: obtaining position data of a vehicle at each moment, where the position data includes longitude and latitude values; using the significance detection CA algorithm to detect the position data at each moment to obtain the significance value at each moment, including: determining the longitude and latitude special values and the longitude and latitude authenticity at the target moment; obtaining the authenticity of equally dividing the position data into multiple data segments at each scale according to the lengths of each preset scale; determining the reference data segment and the preliminary significance value, and further obtaining the significance value; performing interpolation processing according to the magnitude of the significance value to complete the positioning data cleaning based on data analysis. By calculating indexes such as longitude and latitude special values and authenticity, the present invention provides a more reliable basis for calculating the preliminary significance value, effectively reduces interference, and improves the accuracy of invalid data judgment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a positioning data cleaning method and system based on data analysis. Background Art

[0002] With the development of big data technology, especially the popularization of mobile Internet and Internet of Things, positioning data has become a core data type in many applications. As a kind of positioning data, location data may contain problems such as noise, error, and missing due to various reasons, which may stem from the instability of device hardware, sensor errors, data loss during data transmission, etc. If these data errors are not cleaned, it will lead to distortion of downstream analysis, decision-making, and prediction, and even affect important applications such as public safety and traffic scheduling. By cleaning location data, the accuracy and integrity of the data can be improved, thus providing more reliable information for relevant decision-making. There is an existing method for detecting invalid location data that uses the significant detection CA algorithm, which has the advantages of strong real-time performance and strong robustness, and can quickly and accurately identify and filter invalid data during the processing of large-scale location data, significantly improving the reliability and accuracy of the positioning system.

[0003] The patent application document with the publication number CN114296063A discloses a CSMA / CA-based cooperative positioning and ranging method, including the following steps: S1: The sending station detects the state of the wireless channel and uses the backoff algorithm to select whether to send the entire data frame; S2: When the sending station repeats using the backoff timer and sends the data frame waiting for confirmation; S3: When the backoff timer of a sending station counts down to 0, use the RTS / CTS data frame to detect whether the receiving station has received the data frame; S4: The sending station repeats step S1 and continues to execute the next frame.

[0004] However, the above patent application document is not aimed at cleaning positioning data, and does not solve the problem that when analyzing whether the location data at a moment is invalid data using the significant detection CA algorithm, it will judge by calculating the size of its significance value. Due to the existence of potential invalid data, when calculating the preliminary significance value of a data point at one scale, if there are invalid data in other data segments at the same scale, it will reduce the data authenticity of the data segment to which the data point belongs, thus affecting the accuracy of the preliminary significance value of the data point at one scale and the significance value at multiple scales, resulting in errors in the judgment of invalid data. Summary of the Invention

[0005] In order to solve the problem that when the significance detection CA algorithm determines whether the position data at each moment is invalid data, it judges by calculating the magnitude of its significance value. However, due to the existence of potential invalid data, the authenticity of the data segment is reduced, thereby affecting the preliminary significance value of the data points at one scale and the accuracy of the significance value at multiple scales, resulting in errors in the judgment of invalid data, the present invention provides a positioning data cleaning method and system based on data analysis.

[0006] In the first aspect, the present invention provides a positioning data cleaning method based on data analysis, adopting the following technical solution:

[0007] A positioning data cleaning method based on data analysis includes: obtaining the position data of the vehicle at each moment, where the position data includes longitude and latitude values; using significance detection algorithm to detect the position data at each moment to obtain the significance value at each moment; performing interpolation processing according to the magnitude of the significance value to complete the positioning data cleaning based on data analysis; during the process of using the significance algorithm for detection, equally dividing the position data into multiple data segments at each scale according to the preset length of each scale, calculating the preliminary significance value of the target moment at each scale among each moment, and determining the significance value of the target moment according to the preliminary significance value, including: at the target scale among each scale, determining the longitude and latitude special values of the target moment according to the longitude and latitude values of the target moment and several previous moments, and the extreme values of the longitude and latitude values at each moment; determining the authenticity of the longitude and latitude of the target moment according to the longitude and latitude special values and the longitude and latitude values of the target moment and several previous moments; taking the sum of the authenticity of the longitude and latitude as the position authenticity of the target moment; for the data segment not including the target moment, taking the sum of the position authenticity of each moment in the data segment as the authenticity of the data segment; determining the reference data segment according to the magnitude of the authenticity; recording the mean value of the sum of the longitude special value and the latitude special value of each moment in each data segment as the characteristic value of each data segment; determining the preliminary significance value of the target moment at the target scale according to the characteristic values of the data segment to which the target moment belongs and the reference data segment.

[0008] The beneficial effects are as follows: By combining the saliency detection CA algorithm and multi-scale analysis, not only can the accuracy of data be improved, but also invalid data can be effectively filtered through the judgment and interpolation processing of saliency values at each scale, reducing the possibility of misjudgment; During the saliency detection process, comparative analysis of the position data at the target moment and several previous moments is adopted, reducing the impact of invalid data on the overall authenticity of the data segment, thereby avoiding positioning errors caused by invalid data; By analyzing the position data at multiple scales, potential valid information in the data can be captured at different scales, and inaccurate or abnormal values can be effectively isolated, enhancing the adaptability of the method to data changes; By calculating special values of the longitude and latitude values of the target moment and its adjacent moments, combined with the change in its authenticity, the characteristic values of the data segment are further determined, ultimately improving the comprehensive evaluation ability of the moment and the data segment; Judging the saliency value according to the characteristic values of the data segment and the target moment can more accurately screen out reliable data, making the final positioning result more accurate, thereby improving the stability and reliability of the positioning system in practical applications.

[0009] Further, the longitude and latitude special values respectively satisfy:

[0010] ,

[0011] ; In the formula, and are respectively the longitude special value and the latitude special value of the moment , and are respectively the longitude value and the latitude value of the moment , and are respectively the means of the longitude values and the latitude values of several previous moments of the moment , and are respectively the maximum values of the longitude values and the latitude values at each moment, and are respectively the minimum values of the longitude values and the latitude values at each moment, is a hyperparameter, is a standard normalization function, is a minimum value function, is an absolute value symbol.

[0012] The beneficial effects are as follows: By precisely calculating the special values of longitude and latitude, the change of position data can be more accurately reflected, thereby improving the accuracy of data cleaning; Considering the mean value of data at historical moments reduces the interference of instantaneous errors and improves the representativeness of data at the target moment; Normalization and standardization operations eliminate the errors caused by differences in data ranges, ensuring fair comparison of different data; By calculating the difference between the maximum and minimum values, extreme data can be effectively identified and processed, reducing its negative impact on the results.

[0013] Further, the longitude and latitude authenticity degrees respectively satisfy:

[0014] ,

[0015] ; In the formula, and are respectively the longitude authenticity degree and latitude authenticity degree at moment , and are respectively the longitude special value and latitude special value at moment , and are respectively the variances of longitude values and latitude values between moment and several previous moments, and are respectively the variances of longitude values and latitude values of several previous moments at moment , and are respectively the differences between the longitude value and latitude value of moment and the previous moment.

[0016] The beneficial effects are as follows: By calculating the longitude and latitude authenticity degrees, the authenticity of positioning data at each moment can be accurately evaluated, ensuring more reliable data; Based on the comparison of longitude and latitude variances, the credibility of data can be dynamically adjusted, identifying data with large fluctuations and optimizing data processing; Introducing the differences between moments to capture the position change trend helps to identify and adjust mutation data, avoiding its impact on the results; Combining special values, variances, and differences for comprehensive evaluation improves the depth and accuracy of data cleaning.

[0017] Further, determining the reference data segment according to the magnitude of the authenticity degree includes: In response to the authenticity degree being greater than a preset reference threshold, determining that the data segment is a reference data segment.

[0018] Further, the preliminary significance value satisfies:

[0019] ; In the formula, is at moment on the scale The preliminary significance value, is the time at the scale the number of reference data segments, is the time at the scale the eigenvalue of the data segment to which it belongs, is the time at the scale the th eigenvalue of the reference data segment, is the absolute value symbol.

[0020] The beneficial effects are as follows: By calculating the preliminary significance value, the significant changes in the data at different scales at the target time can be effectively evaluated, helping to identify key changes or anomalies in the data; Comparing with multiple reference data segments improves the accuracy and robustness of the significance evaluation and avoids the deviation of a single reference data; By averaging the differences, the influence of individual outliers on the significance calculation is effectively reduced, ensuring that the evaluation results are more stable and reliable.

[0021] Further, the significance value satisfies:

[0022] ; where is the significance value at time , is the number of scales, is the time at the scale the preliminary significance value.

[0023] The beneficial effects are as follows: By averaging the preliminary significance values at multiple scales, the significance at the target time can be more comprehensively evaluated, avoiding the deviation of a single scale; By smoothing the significance values of each scale, the influence of data noise is reduced, ensuring more accurate results; Through normalization and averaging operations, the sensitivity to subtle changes is improved and the detection ability is enhanced.

[0024] Further, the interpolation processing according to the magnitude of the significance value includes: in response to the significance value being greater than a preset invalid threshold, determining that the position data at the target time is invalid data, otherwise, the position data at the target time is real data; For all invalid data, select multiple real data immediately adjacent to the invalid data before, and assign the mean of the longitude and latitude values of the multiple real data to the invalid data to complete the interpolation processing.

[0025] In a second aspect, the present invention provides a positioning data cleaning system based on data analysis, adopting the following technical solutions:

[0026] A positioning data cleaning system based on data analysis, comprising: a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned positioning data cleaning method based on data analysis is implemented.

[0027] By adopting the above technical solution, a computer program is generated for the above-mentioned positioning data cleaning method based on data analysis and stored in the memory to be loaded and executed by the processor, so as to manufacture a terminal device according to the memory and the processor, which is convenient to use.

[0028] The present invention has the following technical effects:

[0029] The traditional significance detection CA algorithm is vulnerable to interference from potential invalid data, resulting in a reduction in the authenticity of data segments, affecting the accuracy of significance values, causing misjudgment of invalid data. By calculating indicators such as special values of longitude and latitude and authenticity, the present invention refines the evaluation of data reliability. At the same time, only data segments with higher authenticity are retained as reference data segments to avoid error propagation, ensuring that the preliminary significance values calculated at each scale of the position data at the target moment are more accurate, providing a more reliable basis for the calculation of preliminary significance values, effectively reducing interference, and improving the accuracy of invalid data judgment; interpolation processing is carried out based on more accurate significance values. The accurate significance values can more clearly identify invalid data, making the interpolation more reasonable, and the cleaned data after filling in the invalid data is more consistent with the actual driving trajectory of the vehicle; multi-scale analysis is used when calculating the preliminary significance value, and the position data is equally divided into multiple data segments at different scales, which can comprehensively consider the data from both local and global perspectives, comprehensively and accurately evaluate the significance at the target moment, and provide rich information for invalid data judgment; the present invention effectively removes invalid data, improves the quality of positioning data, provides a reliable basis for applications such as vehicle positioning and trajectory analysis, enhances the practicability of the system, and helps related industries improve operation and management efficiency. Description of the Drawings

[0030] By referring to the drawings and reading the following detailed description, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become easily understandable. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals are for the same or corresponding parts.

[0031] Figure 1 It is a method flow chart in a positioning data cleaning method based on data analysis according to an embodiment of the present invention.

[0032] Figure 2 It is a method flow chart of step S2 in a positioning data cleaning method based on data analysis according to an embodiment of the present invention. Detailed Embodiments

[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0034] It should be understood that when terms such as "first" and "second" are used in the claims, the description, and the drawings of the present invention, they are only used to distinguish different objects, rather than to describe a specific order. The terms "including" and "comprising" used in the description and claims of the present invention indicate the existence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the existence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0035] An embodiment of the present invention discloses a positioning data cleaning method based on data analysis. Refer to Figure 1 , including steps S1 - S3:

[0036] S1: Obtain the position data of the vehicle at each moment, and the position data includes longitude and latitude values.

[0037] Use GPS to collect the position data of the vehicle at each moment, that is, longitude and latitude coordinate data. Since the longitude data and latitude data corresponding to each moment need to be analyzed separately later, a modulus conversion device is required to perform modulus conversion on the longitude data and latitude data to obtain the longitude value and latitude value at each moment.

[0038] The implementer can set the collection frequency according to the specific implementation situation. For example, 1 s / time.

[0039] S2: Use the significance detection algorithm to detect the position data at each moment, and obtain the significance value at each moment.

[0040] It should be noted that the core purpose of the present invention is to use the significance detection CA algorithm to detect the invalid data in the position data, and then perform interpolation processing on all the invalid data. During the process, when calculating the significance value of the position data at each moment at a certain scale, only the reference data segment with a higher authenticity is retained, which is mainly based on the analysis of the numerical characteristics of the position data at each moment, so as to realize the cleaning of the position data.

[0041] When using significance During the detection process of the algorithm, according to the preset lengths of each scale (exemplarily, the first scale is 21 data points, the upper limit of the extended scale is 5, and 5 data points are selected on both sides of the previous scale for extension each time), the position data is equally divided into multiple data segments at each scale (wherein, the data segment to which the target moment belongs takes the target moment as the center point of the data segment. If the number of data points on either side is insufficient when taking the target moment as the center point, it can be supplemented with the data points on the other side), calculate the preliminary significance value of the target moment at each scale for each moment, and determine the significance value of the target moment according to the preliminary significance value; refer to Figure 2 Step S2 includes steps S201 - S206, specifically as follows:

[0042] S201: At the target scale among each scale, according to the longitude and latitude values of the target moment and several previous moments, as well as the extreme values of the longitude and latitude values at each moment, determine the longitude and latitude special values of the target moment.

[0043] Implementers can set the number of previous moments according to the specific implementation situation. For example, 50. If the position data before the sampling moment is less than 50, the position data of the first 1 minute collected will be used as a reference for analysis and not for actual significance value calculation.

[0044] It should be noted that in this step, the performance of the longitude and latitude values for each moment is analyzed to obtain the longitude and latitude special values for each moment; the quantification of this index is for the subsequent analysis of the authenticity of the position data. Since the invalid data in the position data is usually relatively special and extreme at the numerical performance level, when analyzing this index, the smaller the difference between the longitude and latitude values corresponding to each moment and the maximum / minimum values of the longitude and latitude values in the vehicle's historical position data (at each moment), the more special the longitude and latitude data of the target moment, and the greater its special value. However, since it is possible that a vehicle has traveled to an extreme longitude that it has never reached before in a recent period of time, it is necessary to further analyze the difference between the longitude and latitude values of the target moment and the longitude and latitude values of its adjacent moments. The greater the difference, the more special the longitude and latitude values of the target moment can be explained.

[0045] Specifically, the longitude and latitude special values respectively satisfy:

[0046] ,

[0047] ;

[0048] In the formula, and are respectively the longitude special value and the latitude special value of the moment , and are the longitude value and the latitude value at time respectively, and are the means of the longitude values and the latitude values at several moments before time respectively, and are the maximum values among the longitude values and the latitude values at each moment respectively, and are the minimum values among the longitude values and the latitude values at each moment respectively, is a hyperparameter, is the standard normalization function, is the minimum value function, is the absolute value symbol.

[0049] Implementers can set the hyperparameter according to the specific implementation situation. For example, 0.000001. The existence of the hyperparameter is to prevent from being 0, which may lead to meaningless calculation results.

[0050] Among them, quantifies the difference between the longitude value at time and the extreme value of the longitude values in the historical position data (at each moment) of the vehicle. The smaller this value is, the smaller the difference is, indicating that the longitude value at time is more special, and its special value is larger; represents the difference relationship of the longitude value at time among the longitude values at several moments before it. The larger this value is, the more special the existence of the longitude value at time is among the longitude values at several moments before it. It can also indicate that the difference between the longitude value at time and the extreme value of the longitude values in the historical position data of the vehicle is smaller with a greater credibility. Then the longitude special value at time will be larger; The latitude special value is the same and will not be elaborated here.

[0051] S202: Determine the longitude and latitude authenticity at the target moment according to the longitude and latitude special values, and the longitude and latitude values at the target moment and several moments before it.

[0052] It should be noted that in this step, the change characteristics of the position data at each moment and the position data at several previous moments will be analyzed, and the special values of the position data at each moment will be optimized to obtain the authenticity of the position data at each moment. Specifically, in the analysis, if the contribution of the longitude and latitude values at the target moment to the variance of the longitude and latitude values at several previous moments is more, it can be explained that the possibility of it being interfered and being invalid data is greater, and the authenticity is lower. However, during the vehicle driving process, there will be conventional behaviors such as turning and U-turning. If such a behavior exists at the target moment, it will all cause the contribution of the longitude and latitude values at the target moment to the variance value of the longitude and latitude values at several previous moments to increase. Therefore, in this step, it is also necessary to continue to analyze the difference between the position data at the target moment and the position data at the immediately previous moment. If the difference between the longitude and latitude values at the target moment and the immediately previous moment is greater, it can be explained that the possibility of the longitude and latitude values at the target moment being interfered and being invalid data is greater, and its authenticity is lower.

[0053] Specifically, the longitude and latitude authenticity respectively satisfy:

[0054] ,

[0055] ;

[0056] In the formula, and are respectively the longitude authenticity and latitude authenticity at moment , and are respectively the longitude special value and latitude special value at moment , and are respectively the variance of the longitude values and the variance of the latitude values between moment and the longitude values at several previous moments, and are respectively the variance of the longitude values and the variance of the latitude values at several previous moments of the longitude values at moment , and are respectively the difference between the longitude value at moment and the longitude value at the previous moment and the difference between the latitude value and the latitude value at the previous moment.

[0057] Among them, The larger it is, the larger the longitude special value at moment is, then the greater the possibility that it is interfered and is invalid data, and the lower its authenticity; represents the contribution amount of the longitude value at moment to the variance composed of the longitude values at several previous moments. The larger this value is, the more the contribution, indicating that at moment The greater the possibility that the longitude value is interfered with and is invalid data, the lower the authenticity of the corresponding longitude; The bigger, the more it can describe the moment The more the longitude value contributes to the variance between it and the longitude values ​​at several previous moments, the greater the credibility is, which means that the longitude value at the target moment is more likely to be interfered with and is invalid data, and the lower its longitude authenticity is; the same is true for latitude authenticity, which will not be elaborated here.

[0058] S203: taking the sum of the longitude and latitude authenticity as the position authenticity at the target moment; for a data segment that does not include the target moment, taking the sum of the position authenticity at each moment in the data segment as the authenticity of the data segment.

[0059] It should be noted that after obtaining the location authenticity at the target moment, in this step, the authenticity of all data segments when calculating the significance value of the location data at a scale is calculated based on the indicator.

[0060] S204: Determine a reference data segment according to the degree of authenticity.

[0061] It should be noted that after obtaining the authenticity of the data segments, it is necessary to eliminate the data segments with lower authenticity and only retain the data segments with higher authenticity as reference data segments.

[0062] Specifically, determining the reference data segment according to the degree of authenticity includes:

[0063] In response to the authenticity being greater than a preset reference threshold, it indicates that the possibility that the data in the data segment is interfered with is smaller, the authenticity of the data segment is higher, and the data segment is identified as a reference data segment.

[0064] Implementers can set the reference threshold according to specific implementation circumstances, for example, 0.6.

[0065] S205: Record the average of the sum of the longitude special value and the latitude special value at each time in each data segment as the characteristic value of each data segment; determine the preliminary significance value of the target moment at the target scale according to the characteristic values ​​of the data segment to which the target moment belongs and the reference data segment.

[0066] It should be noted that after the data segments are screened, reference data segments with a high degree of authenticity are obtained. Subsequently, based on the reference data segments, a preliminary significance value of the position data at a scale at the target moment is calculated.

[0067] Specifically, the preliminary significance value satisfies:

[0068] ;

[0069] In the formula, For the moment At scale The preliminary significance value of is the moment At scale The number of reference data segments under is the moment At scale The eigenvalue of the data segment to which it belongs under is the moment At scale Under the The eigenvalue of the th reference data segment is the absolute value symbol.

[0070] S206: Determine the significance value of the target moment according to the preliminary significance value.

[0071] It should be noted that through the above analysis steps, the preliminary significance value of the position data of the target moment at one scale is obtained. Now, the optimized significance detection CA algorithm is used to detect the position data of the target moment, and the mean value of the preliminary significance values of the position data of the target moment at all scales is calculated as its significance value.

[0072] Specifically, the significance value satisfies:

[0073] ;

[0074] In the formula, is the moment The significance value of is the number of scales, is the moment At scale The preliminary significance value of

[0075] S3: Perform interpolation processing according to the magnitude of the significance value to complete the cleaning of positioning data based on data analysis.

[0076] Specifically, the interpolation processing according to the magnitude of the significance value includes:

[0077] In response to the significance value being greater than a preset invalid threshold, it is determined that the position data of the target moment is invalid data; otherwise, the position data of the target moment is real data;

[0078] For all invalid data, select multiple real data immediately adjacent to the invalid data before, and assign the mean value of the longitude and latitude values of the multiple real data to the invalid data to complete the interpolation processing.

[0079] Implementers can set the invalid threshold and the number of immediately preceding real data according to the specific implementation situation. For example, the invalid threshold is 0.75 and the number of real data is 5.

[0080] An embodiment of the present invention also discloses a positioning data cleaning system based on data analysis, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a positioning data cleaning method based on data analysis according to the present invention is implemented.

[0081] The above system also includes other components well-known to those skilled in the art such as a communication bus and a communication interface. Their settings and functions are known in the art, so they will not be described in detail here.

[0082] In the present invention, the aforementioned memory can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. For example, a computer-readable storage medium can be any suitable magnetic storage medium or magneto-optical storage medium, or any other medium that can be used to store the required information and can be accessed by an application program, module, or both. Any such computer storage medium can be part of the device or accessible or connectable to the device.

[0083] Although this specification has shown and described multiple embodiments of the present invention, it is obvious to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art will think of many changes, alterations, and alternative ways without departing from the spirit and idea of the present invention. It should be understood that various alternative solutions to the embodiments of the present invention described herein can be adopted in the process of practicing the present invention.

[0084] The above are all preferred embodiments of the present invention. The protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A positioning data cleaning method based on data analysis, characterized in that: include: Obtaining the location data of the vehicle at each time, wherein the location data includes longitude and latitude values; Using saliency detection The algorithm detects the location data at each moment to obtain the significance value at each moment; interpolation processing is performed according to the magnitude of the significance value to complete the positioning data cleaning based on data analysis; In using significant During the detection process of the algorithm, the position data is equally divided into multiple data segments at each scale according to the length of each preset scale, and the preliminary significance value of the target moment at each time is calculated. According to the preliminary significance value, the significance value of the target moment is determined, including: At the target scale in each scale, the longitude and latitude anomaly factors of the target moment are determined according to the longitude and latitude values ​​of the target moment and several previous moments, as well as the extreme values ​​of the longitude and latitude values ​​at each moment; the longitude and latitude authenticity of the target moment is determined according to the longitude and latitude anomaly factors and the longitude and latitude values ​​of the target moment and several previous moments; the sum of the longitude and latitude authenticity is taken as the position authenticity of the target moment; for a data segment that does not include the target moment, the sum of the position authenticity of each moment in the data segment is taken as the authenticity of the data segment; according to the size of the authenticity, a reference data segment is determined; the mean value of the sum of the longitude anomaly factor and the latitude anomaly factor at each moment in each data segment is recorded as the characteristic value of each data segment; according to the characteristic values ​​of the data segment to which the target moment belongs and the reference data segment, the preliminary significance value of the target moment at the target scale is determined.

2. A positioning data cleaning method based on data analysis according to claim 1, characterized in that: The longitude and latitude anomaly factors satisfy: , ; In the formula, and Separately for the moment The longitude anomaly factor and latitude anomaly factor of and Separately for the moment The longitude and latitude values ​​of and Separately for the moment The average of the longitude and latitude values ​​at several previous moments, and are the maximum values ​​of longitude and latitude at each moment, and are the minimum values ​​of longitude and latitude at each moment, is a hyperparameter, is the standard normalization function, is the minimum function, is the absolute value symbol.

3. The method for cleaning positioning data based on data analysis according to claim 1, characterized in that: The longitude and latitude truths satisfy: , ; In the formula, and Separately for the moment The true longitude and latitude of and Separately for the moment The longitude anomaly factor and latitude anomaly factor of and Separately for the moment The variance of the longitude and latitude values ​​at several previous moments, and Separately for the moment The variance of the longitude and latitude values ​​at several previous moments, and Separately for the moment The difference in longitude and latitude from the previous moment, is the standard normalization function.

4. The method for cleaning positioning data based on data analysis according to claim 1, characterized in that: The step of determining the reference data segment according to the degree of authenticity includes: In response to the authenticity being greater than a preset reference threshold, the data segment is identified as a reference data segment.

5. The method for cleaning positioning data based on data analysis according to claim 1, characterized in that: The preliminary significance value satisfies: ; In the formula, For the moment In scale The initial significance value of For the moment In scale The number of reference data segments under For the moment In scale The characteristic value of the data segment below, For the moment In scale The next The characteristic value of the reference data segment, is the absolute value symbol.

6. The method for cleaning positioning data based on data analysis according to claim 1, characterized in that: The significance value satisfies: ; In the formula, For the moment The significance value of is the number of scales, For the moment In scale The initial significance value of is the standard normalization function.

7. The method for cleaning positioning data based on data analysis according to claim 1, characterized in that: The interpolation processing is performed according to the magnitude of the significance value, comprising: In response to the significance value being greater than a preset invalid threshold, determining that the position data at the target time is invalid data, otherwise, the position data at the target time is true data; For all invalid data, multiple real data immediately before the invalid data are selected, and the average of the longitude and latitude values ​​of the multiple real data is assigned to the invalid data to complete the interpolation process.

8. A positioning data cleaning system based on data analysis, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a positioning data cleaning method based on data analysis according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Cooperative positioning ranging method based on CSMA / CA

    CN114296063A

  • Wind turbine generator abnormal operation data processing method

    CN111522808A

  • Physical examination data significance analysis method based on multiple examination correction

    CN119296785A