Space-time big data credibility evaluation method and system

By standardizing and calculating the similarity of spatiotemporal big data, the problem of low efficiency in the evaluation of spatiotemporal big data in existing technologies has been solved, and efficient and objective credibility assessment has been achieved.

CN121786496APending Publication Date: 2026-04-03NAT SURVEYING & MAPPING PROD QUALITY INSPECTION & TESTING CENT
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies are inefficient in spatiotemporal big data assessment, making it difficult to effectively verify the multi-source distribution, heterogeneity, spatiotemporal correlation, and high noise characteristics of data, resulting in unreliable data affecting decision-making and applications.

Method used

By acquiring surveying and mapping geographic information data, IoT sensing data, and Internet crawling data, we standardize, classify, and filter data sequences of the same type, calculate similarity, construct a data screening matrix, and use similarity to assess data credibility and remove unreliable data.

Benefits of technology

It enables objective, quantitative, and repeatable credibility assessment based on the consistency and synergy of the data itself, improving the efficiency and accuracy of the assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786496A_ABST
    Figure CN121786496A_ABST
Patent Text Reader

Abstract

The invention relates to a time-space big data credibility evaluation method and system, and belongs to the technical field of data credibility evaluation. Performing standardization processing on the space-time big data to obtain standardized space-time big data; classifying the standardized space-time big data to obtain different types of data acquisition sequences; screening out the same type of data acquisition sequences from different sources; calculating the similarity of the same type of data acquisition sequences from different sources; and evaluating the credibility of the data acquisition sequence by using the similarity. According to the method, cross validation is performed on the basis of the similarity among multi-source same-type data acquisition sequences such as surveying and mapping geographic information data, Internet of Things perception data and Internet captured data, judgment is not dependent on a single data source or artificial experience, and the credibility can be evaluated from the consistency and cooperative relationship of the data; and the evaluation result is more objective, quantitative and repeatable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data credibility assessment technology, and in particular to a spatiotemporal big data credibility assessment method and system. Background Technology

[0002] Spatiotemporal big data reflects changes in socio-economic activities and people's production and life. Therefore, it has been widely used in surveying and mapping geographic information data updates, natural resource management, smart city construction, natural disaster monitoring and early warning, environmental monitoring, smart agriculture, and other fields.

[0003] Spatiotemporal big data is characterized by multi-source distribution, heterogeneity, spatiotemporal correlation, and high noise. These characteristics may lead to inconsistencies, inaccuracies, incompleteness, and outdatedness, or data describing the same entity may conflict. These problems inevitably result in a large amount of unreliable data in big data. Without a good and reliable data environment and timely and effective evaluation of the collected data, allowing problematic data to be used in work will bring great risks to the application of big data, mislead decision support and intelligent applications, and unreliable data may even cause serious economic and social losses.

[0004] Verification analysis is a commonly used method for assessing the credibility of spatiotemporal big data. Verification analysis refers to analyzing the quality assurance of spatiotemporal big data in data collection and aggregation, data governance, and data fusion, as well as the accuracy, timeliness, and precision of the spatiotemporal big data itself, by verifying evidence such as original data, process quality control, technical design and summary reports, as well as data organization and naming, data classification and expression, spatiotemporal correlation, attribute and field filling.

[0005] The verification and analysis method currently mainly adopts a combination of quality inspection software and human-computer interaction. In the context of spatiotemporal big data, it has the following main drawbacks: the quality inspection software is mostly focused on checking "surface" quality issues such as format specifications, field completeness, and consistency. For "deeper" credibility characteristics such as spatiotemporal neighborhood relationships, trajectory continuity, and mutual corroboration and conflict of multi-source observations, manual analysis and judgment are still usually required, which is inefficient. Summary of the Invention

[0006] To address the aforementioned problems, the purpose of this invention is to provide a method and system for assessing the credibility of spatiotemporal big data.

[0007] A method for assessing the credibility of spatiotemporal big data includes:

[0008] Step 1: Acquire spatiotemporal big data; the spatiotemporal big data includes: surveying and mapping geographic information data, Internet of Things sensing data, and Internet crawled data;

[0009] Step 2: Standardize the spatiotemporal big data to obtain standardized spatiotemporal big data;

[0010] Step 3: Classify the standardized spatiotemporal big data to obtain different types of data collection sequences;

[0011] Step 4: Filter out data collection sequences of the same type from different sources;

[0012] Step 5: Calculate the similarity of similar data collection sequences from different sources;

[0013] Step 6: Use the similarity to assess the credibility of the data collection sequence.

[0014] Preferably, step 2: standardizing the spatiotemporal big data to obtain standardized spatiotemporal big data includes:

[0015] Step 2.1: Calculate the mean and variance of the spatiotemporal big data;

[0016] Step 2.2: Standardize the spatiotemporal big data based on its mean and variance to obtain standardized spatiotemporal big data; the standardization process is as follows:

[0017]

[0018] Among them, z (i) Let x represent the i-th standardized spatiotemporal big data, μ represent the mean of all spatiotemporal big data, σ represent the standard deviation of all spatiotemporal big data, and x represent the standard deviation of all spatiotemporal big data. i Let N represent the i-th spatiotemporal big data, and let N represent the total number of spatiotemporal big data.

[0019] Preferably, step 5: calculating the similarity of similar data collection sequences from different sources includes:

[0020] Formula used:

[0021]

[0022] Calculate the similarity between similar data collection sequences from different sources; where r(A,B) represents the similarity between data collection sequence A and data collection sequence B, and a i This represents the i-th element in the data collection sequence A. b represents the mean of the data collection sequence A. i This represents the i-th element in the data acquisition sequence B. This represents the mean of data collection sequence B.

[0023] Preferably, step 6: evaluating the credibility of the data collection sequence using the similarity includes:

[0024] Step 6.1: Construct a data screening matrix based on the similarity of similar data collection sequences from different sources;

[0025] Step 6.2: Calculate the threshold using the data screening matrix;

[0026] Step 6.3: Mark data collection sequences with similarity less than or equal to the threshold as "trustworthy", and mark data collection sequences with similarity greater than the threshold as "untrustworthy".

[0027] Preferably, in step 6.1, the data screening matrix is:

[0028]

[0029] in, D represents the data screening matrix. ij Let represent the element in the i-th row and j-th column of the data screening matrix, where 1 ≤ i ≤ L and 1 ≤ j ≤ L.

[0030] Preferably, step 6.2: calculating the threshold using the data screening matrix includes:

[0031] The threshold is calculated based on the extreme values ​​in the data screening matrix; the formula for calculating the threshold is:

[0032]

[0033] Among them, t fine D represents the threshold. max This represents the maximum value in the data screening matrix.

[0034] Preferably, after step 6.3, the method further includes:

[0035] Calculate the timeliness of the filtered data collection sequences and remove data collection sequences whose timeliness is outside the set range; the formula for calculating data timeliness is as follows:

[0036] Timelines = e -λ·Δt

[0037] Where Timelines represents the timeliness of the data, λ represents the attenuation coefficient, and Δt represents the time interval from data acquisition to the present.

[0038] This invention also provides a spatiotemporal big data credibility assessment system, comprising:

[0039] The data acquisition module is used to acquire spatiotemporal big data; the spatiotemporal big data includes: surveying and mapping geographic information data, Internet of Things sensing data, and Internet crawled data;

[0040] The standard processing module is used to standardize spatiotemporal big data to obtain standardized spatiotemporal big data.

[0041] The data classification module is used to classify standardized spatiotemporal big data to obtain different types of data collection sequences;

[0042] The filtering module is used to filter out data collection sequences of the same type from different sources;

[0043] The similarity calculation module is used to calculate the similarity between similar data collection sequences from different sources.

[0044] A credibility assessment module is used to assess the credibility of the data collection sequence using the similarity.

[0045] The present invention also provides an electronic device, including a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the transceiver, the memory, and the processor are connected via the bus, characterized in that the computer program, when executed by the processor, implements the steps in the above-described spatiotemporal big data credibility assessment method.

[0046] The present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the above-described spatiotemporal big data credibility assessment method.

[0047] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0048] This invention relates to a spatiotemporal big data credibility assessment method. Compared with the prior art, this invention is based on cross-validation of the similarity between multiple data collection sequences of the same type, such as surveying and mapping geographic information data, Internet of Things sensing data, and Internet crawling data. It does not rely on a single data source or human experience judgment, and can assess credibility from the consistency and synergy of the data itself, making the assessment results more objective, quantitative and repeatable.

[0049] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 A flowchart of a spatiotemporal big data credibility assessment method provided by the present invention;

[0052] Figure 2 A schematic diagram of a spatiotemporal big data credibility assessment system provided by the present invention. Detailed Implementation

[0053] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0054] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0055] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0056] Please see Figure 1 A spatiotemporal big data credibility assessment method includes:

[0057] Step 1: Acquire spatiotemporal big data; the spatiotemporal big data includes: surveying and mapping geographic information data, Internet of Things sensing data, and Internet crawled data;

[0058] Step 2: Standardize the spatiotemporal big data to obtain standardized spatiotemporal big data;

[0059] Step 2 includes:

[0060] Step 2.1: Calculate the mean and variance of the spatiotemporal big data;

[0061] Step 2.2: Standardize the spatiotemporal big data based on its mean and variance to obtain standardized spatiotemporal big data; the standardization process is as follows:

[0062]

[0063] Among them, z (i) Let x represent the i-th standardized spatiotemporal big data, μ represent the mean of all spatiotemporal big data, σ represent the standard deviation of all spatiotemporal big data, and x represent the standard deviation of all spatiotemporal big data. i Let N represent the i-th spatiotemporal big data, and let N represent the total number of spatiotemporal big data.

[0064] Step 3: Classify the standardized spatiotemporal big data to obtain different types of data collection sequences;

[0065] Step 4: Filter out data collection sequences of the same type from different sources;

[0066] Step 5: Calculate the similarity of similar data collection sequences from different sources;

[0067] In step 5, the following formula is used:

[0068]

[0069] Calculate the similarity between similar data collection sequences from different sources; where r(A,B) represents the similarity between data collection sequence A and data collection sequence B, and a i This represents the i-th element in the data collection sequence A. b represents the mean of the data collection sequence A. i This represents the i-th element in the data acquisition sequence B. This represents the mean of data collection sequence B.

[0070] Step 6: Use the similarity to assess the reliability of the data collection sequence;

[0071] Furthermore, step 6 includes:

[0072] Step 6.1: Construct a data screening matrix based on the similarity of similar data collection sequences from different sources; the data screening matrix is ​​as follows:

[0073]

[0074] in, D represents the data screening matrix. ij Let represent the element in the i-th row and j-th column of the data screening matrix, where 1 ≤ i ≤ L and 1 ≤ j ≤ L.

[0075] Step 6.2: Calculate the threshold using the data screening matrix;

[0076] In step 6.2, the threshold is calculated based on the extreme values ​​in the data screening matrix; the formula for calculating the threshold is:

[0077]

[0078] Among them, t fine D represents the threshold. max This represents the maximum value in the data screening matrix.

[0079] Step 6.3: Mark data collection sequences with similarity less than or equal to the threshold as "trustworthy", and mark data collection sequences with similarity greater than the threshold as "untrustworthy".

[0080] Following step 6.3, the following is also included:

[0081] Calculate the timeliness of the filtered data collection sequences, and remove data collection sequences whose timeliness is outside the set range (Timelines ≥ 60%). The formula for calculating data timeliness is as follows:

[0082] Timelines = e -λ·Δt

[0083] Here, Timeline represents the timeliness of the data, λ represents the decay coefficient, which controls the rate at which the reliability of the data decreases over time and needs to be calibrated according to the specific scenario. Δt represents the time interval from data collection to the present, and it must be consistent with the time unit of λ (such as month or year).

[0084] Different fields have varying sensitivities to data timeliness. Through research experiments, with λ set to 0.1 for spatiotemporal information data and the update cycle used as the evaluation parameter, the data timeliness T was calculated using a data age of 1-10 years. l The calculation results are shown in Table 1 below.

[0085] Table 1

[0086]

[0087]

[0088] The tables above are combined, and the values ​​are rounded to obtain Table 2 below.

[0089] Table 2 Comparison of Data Timeliness and Reliability

[0090]

[0091] This invention cross-validates data based on the similarity between multiple data collection sequences of the same type, such as surveying and mapping geographic information data, IoT sensing data, and Internet crawling data. It does not rely on a single data source or human experience judgment, and can assess credibility based on the consistency and synergy of the data itself, making the assessment results more objective, quantitative, and repeatable.

[0092] This invention also provides a spatiotemporal big data credibility assessment system, comprising:

[0093] The data acquisition module is used to acquire spatiotemporal big data; the spatiotemporal big data includes: surveying and mapping geographic information data, Internet of Things sensing data, and Internet crawled data;

[0094] The standard processing module is used to standardize spatiotemporal big data to obtain standardized spatiotemporal big data.

[0095] The data classification module is used to classify standardized spatiotemporal big data to obtain different types of data collection sequences;

[0096] The filtering module is used to filter out data collection sequences of the same type from different sources;

[0097] The similarity calculation module is used to calculate the similarity between similar data collection sequences from different sources.

[0098] A credibility assessment module is used to assess the credibility of the data collection sequence using the similarity.

[0099] Compared with the prior art, the beneficial effects of the spatiotemporal big data credibility assessment system provided by the present invention are the same as the beneficial effects of the spatiotemporal big data credibility assessment method described in the above technical solution, and will not be repeated here.

[0100] The present invention also provides an electronic device, including a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor. The transceiver, the memory, and the processor are connected via the bus. The computer program, when executed by the processor, implements the steps in the aforementioned spatiotemporal big data credibility assessment method. Compared with the prior art, the beneficial effects of the electronic device provided by the present invention are the same as those of the aforementioned spatiotemporal big data credibility assessment method, and will not be elaborated upon here.

[0101] The present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the steps in the above-described spatiotemporal big data credibility assessment method. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present invention are the same as the beneficial effects of the spatiotemporal big data credibility assessment method described in the above-described technical solution, and will not be repeated here.

[0102] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for assessing the credibility of spatiotemporal big data, characterized in that, include: Step 1: Acquire spatiotemporal big data; The spatiotemporal big data includes: surveying and mapping geographic information data, Internet of Things sensing data, and Internet crawling data; Step 2: Standardize the spatiotemporal big data to obtain standardized spatiotemporal big data; Step 3: Classify the standardized spatiotemporal big data to obtain different types of data collection sequences; Step 4: Filter out data collection sequences of the same type from different sources; Step 5: Calculate the similarity of similar data collection sequences from different sources; Step 6: Use the similarity to assess the credibility of the data collection sequence.

2. The spatiotemporal big data credibility assessment method according to claim 1, characterized in that, Step 2: Standardize the spatiotemporal big data to obtain standardized spatiotemporal big data, including: Step 2.1: Calculate the mean and variance of the spatiotemporal big data; Step 2.2: Standardize the spatiotemporal big data based on its mean and variance to obtain standardized spatiotemporal big data; the standardization process is as follows: Among them, z (i) Let x represent the i-th standardized spatiotemporal big data, μ represent the mean of all spatiotemporal big data, σ represent the standard deviation of all spatiotemporal big data, and x represent the standard deviation of all spatiotemporal big data. i Let N represent the i-th spatiotemporal big data, and let N represent the total number of spatiotemporal big data.

3. The spatiotemporal big data credibility assessment method according to claim 2, characterized in that, Step 5: Calculate the similarity of similar data collection sequences from different sources, including: Formula used: Calculate the similarity between similar data collection sequences from different sources; where r(A,B) represents the similarity between data collection sequence A and data collection sequence B, and a i This represents the i-th element in the data collection sequence A. b represents the mean of the data collection sequence A. i This represents the i-th element in the data acquisition sequence B. This represents the mean of data collection sequence B.

4. The spatiotemporal big data credibility assessment method according to claim 3, characterized in that, Step 6: Evaluating the credibility of the data collection sequence using the similarity score includes: Step 6.1: Construct a data screening matrix based on the similarity of similar data collection sequences from different sources; Step 6.2: Calculate the threshold using the data screening matrix; Step 6.3: Mark data collection sequences with similarity less than or equal to the threshold as "trustworthy", and mark data collection sequences with similarity greater than the threshold as "untrustworthy".

5. The spatiotemporal big data credibility assessment method according to claim 4, characterized in that, In step 6.1, the data screening matrix is ​​as follows: in, D represents the data screening matrix. ij Let represent the element in the i-th row and j-th column of the data screening matrix, where 1 ≤ i ≤ L and 1 ≤ j ≤ L.

6. The spatiotemporal big data credibility assessment method according to claim 5, characterized in that, Step 6.2: Calculating the threshold using the data screening matrix includes: The threshold is calculated based on the extreme values ​​in the data screening matrix; the formula for calculating the threshold is: Among them, t fine D represents the threshold. max This represents the maximum value in the data screening matrix.

7. The spatiotemporal big data credibility assessment method according to claim 6, characterized in that, Following step 6.3, the following is also included: Calculate the timeliness of the filtered data collection sequences and remove data collection sequences whose timeliness is outside the set range; the formula for calculating data timeliness is as follows: Timelines=e -λ·Δt Where Timelines represents the timeliness of the data, λ represents the attenuation coefficient, and Δt represents the time interval from data acquisition to the present.

8. A spatiotemporal big data credibility assessment system, characterized in that, include: The data acquisition module is used to acquire spatiotemporal big data; The spatiotemporal big data includes: surveying and mapping geographic information data, Internet of Things sensing data, and Internet crawling data; The standard processing module is used to standardize spatiotemporal big data to obtain standardized spatiotemporal big data. The data classification module is used to classify standardized spatiotemporal big data to obtain different types of data collection sequences; The filtering module is used to filter out data collection sequences of the same type from different sources; The similarity calculation module is used to calculate the similarity between similar data collection sequences from different sources. A credibility assessment module is used to assess the credibility of the data collection sequence using the similarity.

9. An electronic device comprising a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the transceiver, the memory, and the processor are connected via the bus, characterized in that, When the computer program is executed by the processor, it implements the steps in the spatiotemporal big data credibility assessment method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps in the spatiotemporal big data credibility assessment method as described in any one of claims 1-7.