A method and system for assessing the quality of spatiotemporal information contained in crowd-sensing data.

By constructing a cross-modal unified quality factor system and a standardized evaluation mechanism, the problems of fragmentation and insufficient adaptability in the quality evaluation of collective intelligence sensing data have been solved, enabling accurate and efficient evaluation of multi-source data and supporting the dynamic updating of spatiotemporal knowledge graphs.

CN122086879APending Publication Date: 2026-05-26FUZHOU UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUZHOU UNIV
Filing Date
2026-02-14
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies for assessing the quality of collectively sensed data suffer from fragmented quality factor dimensions, a lack of unified cross-modal measurement standards, a lack of a general extraction framework, and insufficient automation and adaptability, making it difficult to meet the real-time and continuous requirements of large-scale spatiotemporal knowledge graphs.

Method used

A unified quality factor system across modalities is constructed, an adaptive quality factor extraction and calculation framework is designed, and an entropy weight method is used to construct a standardized evaluation mechanism to achieve unified evaluation and screening of multi-source data.

Benefits of technology

It enables accurate and efficient quality assessment of collective intelligence sensing data, supports the optimization and complementarity of multi-source data, improves assessment efficiency and adaptability, and meets the dynamic update requirements of spatiotemporal knowledge graphs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122086879A_ABST
    Figure CN122086879A_ABST
Patent Text Reader

Abstract

This invention proposes a method and system for assessing the quality of spatiotemporal information contained in crowdsourced sensing data, aiming to solve the problems of fragmented quality factors and the lack of a unified cross-modal assessment framework in existing technologies. The method includes: Step S1, constructing a common and individual cross-modal quality factor system. Common factors cover richness, accuracy, timeliness, and reliability, while individual factors adapt to six types of multimodal data features, including professional spaces and IoT sensing. Step S2, by integrating spatiotemporal statistical natural language processing, deep learning, and other technologies, designing differentiated algorithms to achieve automated extraction and calculation of quality factors. Step S3, based on the entropy weight method, completing factor normalization, objective weighting, and comprehensive scoring, dividing the data into four quality levels. This invention achieves a unified and accurate assessment of the quality of spatiotemporal information contained in multimodal data, supports cross-type quantitative comparison, and provides core support for high-quality data source selection and the construction and updating of spatiotemporal knowledge graphs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention proposes a method and system for evaluating the quality of spatiotemporal information contained in crowd-sensing data, which relates to the field of intelligent algorithm processing. Background Technology

[0002] In the construction and continuous updating of large-scale spatiotemporal knowledge graphs, collectively sensed data serves as the core foundation, and the quality of the spatiotemporal information it contains directly determines the reliability and application effectiveness of the knowledge graph. This type of data encompasses six categories of multimodal data: professional spatial data, IoT sensing data, mobile trajectory data, geographic semantic web, scientific literature, and online text. The acquisition mechanisms, organizational structures, semantic expressions, and update frequencies of different data sources vary significantly, resulting in inconsistent quality and posing a severe challenge to information fusion and quality control of cross-modal data. Therefore, constructing a unified evaluation system adaptable to the characteristics of multimodal data and achieving accurate extraction and calculation of quality factors is a key technical problem urgently needing to be solved in the current field of spatiotemporal knowledge services.

[0003] Although existing quality assessment technologies have evolved from a single "data layer" to a multi-level "data-information-knowledge" approach, they still have the following two significant technical shortcomings when dealing with complex crowd sensing data:

[0004] (1) The quality factors are fragmented and lack a unified cross-modal measurement. Early methods were only for traditional spatiotemporal data (i.e., professional spaces, IoT sensing, and mobile trajectories), focusing on using classic indicators from surveying and geographic information science (such as geometric accuracy, resolution, and topological consistency). Subsequent methods, targeting ubiquitous spatiotemporal data (i.e., semantic web, scientific literature, and online text), introduced techniques such as natural language processing (e.g., LDA topic modeling) and symbolic logic reasoning (e.g., SPARQL rule verification). However, the observation accuracy of the physical world (e.g., GPS positioning error) and the semantic quality of social perception (e.g., text relevance and source credibility) belong to different dimensional measurement spaces, lacking a unified mapping mechanism and normalization standard. This data silo-like evaluation status makes it impossible for the system to horizontally compare the value of data from different sources under the same dimension, making it difficult to support the optimization and complementarity of multi-source data.

[0005] (2) A general framework for extracting quality factors is lacking, and its automation and adaptability are insufficient. Existing evaluation schemes mostly rely on static rule bases or customized implementations of general standards (such as ISO 19157). The former heavily relies on the prior knowledge of domain experts, resulting in high rule writing costs, difficult maintenance, and poor transferability; the latter, while possessing generality, struggles to dynamically balance the weighting relationship between common indicators (such as completeness and timeliness) and individual indicators (such as location drift in trajectory sampling and semantic ambiguity in text) of the six types of crowd sensing data. When faced with massive, streaming six types of crowd sensing data, existing methods are unable to meet the real-time and continuous requirements of spatiotemporal knowledge graphs for data filtering, severely restricting the dynamic update efficiency of the graph. Summary of the Invention

[0006] In view of this, in order to address the core pain points of existing technologies such as fragmented quality factor dimensions, lack of unified cross-modal measurement standards, absence of a general extraction framework, and insufficient automation and adaptability, this invention proposes a method and system for evaluating the quality of spatiotemporal information contained in collective intelligence sensing data.

[0007] This invention aims to construct a complete technical solution from a unified factor system to an adaptive extraction algorithm to a standardized comprehensive evaluation, so as to achieve accurate and efficient quality evaluation of the spatiotemporal information contained in six types of collective intelligent sensing data: professional spatial data, IoT sensing data, mobile trajectory data, geographic semantic web, scientific and technological literature and online text, and provide high-quality data support for the construction and dynamic updating of large-scale spatiotemporal knowledge graphs.

[0008] This invention proposes a method and system for evaluating the quality of spatiotemporal information contained in crowd-sensing data, comprising the following:

[0009] This invention proposes a method for evaluating the quality of spatiotemporal information contained in crowd-sensing data, characterized by comprising the following:

[0010] Step S1: Construct a cross-modal unified quality factor system, including building a cross-modal unified quality factor system from two value levels: data source and data content. At the same time, it constructs common quality factors that cover six types of collective intelligence perception data and personalized quality factors that adapt to various types of data. Among them, the common quality factors include richness, accuracy, timeliness, and reliability, while the personalized quality factors are customized according to the differences in different data modalities.

[0011] Step S2: Design an adaptive quality factor extraction and calculation framework. Addressing the heterogeneity of the six types of swarm intelligence sensing data in terms of source, spatiotemporal representation, and structural form, this framework comprehensively utilizes technologies including spatial computation, statistical analysis, deep learning, symbolic query, and natural language processing to achieve automated and quantitative extraction of various quality factors. The adaptive quality factor extraction and calculation framework has an extensible interface, supporting flexible invocation of corresponding calculation modules based on different data modalities.

[0012] Step S3: Establish a standardized comprehensive evaluation mechanism. Use the entropy weight method to construct a standardized evaluation process. Through four steps—normalization, weighting, comprehensive scoring, and grade classification—achieve the fusion calculation of quality factors with different dimensions and cross-data type comparison output. Finally, form a unified quality score and grade classification to complete the quantitative comparison and screening of the quality of spatiotemporal information contained in multi-source data.

[0013] Further, step S1 includes the following:

[0014] Step S11: Define common quality factors, covering four core dimensions that need to be considered for all six types of crowd-sensing data. Each dimension has a clearly defined sub-indicator to ensure the comprehensiveness and universality of the evaluation.

[0015] Richness: Characterizes the completeness and detail of the spatiotemporal information contained in the data, specifically including three sub-indicators: information richness, spatiotemporal range, and spatiotemporal granularity;

[0016] Accuracy: Reflects the degree to which data matches objective reality, including sub-indicators such as physical observation error, data anomaly ratio, and semantic and logical consistency;

[0017] Timeliness: Reflects the freshness and effectiveness of data information, with data update / release time and information validity status as core sub-indicators;

[0018] Reliability: Ensuring the credibility and influence of data sources, including two sub-indicators: the qualifications of the publisher and the popularity of information dissemination;

[0019] Step S12: Define the individual quality factors, which are customized for the structural characteristics of different modal data. The individual quality factors include the following:

[0020] Professional spatial data: Attribute item standardization;

[0021] IoT sensing data: Data record integrity;

[0022] Movement trajectory data: trajectory loss rate;

[0023] Geographic Semantic Web data: proportion of spatiotemporal tuples, ontology standardization;

[0024] Scientific literature data: the proportion of spatiotemporal knowledge;

[0025] Online text data: the proportion of spatiotemporal knowledge.

[0026] Further, step S2 includes the following:

[0027] Step S21: Extraction of common factor richness, including three sub-indicators: information richness, spatiotemporal range, and spatiotemporal granularity. Customized calculation logic is designed for different data types, and finally, a statistical analysis method based on information entropy is performed, including the following:

[0028] Step S211: Topic richness extraction includes: using the principle of information entropy to quantify the diversity and information content of topics contained in the dataset. The extraction process is from topic x identification to probability p(x) calculation to information entropy H(x) solution.

[0029] ;(1);

[0030] Step S212: Spatiotemporal range extraction includes: establishing a spatial grid index, calculating the actual spatial range Area covered by the data object, and then comparing it with the theoretical global spatial range Area of ​​the scene to which the data belongs. global By comparing the values, the spatial coverage range can be obtained. spa ;

[0031] Iterate through the timestamps of the data items to determine the earliest observation time t. min With the latest observation time t max Calculate the absolute time span ∆t=t max -t min An exponential decay model is used to quantify the richness of the time span (Range). temp , where λ t Attenuation coefficient:

[0032] (2);

[0033] (3);

[0034] Step S213: Spatiotemporal granularity extraction includes: firstly, parsing the explicit resolution based on the metadata regular expression; if the metadata lacks this factor, then performing reverse calculation based on the spatiotemporal range Area, ∆t, and data volume |data| to obtain the following expression:

[0035] (4);

[0036] Here, Resolution represents the spatiotemporal granularity extraction function.

[0037] Furthermore, step S2 also includes the following:

[0038] Step S22: Accuracy extraction of common factors, using methods such as error extraction and anomaly detection for different data types, including the following:

[0039] Step S221: Measurement and time error extraction includes parsing metadata or associated authoritative error analysis reports to obtain data error values, including plane error, elevation error, sensor measurement error, positioning error, time synchronization error, etc.

[0040] Step S222: Outlier and location drift detection includes: using a three-level anomaly detection algorithm to jointly identify isolated anomalies, periodic anomalies, and trend drifting anomalies; using a location drift detection algorithm to identify drift points with spatial positional shifts; and finally calculating the number and proportion of anomalies and drifts.

[0041] The three-level anomaly detection algorithm includes: Hampel filter, STL periodicity decomposition, and sliding window statistical detection;

[0042] The positioning drift detection algorithm integrates traditional filtering, velocity and distance and angle threshold detection, and spatiotemporal clustering correction.

[0043] Step S223: Semantic Consistency Calculation: For Semantic Web or text-based data (i.e., scientific and technological documents, online text), knowledge graph embedding representation and text embedding representation are used respectively. The probability of the triple (h, r, t) being true is quantified by f(h, r, t) or the text similarity is calculated using Sim(Text, Text). * To achieve quantitative calculation, the expression is as follows:

[0044] (5);

[0045] (6);

[0046] Where sigmoid() represents the normalization function;

[0047] In the triplet, The embedding vector representing the head entity. An embedding vector representing a relation. The embedding vector representing the tail entity. The embedding vector representing the document;

[0048] Semantic Web or text-based data includes scientific and technological literature and online text;

[0049] Step S23: Evaluation of the timeliness of common factors. Based on the time reflected by the data, the update and release time, and the validity status of the information, calculate the current timeliness decay, including the following:

[0050] Step S231: Data Reflection Time Extraction: Extract the latest observation time t from the time field of the data content. max Calculate the lag difference (T) between it and the current time. current -t last );

[0051] Step S232: Dataset Publication Time Extraction: Parse metadata to obtain the dataset's publication or last update time T. release Calculate the lag difference (T) between it and the current time. current - T release );

[0052] Step S233: Information Validity Status Extraction: For information with time constraints, such as policies, announcements, and notices, in scientific and technological literature and online texts, a RoBERTa-wwm+CRF sequence labeling model is constructed to achieve high-precision extraction of time information, i.e., extracting the effective time T. start and the termination time T end Combined with the current time T current Determine the validity of the information using Valid(T);

[0053] (7);

[0054] Where Ⅱ() represents an indicator function, which has a value of 1 when the condition is met and 0 otherwise.

[0055] Furthermore, step S2 also includes the following:

[0056] Step S24: Extracting common factors for reliability, focusing on the credibility of data sources and the dissemination influence of datasets, and designing evaluation methods for each indicator, including the following:

[0057] Step S241: Publisher qualifications and type extraction includes: using a pre-trained model to extract the name description text. name Embedded representations are used, and the TextCNN text classification model is concatenated to identify four trust levels of the publisher. publisher The expression is as follows:

[0058] (8);

[0059] Where argmax represents the text extraction classification model (TextCNN()), BERT() represents the pre-trained model, W represents the weight matrix of the classifier, and b represents the bias vector.

[0060] Step S242: Calculate the spread popularity: Call the public interface of the relevant data platform and use regular expressions, XPath parsing or structured interface methods to obtain multiple factors reflecting the influence of data spread, including downloads, citations, clicks, reposts and comments.

[0061] Furthermore, step S2 also includes the following:

[0062] Step S25, the extraction and calculation of individual quality factors, proposes targeted extraction and calculation methods according to the unique quality factor dimensions of each data type, including the following:

[0063] Step S251: The normalization calculation of attribute items in the professional space includes: based on the standard specifications of the scene domain or referring to authoritative datasets, dividing the attribute fields into core attributes attr. core and auxiliary attribute attr aux Norm for setting the normalization of attr based on the attributes of the current dataset. attr , where λ aux The adaptation weight coefficient for auxiliary attributes;

[0064] (9)

[0065] Step S252: The data record integrity calculation for IoT sensing includes: traversing and querying the data, recording missing or invalid values, and calculating their proportion. ;

[0066] Step S253: The trajectory loss rate calculation includes: based on the officially stated theoretical sampling interval. * Calculate [t] within the time period min , t max Expected value of trajectory points ∆t / interval * And calculate the loss rate Ratio(loss) based on the actual trajectory data:

[0067] (10)

[0068] Step S254: Ontology normative calculation of Geographic Semantic Web data includes: focusing on the degree of conformity between the ontology model of Geographic Semantic Web and mainstream geographic information ontology standards, adopting a quantitative method of compliance verification from term to relation r to attribute attr and weighted fusion, and using query language to achieve accurate extraction and calculation, resulting in:

[0069] (11);

[0070] Step S255: Calculation of the proportion of spatiotemporal knowledge or tuples in geographic semantic web, scientific and technological literature, and online text includes: To measure the density of spatiotemporal information contained in the data, two suitable extraction and calculation methods are designed for the structural differences of the three types of data to achieve a standardized assessment of the proportion of spatiotemporal knowledge; that is, for the semantic web, predicate matching is performed through query statements to determine whether spatiotemporal tuples contain time attributes, spatial attributes, or geographical relationships; for scientific and technological literature and online text, the HanLP named entity recognition algorithm is used to label time entities and location entities from the text, count the number of text units containing at least one spatiotemporal entity or description, and calculate their distribution density in the whole text.

[0071] (12);

[0072] Where unit is a tuple or text unit, and Ratio(spatial, temporal) represents the proportion of spatiotemporal knowledge or tuples.

[0073] Further, step S3 includes the following:

[0074] Step S31: Quality factor normalization processing includes: For the various quality factors ind extracted in step S2, the range normalization method is first used to eliminate dimensional differences. If the factor ind is a negative indicator, it needs to be positively normalized to ensure that all factors are in the same direction.

[0075] (12);

[0076] (13)

[0077] Among them, ind min with ind max These are the minimum and maximum values ​​of the factor across all datasets to be evaluated, respectively.

[0078] Step S32 Entropy Weighting Method: To avoid subjective weighting bias, the entropy weighting method is adopted. Based on the dispersion of the distribution of each common quality factor, its objective weight is calculated, including the following:

[0079] Step S321: Construct a common quality factor matrix: Given n datasets to be evaluated, each containing m normalized quality factors, construct... Dimensional quality factor matrix:

[0080] (15);

[0081] Step S322 Calculate the scaling matrix: To avoid zero values ​​in subsequent logarithmic operations, a very small constant is introduced: ; Calculate the weight p of the i-th dataset under the j-th factor. ij :

[0082] (16)

[0083] Step S323 Calculate the information entropy of the commonality factor: Information entropy e j E reflects the uncertainty of the j-th quality factor. j The smaller the value, the greater the variation of the factor across different datasets.

[0084] (16)

[0085] Step S324: Calculate the objective weights of each common factor: First, calculate the difference coefficient (1-e) of the j-th factor. j A larger value indicates a greater contribution of the factor to the quality assessment; subsequently, the difference coefficient is normalized to obtain the final weight w. j for:

[0086] (17)

[0087] Step S33: Calculate the common factor score: For each dataset, sum the normalized values ​​of the common quality factors with their corresponding weights, and convert the sum to a percentage score to obtain the common factor score for that dataset. com :

[0088] (18)

[0089] Step S34: Calculation of Personality Factor Scores: The theoretical value normalization method is used to process the data. All applicable personality factors are normalized according to formula (12) in step S31 and averaged to convert them into a percentage-based personality factor score. per .

[0090] Furthermore, step S3 also includes the following:

[0091] The stratified weighted composite score in step S35 includes: using a stratified weighted method to integrate the scores of common factors and individual factors to calculate the final composite quality score, whereby... , Given weights;

[0092] (19)

[0093] The quality grading in step S35 includes: based on the overall quality score. i Based on the actual quality requirements for constructing spatiotemporal knowledge graphs, dataset i is divided into the following four quality levels, and corresponding usage and selection strategies are formulated:

[0094] When Score i A score of ≥90 indicates that the dataset is preferred and can be directly used for the construction and updating of spatiotemporal knowledge graphs.

[0095] When 75≤ Score i When the value is less than 90, it indicates that the dataset is qualified and usable, meaning that the dataset has been slightly optimized or verified in a specific application scenario based on the requirements.

[0096] When 60≤ Score i If the score is less than 75, it indicates that the dataset basically meets the requirements, but it also indicates that the dataset has obvious defects. It is recommended to restrict its use and supplement the data or correct the quality defects.

[0097] When Score i If the score is less than 60, or if the key quality factor fails to meet the standard, it indicates that the dataset does not meet the basic quality requirements and is prohibited from being used for the construction of spatiotemporal knowledge graphs.

[0098] According to a second aspect of the present invention, the present invention provides a system for assessing the quality of spatiotemporal information contained in crowd sensing data, comprising an electronic device, wherein the electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements a method for assessing the quality of spatiotemporal information contained in crowd sensing data as described in any one of the present invention.

[0099] According to a third aspect of the present invention, the present invention provides a system for assessing the quality of spatiotemporal information contained in crowd-sensing data, comprising a computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements any of the methods for assessing the quality of spatiotemporal information contained in crowd-sensing data as described in the present invention.

[0100] The present invention has the following advantages:

[0101] The method of the present invention has the following advantages compared with existing methods:

[0102] This invention breaks through the bottleneck of unified cross-modal evaluation: addressing the pain points of fragmented quality factors and lack of unified measurement standards in existing technologies, it constructs a "common + individual" cross-modal quality factor system, which not only covers the core quality dimensions of six types of collective intelligence sensing data, but also adapts to the unique attributes of various types of data. Through a unified normalization standard, it realizes cross-dimensional mapping between physical observation accuracy and semantic quality, and supports cross-modal quality horizontal comparison.

[0103] This invention enhances automated adaptation capabilities: it overcomes the problems of existing solutions relying on static rules and having poor transferability by designing a differentiated extraction algorithm framework that integrates spatiotemporal statistical analysis, deep learning, and natural language processing technologies. It enables automated and quantitative extraction of quality factors for multi-source heterogeneous data and has an extensible interface to support rapid adaptation to new data types.

[0104] This invention achieves objective, accurate, and efficient evaluation: it avoids the problems of subjective weighting bias and complex and time-consuming processes in existing methods, and adopts the entropy weight method to objectively calculate factor weights based on the degree of data dispersion. It improves evaluation efficiency through a standardized process of "normalization-weighting-scoring-grading" to achieve near real-time performance; it refines the four-level quality level classification and application suggestions, provides a clear basis for data source selection, and provides reliable support for the construction of spatiotemporal knowledge graphs. Attached Figure Description

[0105] Figure 1 This is a schematic diagram of the steps of the present invention.

[0106] Figure 2 This is a schematic diagram of the quality assessment process based on multi-source swarm intelligence sensing data according to the present invention. Detailed Implementation

[0107] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings.

[0108] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0109] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments of the present invention; as used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise; furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0110] like Figures 1 to 2 As shown, this invention proposes a method and system for evaluating the quality of spatiotemporal information contained in crowd-sensing data, characterized by comprising the following:

[0111] The present invention aims to propose a quality assessment method based on professional spatial data, IoT sensing, mobile trajectory, geographic semantic web, scientific and technological literature, and online text. Detailed implementation details of some quality factors in the above invention include the following:

[0112] Step S21: Common Factor Richness Extraction

[0113] Step S211: Theme Richness Extraction

[0114] In one embodiment of the present invention, in professional spatial data, vector data calculates the frequency p(x) by counting the frequency of each attribute topic in the dataset through attribute items; while raster data calculates the frequency p(x) by dividing the number of raster cells occupied by the category represented by "value" by the total number of raster cells occupied by the total number of categories. For example, the "value" column of each row represents the terrain type of a raster. The "value" column is extracted to obtain a list of all terrain types. The number of times each terrain type appears is counted, and the total number of all terrain types (i.e., the total number of raster cells) is calculated. That is, the number of times each type appears is divided by the total number to obtain its frequency p(x); while for POI data with two or more key attribute fields, the attributes are combined, all locations are traversed, the type of each location is split, (type, city) pairs are generated, the total number is counted, and the frequency of the combined attribute pairs is calculated to calculate the frequency p(x); then the entropy value is calculated using the information entropy formula (1).

[0115] In one embodiment of the present invention, in order to reduce the number of unique values ​​in IoT sensing data, the measured values ​​are discretized; an equal-frequency binning method is used to avoid bin vacancy problems caused by uneven value ranges; the number of bins k is determined using the statistical Sturges formula, the data points are divided into k bins, and the frequency p(i) of each bin is calculated. Finally, the entropy value is calculated according to the information entropy formula (1), and the relevant formula includes the following:

[0116] (20);

[0117] ;(twenty one);

[0118] Where n is the number of observation points (samples) in the dataset.

[0119] In one embodiment of the present invention, the movement trajectory data is analyzed based on the number of times L of the discrete set of locations (such as POIs, grids) visited by the user. i And the total number of times users visited all locations L total Calculate the probability p(i) of a user visiting a specific location, and calculate its entropy value according to the information entropy formula (1).

[0120] ;(twenty two);

[0121] In one embodiment of the present invention, in the geographic semantic web data, the rdf:type or gn:featureClass of all entities is extracted, the total number of entities N in the statistical data set and the number of times each category appears n are counted, the probability p(i) is calculated, and then its entropy value is calculated according to the information entropy formula (1).

[0122] ;(twenty three);

[0123] In one embodiment of the present invention, in scientific and technological literature data, documents are segmented, a sufficient number of documents are generated to train the model, and the model is based on the topic-specific word consistency index C. v Select the optimal number of topics n, and calculate the term frequency p(x) of the topic words based on LDA (Latent Dirichlet Allocation) topic extraction using TF-IDF (Term Frequency-Inverse Document Frequency). Then, combine it with the information entropy formula (1) to calculate its entropy value.

[0124] ;(twenty four);

[0125] In one embodiment of the present invention, in the network text data, the data is first preprocessed by text cleaning, word segmentation and removal of common words, and then a dictionary is constructed to count the word frequency p(x), and the information entropy value is calculated according to the information entropy formula (1).

[0126] Step S212 Spatiotemporal range extraction:

[0127] In one embodiment of the present invention, for textual data, spatial range information naming recognition and structuring are performed using natural language processing algorithms. A RoBERTa-wwm+CRF sequence labeling model is constructed to extract the temporal and spatial ranges of textual data: the textual data to be processed is acquired and normalized, the normalization process including at least denoising, sentence segmentation, length truncation and padding, and retaining the original character position index; a sequence labeling tag system is constructed, including a temporal range label set and a spatial range label set, and BIO annotation is used to label training samples to obtain supervised data; the RoBERTa-wwm tokenizer is used to segment the text and generate token sequences and their alignment mapping with the original characters, the token sequences are converted into model input vectors and input into the RoBERTa-wwm encoder to obtain each token. The contextual semantic representation is obtained; the contextual semantic representation is input into a linear layer to obtain the emission score of each tag corresponding to each token, and then input into a CRF layer to perform global optimal tag sequence decoding in combination with tag transition constraints, outputting a tag sequence of time range and spatial range; based on the tag sequence, adjacent tokens of the same type are merged into entity fragments and written back as the original text character span to obtain structured extraction results, wherein the time range results include time interval fragments and their start and end boundaries, and the spatial range results include place names / geographical entities or spatial range phrases; the training set is used to jointly train RoBERTa-wwm and CRF, and the model parameters are fixed after parameter tuning based on the validation set, which is used to automatically extract the time range and spatial range of scientific and technological documents and online texts.

[0128] (25);

[0129] (26);

[0130] (27)

[0131] Where: h i The context semantic vector output by the RoBERTa-wwm encoder for the i-th token; W and h are the parameters of the linear mapping layer; s i Let s be the emission score vector for each tag corresponding to the nth Token. i , y y represents the normalized score of the nth token labeled as tag y; Score(x, y) is the CRF sequence score; y* (the optimal tag sequence output during the inference phase).

[0132] Step S213 Spatiotemporal granularity extraction:

[0133] In one embodiment of the present invention, for the quality factors present in the metadata, the system parses the explicit resolution based on the metadata regular expression.

[0134] In one embodiment of the present invention, if the metadata lacks this factor, reverse calculation is performed, and the trajectory point sampling frequency Freq is calculated. traj It is obtained from the ratio of the total number of sampling points to the total duration Δt.

[0135] (28)

[0136] In one embodiment of the present invention, information resolution is achieved by performing Chinese word segmentation on the text to be evaluated (Text) and filtering out preset punctuation marks to obtain a word sequence, while simultaneously performing part-of-speech tagging. Subsequently, multi-dimensional granular indicators are extracted from the text. To construct a concise evaluation model, the seven indicators are categorized into three groups based on their weighted contributions: the core indicator, detail word density D. detail Assign a weight of 0.20; then calculate lexical diversity TTR, word frequency distribution entropy H, and digit density D. num and conjunction count C con A uniform weight of 0.15 was assigned; the average word length was L. avg ; and TF-IDF variance V tfidf Assign a weight of 0.10. Finally, standardize each indicator and subtract it into the following formula for weighted summation to obtain a granularity score between 0 and 1. Based on the score, the quality factors are divided into: coarse granularity (0-0.4), medium granularity (0.4-0.7), and fine granularity (0.7-1).

[0137] (29);

[0138] Step S22: Accurate extraction of common factors:

[0139] Step S221 Measurement / Time Error Extraction:

[0140] In one embodiment of the present invention, for each error present in the metadata, the system parses out the specific error value based on the metadata regular expression.

[0141] Step S222 Outlier / Location Drift Detection:

[0142] In one embodiment of the present invention, outlier detection is performed using a three-level algorithm: the first is Hampel point anomaly, which identifies mutations / glitches (such as interference or packet loss) within a sliding window based on a dynamic threshold of median and MAD; the second is STL periodic anomaly, which decomposes the sequence into trend / period / residual, and identifies anomalies when the residual exceeds the high quantile (e.g., 99.5%), used to detect fluctuations outside the periodic regularity (e.g., periodic drift or overload); the third is window statistical drift anomaly, which compares the changes in mean / variance of adjacent windows, and identifies drift when the value exceeds the threshold, used to identify chronic failures and baseline drift (e.g., aging or gradual environmental changes). The combined formula of the three-level outlier detection algorithm is as follows:

[0143] (30)

[0144] Where: x t The original sequence; Robustness scale for Hampel point anomalies; STL decomposition into x t =s t +τ t +r t r t For residuals (periodic anomalies are reflected in the residuals); σ is the mean of the sliding window (trend drift is characterized by the difference in statistics between adjacent windows); r , σ μ The corresponding scale (which can be the standard deviation or the robust scale) is λ, and the uniform threshold is λ.

[0145] The percentage of outliers is calculated by combining the number of outliers obtained from the outlier detection algorithm with the total amount of data.

[0146] (31);

[0147] In one embodiment of the present invention, a positioning drift detection algorithm (integrating traditional filtering, velocity / distance / angle threshold detection, and spatiotemporal clustering correction) is used to identify drift points with spatial position shifts. The drift point fusion detection formula is as follows:

[0148] (32);

[0149] in: For the original positioning p t Position after traditional filtering;

[0150] Where Δv t Δd t ,Δα t Indicates: by The calculated changes in speed / distance / heading angle;

[0151] θv, θd, θα: corresponding thresholds;

[0152] Clust() represents the spatiotemporal clustering correction operator.

[0153] ;

[0154] The positioning accuracy of the data is calculated based on the number of drift points obtained from the positioning drift detection algorithm and the total amount of data.

[0155] ;(33);

[0156] Step S223 Semantic consistency calculation:

[0157] In one embodiment of the present invention, when the object to be evaluated is semantic web data, a knowledge graph embedding representation method (such as TransE, RotatE) is used to obtain the set of triples to be evaluated, 𝑇. The entity and relation are vectorized into a representation T={(h,r,t)} using a knowledge graph embedding representation method (such as TransE, RotatE). Based on the embedding model, the validity score of any triple (ℎ, 𝑟, 𝑡) is calculated and mapped to obtain the validity probability. Then, the validity probabilities of the triples are aggregated to obtain the semantic web semantic consistency score. Taking TransE as an example, the calculation method of the validity probability and the aggregated score is as follows:

[0158] (34);

[0159] ;(35);

[0160] In the triplet, The embedding vector representing the head entity. An embedding vector representing a relation. The embedding vector representing the tail entity. The embedding vector representing the document;

[0161] || represents the L2 norm; sigmoid is the normalization mapping function used to map the distance result to the (0,1) interval as the probability value of the triple semantics; Score KG Semantic consistency score of the Semantic Web.

[0162] When the data to be evaluated is text-based, text embedding representation (such as BERT) is used to divide the text into several semantic units according to sentences or semantic paragraphs, obtaining the set of texts to be evaluated, D={d i} and the authoritative reference text set A={a i The text to be evaluated and the authoritative reference text are encoded into vector representations using text embedding representation methods (such as BERT). The semantic similarity between the text to be evaluated and the authoritative reference text is calculated, and the semantic similarity is aggregated to obtain the text semantic consistency score. The calculation method of the text matching and aggregation score is shown in the following formula.

[0163] ;(36);

[0164] (37)

[0165] in: This is a vector representation of the semantic unit of the text to be evaluated; A vector representation of authoritative reference text; a * (i) For each text to be evaluated Retrieves the closest reference text from a set of authoritative sources; Score Text Text semantic consistency score.

[0166] Step S23: Timeliness evaluation of common factors:

[0167] Step S233: Extracting valid information status:

[0168] For information with time constraints, such as policies, announcements, and notices, contained in scientific and technological literature and online texts, a RoBERTa-wwm+CRF sequence labeling model is constructed to achieve a high-precision extraction method for time information, which is the same as the spatiotemporal range extraction method in step S212.

[0169] Step S24: Reliability extraction of common factors:

[0170] Step S241: Extracting Publisher Qualifications / Types:

[0171] In one embodiment of the present invention, a parallel dual-path feature fusion model of "BERT+TextCNN" is constructed, combining global semantic understanding and local pattern recognition to achieve accurate quantification of the trustworthiness level of the publisher's identity. This involves obtaining the publisher's name and description text (Text). name This data is then fed into the pre-trained BERT model, where a multi-layer self-attention mechanism is used to extract the global semantic vector representation of the text. The vector corresponding to the start flag [CLS] output from the last layer of BERT is extracted as the global feature vector. , where d1 is the hidden layer dimension of BERT. Text name Input is processed into the TextCNN convolutional neural network. Multi-scale convolutional kernels capture local key patterns with eligibility features (such as suffixes like "official," "certified," and "bureau / department") in the text. After convolution and global max pooling operations, the output is a local key feature vector. d2 represents the total number of features extracted by the convolution kernel. Finally, the feature vectors output from the two branches are concatenated and fused to construct an enhanced publisher feature representation vector. This fused feature vector is then input into a linear classification layer, where it undergoes Softmax normalization and maximum value optimization to ultimately determine the publisher's rank category.

[0172] (38)

[0173] Where: BERT(·) is the global semantic vector mapping function performed by the BERT pre-trained model; TextCNN(·) is the local feature mapping function performed by the convolutional neural network; W is the weight matrix of the classification layer; b is a learnable bias term; class publisher The publisher is categorized into four levels: government or official authority (1 point), certified group (0.85 points) or representative, certified individual (0.7 points), and ordinary individual (0.55 points).

[0174] Step S242: Calculation of heat propagation:

[0175] In one embodiment of the present invention, regular expressions, XPath parsing, or structured interface methods are used to obtain multiple factors reflecting the influence of data dissemination, including but not limited to downloads, citations, clicks, reposts, and comments.

[0176] The algorithm for extracting the propagation heat factor is shown in Table 1 below:

[0177] Table 1 Algorithm for Extracting Propagation Heat Factor

[0178] Algorithm S242 Propagation Heat Factor Extraction Input: The object to be evaluated, obj (including URL or platform ID); Public data interfaces / page access methods APIs (webpage source code, community platform interfaces, citation databases); Parsing rules (Regex, XPath, FieldPath) Output: A set of factors contributing to the spread of information, F (one or more of the following: downloads, citations, clicks, shares, and comments). 1: F<- empty_map 2: src<- Fetch(obj, API) / / Retrieves publicly returned content 3: type<- DetectType(src) / / type ∈ {HTML, STRUCTURED} 4: if (type == HTML) then 5: html<- src.content 6: F["download"]<- Parse(html, Rule["download"], method ∈ {Regex, XPath}) 7: F["citation"]<- Parse(html, Rule["citation"], method ∈ {Regex, XPath}) 8: F["view"]<- Parse(html, Rule["view"], method ∈ {Regex, XPath}) 9: F["share"]<- Parse(html, Rule["share"], method ∈ {Regex, XPath}) 10: F["comment"]<- Parse(html, Rule["comment"], method ∈ {Regex, XPath}) 11: else 12: data<- src.content / / Returns structured data such as JSON / XML 13: F["download"]<- Parse(data, Rule["download"], method == FieldPath) 14: F["citation"]<- Parse(data, Rule["citation"], method == FieldPath) 15: F["view"]<- Parse(data, Rule["view"], method == FieldPath) 16: F["share"]<- Parse(data, Rule["share"], method == FieldPath) 17: F["comment"]<- Parse(data, Rule["comment"], method == FieldPath) 18: end if 19: return F

[0179] The detailed code for the propagation heat factor extraction algorithm is presented in Table 1.

[0180] Step S25: Extraction and Calculation of Individual Quality Factors:

[0181] Step S254: Ontology normalization calculation of geographic semantic web data:

[0182] In one embodiment of the present invention, feature detection is performed on the dataset under test using the SPARQL graph database query language. At the metadata specification level, it searches for the presence of the `rdf:type` predicate to define entity categories and verifies the usage of RDFS core terms such as `rdfs:isDefinedBy` and `rdfs:seeAlso`. At the geographic domain specification level, it matches the `gn:Feature` class and core attribute predicates such as `gn:name` and `gn:countryCode` in the GeoNames Ontology specification. At the spatial vocabulary specification level, it searches for W3C Geo (WGS84) related attributes such as `wgs84:lat` and `wgs84:long`, and performs data type validation on the extracted coordinate literals to determine if they conform to the floating-point (Float) numerical specification. Furthermore, regarding the interconnectivity of Linked Open Data (LOD), it statistically analyzes external association predicates such as `foaf:page` and uses regular expressions to perform protocol validity checks on the referenced Uniform Resource Identifiers (URIs) to verify whether they conform to preset access protocol standards. Based on the above detection results, a weighted fusion method is used to calculate the ontology normativity index Match(ontology).

[0183] Step S255: Calculation of the proportion of spatiotemporal knowledge (or tuples) in geographic semantic web, scientific literature, and online text:

[0184] In one embodiment of the present invention, the HanLP named entity recognition algorithm is employed. The model's shared representation layer extracts the deep semantic vector of the text's context, and its joint inference interface for word segmentation and part-of-speech tagging is invoked. The model captures long-distance dependencies through a self-attention mechanism, outputting a sequence W of terms, and simultaneously generating a sequence W for each term. j Generate corresponding part-of-speech tags t j A spatiotemporal feature mapping table is constructed, and terms are classified according to their attributes based on the part-of-speech tagging results: Entities tagged with time words (t), time adverbs, or those with time reference are identified as time terms. Entities tagged with place names (ns), organization names (nt), locative words (f), or latitude and longitude descriptors are identified as spatial terms. If a text unit contains at least one entity of the above categories, or contains feature morphemes with spatiotemporal indication, it is marked as a text unit containing spatiotemporal information, and its distribution density in the entire text is calculated.

[0185] (39)

[0186] In addition to the above, the present invention also has related embodiments, including the following:

[0187] This invention randomly selected one set of data each from the collective intelligence sensing dataset: professional space, IoT sensing, mobile trajectory, semantic web, online text, and scientific literature. The quality factor extraction results for each data set are shown in Tables 2 to 5. Using the aforementioned comprehensive score and quality assessment calculation method, the final results are shown in Table 2. This data quality level assessment provides a decision on whether to use the data. Common factor quality factor values ​​and individual factor quality factor values ​​are shown in Tables 3 and 4. Table 5 presents the quality assessment results for each data set. From Tables 2 to 5, it can be seen that the method and system proposed in this invention have strong practicality.

[0188] In one embodiment of the present invention, the selected data is described in Table 2:

[0189] Table 2. Description of Data Used in Examples

[0190]

[0191] The data used in the examples are described in Table 2.

[0192] In one embodiment of the present invention, the common factor quality factor values ​​are shown in Table 3:

[0193] Table 3 Common Factor Quality Factor Values

[0194]

[0195] The common factor quality factor values ​​are presented in Table 3.

[0196] In one embodiment of the present invention, the quality factor values ​​of the personality factor are shown in Table 4:

[0197] Table 4. Quality Factor Values ​​of Personality Factors

[0198]

[0199] The values ​​of the personality factor and the quality factor are shown in Table 4.

[0200] In one embodiment of the present invention, the data quality assessment results are shown in Table 5:

[0201] Table 5. Results of data quality assessment

[0202]

[0203] The results of the data quality assessment are presented in Table 5.

[0204] If the functions described in this invention are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0205] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

[0206] The above are preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention shall fall within the protection scope of the present invention.

Claims

1. A method for assessing the quality of spatiotemporal information contained in crowd-sensing data, characterized in that, Includes the following: Step S1: Construct a cross-modal unified quality factor system, including building a cross-modal unified quality factor system from two value levels: data source and data content. At the same time, it constructs common quality factors that cover six types of collective intelligence perception data and personalized quality factors that adapt to various types of data. Among them, the common quality factors include richness, accuracy, timeliness, and reliability, while the personalized quality factors are customized according to the differences in different data modalities. Step S2: Design an adaptive quality factor extraction and calculation framework. Addressing the heterogeneity of the six types of swarm intelligence sensing data in terms of source, spatiotemporal representation, and structural form, this framework comprehensively utilizes technologies including spatial computation, statistical analysis, deep learning, symbolic query, and natural language processing to achieve automated and quantitative extraction of various quality factors. The adaptive quality factor extraction and calculation framework has an extensible interface, supporting flexible invocation of corresponding calculation modules based on different data modalities. Step S3: Establish a standardized comprehensive evaluation mechanism. Use the entropy weight method to construct a standardized evaluation process. Through four steps—normalization, weighting, comprehensive scoring, and grade classification—achieve the fusion calculation of quality factors with different dimensions and cross-data type comparison output. Finally, form a unified quality score and grade classification to complete the quantitative comparison and screening of the quality of spatiotemporal information contained in multi-source data.

2. The method for assessing the quality of spatiotemporal information contained in crowd sensing data according to claim 1, characterized in that, Step S1 includes the following: Step S11: Define common quality factors, covering four core dimensions that need to be considered for all six types of crowd-sensing data. Each dimension has a clearly defined sub-indicator to ensure the comprehensiveness and universality of the evaluation. Richness: Characterizes the completeness and detail of the spatiotemporal information contained in the data, specifically including three sub-indicators: information richness, spatiotemporal range, and spatiotemporal granularity; Accuracy: Reflects the degree to which data matches objective reality, including sub-indicators such as physical observation error, data anomaly ratio, and semantic and logical consistency; Timeliness: Reflects the freshness and effectiveness of data information, with data update / release time and information validity status as core sub-indicators; Reliability: Ensuring the credibility and influence of data sources, including two sub-indicators: the qualifications of the publisher and the popularity of information dissemination; Step S12: Define the individual quality factors, which are customized for the structural characteristics of different modal data. The individual quality factors include the following: Professional spatial data: Attribute item standardization; IoT sensing data: Data record integrity; Movement trajectory data: trajectory loss rate; Geographic Semantic Web data: proportion of spatiotemporal tuples, ontology standardization; Scientific literature data: the proportion of spatiotemporal knowledge; Online text data: the proportion of spatiotemporal knowledge.

3. The method for assessing the quality of spatiotemporal information contained in crowd sensing data according to claim 1, characterized in that, Step S2 includes the following: Step S21: Extraction of common factor richness, including three sub-indicators: information richness, spatiotemporal range, and spatiotemporal granularity. Customized calculation logic is designed for different data types, and finally, a statistical analysis method based on information entropy is performed, including the following: Step S211: Topic richness extraction includes: using the principle of information entropy to quantify the diversity and information content of topics contained in the dataset. The extraction process is from topic x identification to probability p(x) calculation to information entropy H(x) solution. ;(1); Step S212: Spatiotemporal range extraction includes: establishing a spatial grid index, calculating the actual spatial range Area covered by the data object, and then comparing it with the theoretical global spatial range Area of ​​the scene to which the data belongs. global By comparing the values, the spatial coverage range can be obtained. spa ; Iterate through the timestamps of the data items to determine the earliest observation time t. min With the latest observation time t max Calculate the absolute time span ∆t=t max -t min An exponential decay model is used to quantify the richness of the time span (Range). temp , where λ t Attenuation coefficient: ;(2); ;(3); Step S213: Spatiotemporal granularity extraction includes: firstly, parsing the explicit resolution based on the metadata regular expression; if the metadata lacks this factor, then performing reverse calculation based on the spatiotemporal range Area, ∆t, and data volume |data| to obtain the following expression: ;(4); Here, Resolution represents the spatiotemporal granularity extraction function.

4. The method for assessing the quality of spatiotemporal information contained in crowd-sensing data according to claim 3, characterized in that, Step S2 also includes the following: Step S22: Accuracy extraction of common factors, using methods such as error extraction and anomaly detection for different data types, including the following: Step S221: Measurement and time error extraction includes parsing metadata or associated authoritative error analysis reports to obtain data error values, including plane error, elevation error, sensor measurement error, positioning error, time synchronization error, etc. Step S222: Outlier and location drift detection includes: using a three-level anomaly detection algorithm to jointly identify isolated anomalies, periodic anomalies, and trend drifting anomalies; using a location drift detection algorithm to identify drift points with spatial positional shifts; and finally calculating the number and proportion of anomalies and drifts. The three-level anomaly detection algorithm includes: Hampel filter, STL periodicity decomposition, and sliding window statistical detection; The positioning drift detection algorithm integrates traditional filtering, velocity and distance and angle threshold detection, and spatiotemporal clustering correction. Step S223: Semantic Consistency Calculation: For Semantic Web or text-based data (i.e., scientific and technological documents, online text), knowledge graph embedding representation and text embedding representation are used respectively. The probability of the triple (h, r, t) being true is quantified by f(h, r, t) or the text similarity Sim(Text, Text) is calculated. ∗ To achieve quantitative calculation, the expression is as follows: );(5); ;(6); Where sigmoid() represents the normalization function; In the triplet, The embedding vector representing the head entity. An embedding vector representing a relation. The embedding vector representing the tail entity. The embedding vector representing the document; Semantic Web or text-based data includes scientific and technological literature and online text; Step S23: Evaluation of the timeliness of common factors. Based on the time reflected by the data, the update and release time, and the validity status of the information, calculate the current timeliness decay, including the following: Step S231: Data Reflection Time Extraction: Extract the latest observation time t from the time field of the data content. max Calculate the lag difference (T) between it and the current time. current -t last ); Step S232: Dataset Publication Time Extraction: Parse metadata to obtain the dataset's publication or last update time T. release Calculate the lag difference (T) between it and the current time. current - T release ); Step S233: Information Validity Status Extraction: For information with time constraints, such as policies, announcements, and notices, in scientific and technological literature and online texts, a RoBERTa-wwm+CRF sequence labeling model is constructed to achieve high-precision extraction of time information, i.e., extracting the effective time T. start and the termination time T end Combined with the current time T current Determine the validity of the information using Valid(T); ;(7); Where Ⅱ() represents an indicator function, which has a value of 1 when the condition is met and 0 otherwise.

5. The method for assessing the quality of spatiotemporal information contained in crowd sensing data according to claim 4, characterized in that, Step S2 also includes the following: Step S24: Extracting common factors for reliability, focusing on the credibility of data sources and the dissemination influence of datasets, and designing evaluation methods for each indicator, including the following: Step S241: Publisher qualifications and type extraction includes: using a pre-trained model to extract the name description text. name Embedded representations are used, and the TextCNN text classification model is concatenated to identify four trust levels of the publisher. publisher The expression is as follows: ;(8); Where argmax represents the text extraction classification model (TextCNN()), BERT() represents the pre-trained model, W represents the weight matrix of the classifier, and b represents the bias vector. Step S242: Calculate the spread popularity: Call the public interface of the relevant data platform, and use regular expressions, XPath parsing or structured interface methods to obtain multiple factors reflecting the influence of data spread, including downloads, citations, clicks, reposts and comments.

6. The method for assessing the quality of spatiotemporal information contained in crowd sensing data according to claim 5, characterized in that, Step S2 also includes the following: Step S25, the extraction and calculation of individual quality factors, proposes targeted extraction and calculation methods according to the unique quality factor dimensions of each data type, including the following: Step S251: The normalization calculation of attribute items in the professional space includes: based on the standard specifications of the scene domain or referring to authoritative datasets, dividing the attribute fields into core attributes attr. core and auxiliary attribute attr aux Norm for setting the normalization of attr based on the attributes of the current dataset. attr , where λ aux The adaptation weight coefficient for auxiliary attributes; ;(9); Step S252: The data record integrity calculation for IoT sensing includes: traversing and querying the data, recording missing or invalid values, and calculating their proportion. ; Step S253: The trajectory loss rate calculation includes: based on the officially stated theoretical sampling interval. ∗ Calculate [t] within the time period min , t max Expected value of trajectory points ∆t / interval ∗ And calculate the loss rate Ratio(loss) based on the actual trajectory data: ;(10); Step S254: Ontology normative calculation of Geographic Semantic Web data includes: focusing on the degree of conformity between the ontology model of Geographic Semantic Web and mainstream geographic information ontology standards, adopting a quantitative method of compliance verification from term to relation r to attribute attr and weighted fusion, and using query language to achieve accurate extraction and calculation, resulting in: ;(11); Step S255: Calculation of the proportion of spatiotemporal knowledge or tuples in geographic semantic web, scientific and technological literature, and online text includes: To measure the density of spatiotemporal information contained in the data, two suitable extraction and calculation methods are designed for the structural differences of the three types of data to achieve a standardized assessment of the proportion of spatiotemporal knowledge; that is, for the semantic web, predicate matching is performed through query statements to determine whether spatiotemporal tuples contain time attributes, spatial attributes, or geographical relationships; for scientific and technological literature and online text, the HanLP named entity recognition algorithm is used to label time entities and location entities from the text, count the number of text units containing at least one spatiotemporal entity or description, and calculate their distribution density in the whole text. ;(12); Where unit is a tuple or text unit, and Ratio(spatial, temporal) represents the proportion of spatiotemporal knowledge or tuples.

7. The method for assessing the quality of spatiotemporal information contained in crowd sensing data according to claim 1, characterized in that, Step S3 includes the following: Step S31: Quality factor normalization processing includes: For the various quality factors ind extracted in step S2, the range normalization method is first used to eliminate dimensional differences. If the factor ind is a negative indicator, it needs to be positively normalized to ensure that all factors are in the same direction. ;(12); ;(13); Among them, ind min with ind max These are the minimum and maximum values ​​of the factor across all datasets to be evaluated, respectively. Step S32 Entropy Weighting Method: To avoid subjective weighting bias, the entropy weighting method is adopted. Based on the dispersion of the distribution of each common quality factor, its objective weight is calculated, including the following: Step S321: Construct a common quality factor matrix: Given n datasets to be evaluated, each containing m normalized quality factors, construct... Dimensional quality factor matrix: ;(15); Step S322 Calculate the scaling matrix: To avoid zero values ​​in subsequent logarithmic operations, a very small constant is introduced: ; Calculate the j-th factor under the j-th factor The proportion of each dataset : ;(16); Step S323 Calculate the information entropy of the commonality factor: Information entropy e j E reflects the uncertainty of the j-th quality factor. j The smaller the value, the greater the variation of the factor across different datasets. ;(16); Step S324: Calculate the objective weights of each common factor: First, calculate the difference coefficient (1-e) of the j-th factor. j A larger value indicates a greater contribution of the factor to the quality assessment; subsequently, the difference coefficient is normalized to obtain the final weight w. j for: ;(17); Step S33: Calculate the common factor score: For each dataset, sum the normalized values ​​of the common quality factors with their corresponding weights, and convert the sum to a percentage score to obtain the common factor score for that dataset. com : ;(18); Step S34: Calculation of Personality Factor Scores: The theoretical value normalization method is used to process the data. All applicable personality factors are normalized according to formula (12) in step S31 and averaged to convert them into a percentage-based personality factor score. per .

8. The method for assessing the quality of spatiotemporal information contained in crowd sensing data according to claim 7, characterized in that, Step S3 also includes the following: The stratified weighted composite score in step S35 includes: using a stratified weighted method to integrate the scores of common factors and individual factors to calculate the final composite quality score, whereby... , Given weights; ;(19); The quality grading in step S35 includes: based on the overall quality score. i Based on the actual quality requirements for constructing spatiotemporal knowledge graphs, dataset i is divided into the following four quality levels, and corresponding usage and selection strategies are formulated: When Score i A score of ≥90 indicates that the dataset is preferred and can be directly used for the construction and updating of spatiotemporal knowledge graphs. When 75≤ Score i When the value is less than 90, it indicates that the dataset is qualified and usable, meaning that the dataset has been slightly optimized or verified in a specific application scenario based on the requirements. When 60≤ Score i If the score is less than 75, it indicates that the dataset basically meets the requirements, but it also indicates that the dataset has obvious defects. It is recommended to restrict its use and supplement the data or correct the quality defects. When Score i If the score is less than 60, or if the key quality factor fails to meet the standard, it indicates that the dataset does not meet the basic quality requirements and is prohibited from being used for the construction of spatiotemporal knowledge graphs.

9. A system for assessing the quality of spatiotemporal information contained in crowd-sensing data, comprising an electronic device, wherein the electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements a method for assessing the quality of spatiotemporal information contained in crowd sensing data as described in any one of claims 1 to 8.

10. A system for assessing the quality of spatiotemporal information contained in crowd-sensing data, comprising a computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements any one of the methods for assessing the quality of spatiotemporal information contained in crowd-sensing data as described in claims 1 to 8.