A method for urban data collection for the Internet of Things

By combining spatial distribution and geometric feature similarity with LiDAR for secondary screening, the problem of interference in neighborhood point screening by neighborhood radius in existing technologies is solved, thereby improving the accuracy and reliability of urban data collection and supporting the construction of smart cities.

CN120763486BActive Publication Date: 2025-11-14XIAN XINGXUN INTELLIGENT COMM TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511269599.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-11-14
Estimated Expiration
2045-09-08

AI Technical Summary

Technical Problem

Existing statistical filtering methods rely solely on neighborhood radius to select neighborhood points in urban environments, which makes them susceptible to interference points, leading to reduced accuracy and reliability of data collection.

Method used

A comprehensive scan using a laser mapping radar is employed to acquire mapping data points. Neighboring points are initially identified by setting a neighborhood radius, and secondary screening is performed by combining spatial distribution and geometric feature similarity. Local statistical feature parameters are calculated to identify and eliminate outliers.

Benefits of technology

This improves the accuracy and reliability of urban data collection, ensuring that the data more accurately reflects the actual situation of the city and provides reliable data support for the construction of smart cities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763486B_ABST
    Figure CN120763486B_ABST
Patent Text Reader

Abstract

This invention relates to the field of data processing technology, specifically to a method for urban data acquisition for the Internet of Things (IoT). The method includes: using a lidar system to comprehensively scan the city and acquire mapping data points in three-dimensional coordinates; for each mapping data point, determining preliminary neighboring points with a preset neighborhood radius; analyzing the similarity between the preliminary neighboring points and the mapping data point in spatial distribution and geometric features to obtain a comprehensive similarity between the preliminary neighboring points and the mapping data point; and performing a secondary screening of the preliminary neighboring points to determine final neighboring points. Based on the final neighboring points, calculating local statistical characteristic parameters for each mapping data point to identify and eliminate outliers, achieving statistical filtering and noise reduction. This effectively removes interference points, ensuring the effectiveness of statistical filtering and improving the accuracy and reliability of urban data acquisition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology. Specifically, it relates to a method for urban data collection for the Internet of Things (IoT). Background Technology

[0002] In today's digital age, the Internet of Things (IoT) technology, by building a vast network, tightly connects various devices and systems in a city, enabling real-time data collection, efficient transmission, and in-depth processing, thus becoming a core driving force for smart city construction. However, during the process of collecting urban mapping data points using lidar technology, the complexity of the urban environment, insufficient reflectivity, obstructions, and adverse weather conditions such as rain and snow can all interfere with the accuracy of data collection, resulting in noisy data points.

[0003] Statistical filtering is a method for filtering out outliers based on the local statistical characteristics of each data point to achieve noise reduction. It calculates statistical parameters such as the average distance, standard deviation, and variance between a data point and its neighbors to comprehensively identify and remove outliers, thereby optimizing data quality. When obtaining the neighbors of each data point, a uniform and fixed neighborhood radius is typically preset for all data points, centered on each data point. Based on these neighbors, the local statistical characteristic parameters of the data point are determined, and outliers are then filtered out to achieve noise reduction and improve data quality.

[0004] For example, the existing Chinese patent application document with publication number CN118447163A discloses a point cloud mapping method based on statistical filtering and KD tree indexing, which includes the following steps: acquiring the original point cloud image captured and sampled from the environment, and performing statistical filtering on the original point cloud image to obtain the filtered point cloud image; performing plane fitting on the filtered point cloud image to obtain the fitted plane point cloud; and using KD tree indexing to recover the RGB information of the fitted plane point cloud to obtain a color point cloud image containing RGB information.

[0005] However, in complex urban environments, the statistical filtering methods described above, which rely solely on neighborhood radii to select neighborhood points, have significant limitations. Cities contain abundant background noise and objects with uneven reflection, such as glass curtain walls and highly reflective building facades. These interference sources severely disrupt the neighborhood point selection process, resulting in a large number of irrelevant points mixed in with the selected neighborhood points. Due to the presence of these interference points, the local statistical feature parameters determined based on the neighborhood points become biased and inaccurate. Inaccurate local statistical feature parameters prevent statistical filtering from accurately removing outliers, affecting the filtering effect and ultimately reducing the accuracy and reliability of urban data collection. Summary of the Invention

[0006] To address the problem that existing statistical filtering methods for filtering and denoising urban data rely solely on neighborhood radii to select neighboring points, making them susceptible to interference and reducing the accuracy and reliability of urban data collection, this invention proposes an urban data collection method for the Internet of Things (IoT), comprising:

[0007] A comprehensive scan of the city is conducted using a laser mapping radar to acquire mapping data points within the city limits; these mapping data points are in three-dimensional coordinates.

[0008] For any acquired survey data point, a neighborhood radius is set with the survey data point as the center, and all survey data points within the neighborhood radius are determined as the initial neighborhood points of the survey data point.

[0009] Based on the similarity of each preliminary neighbor point to the mapping data point in terms of spatial distribution and geometric features, the comprehensive similarity between each preliminary neighbor point and the mapping data point is determined. Based on the comprehensive similarity, the preliminary neighbor points are further filtered to determine the final neighbor points of the mapping data point.

[0010] The local statistical feature parameters of each survey data point are calculated based on all the final neighboring points of each survey data point. Based on the local statistical feature parameters of each survey data point, it is determined whether the survey data point is an outlier and the outlier is removed to achieve statistical filtering.

[0011] All survey data points that have undergone statistical filtering are used as the final survey data points to complete the urban data collection.

[0012] The aforementioned technical solution utilizes comprehensive scanning of mapping data points by LiDAR to accurately depict the spatial structure of a city, providing a rich and intuitive spatial data foundation for smart city construction. Furthermore, it employs a method of setting a neighborhood radius centered on a single mapping data point to initially determine neighboring points. This provides a relatively reasonable initial data range for subsequent analysis of the local characteristics of each data point, a common method in statistical filtering for preliminary neighbor point selection, preparing for further filtering and calculation of local statistical characteristic parameters. Moreover, addressing the limitations of traditional methods that solely rely on neighborhood radius for neighbor point selection, it incorporates considerations of spatial distribution and geometric feature similarity. Through comprehensive similarity-based secondary filtering, it more accurately identifies neighboring points truly relevant to the current mapping data point, avoiding the inclusion of interfering points. This ensures that the final determined neighboring points for each mapping data point more accurately reflect its true local characteristics. Furthermore, by using the final neighborhood points determined through secondary screening to calculate local statistical feature parameters, the local statistical feature parameters become more reliable, enabling more accurate identification and removal of outliers, thus improving the filtering effect. This allows the processed data to more accurately reflect the actual situation of the city, improves the accuracy of urban data collection, and provides more reliable data for subsequent smart city construction applications.

[0013] Furthermore, the local statistical characteristic parameters of each mapping data point include: the mean distance between each mapping data point and all its final neighboring points.

[0014] Furthermore, one method for determining the final neighboring points of a mapping data point by performing a secondary screening of the initial neighboring points based on the comprehensive similarity is as follows:

[0015] Set a threshold for overall similarity;

[0016] If the normalized value of the comprehensive similarity between a preliminary neighbor point and the survey data point is less than the threshold among all preliminary neighbor points of each survey data point, the preliminary neighbor point is removed, and the remaining preliminary neighbor points are taken as the final neighbor points of the survey data point.

[0017] The above technical solution performs a secondary screening operation on the preliminary neighborhood points obtained by traditional methods based on a preset threshold, effectively eliminating a large number of interfering points and making the finally determined neighborhood points more representative and accurate.

[0018] Furthermore, another method for determining the final neighboring points of a mapping data point by performing a secondary screening based on the comprehensive similarity is as follows:

[0019] Calculate the mean of the overall similarity between all preliminary neighboring points and the mapping data point, and use this mean as the threshold for overall similarity;

[0020] If the normalized value of the comprehensive similarity between a preliminary neighbor point and the survey data point is less than the threshold among all preliminary neighbor points of each survey data point, the preliminary neighbor point is removed, and the remaining preliminary neighbor points are taken as the final neighbor points of the survey data point.

[0021] The above technical solution dynamically determines the threshold by calculating the average of the comprehensive similarity between all preliminary neighboring points and the central mapping data point, rather than using a fixed preset value. This method sets the screening criteria based on the characteristics of the data itself, and can adapt to the neighborhood situation of different mapping data points. Using this threshold, interference points can be eliminated more accurately, improving the correlation and consistency between the final neighboring points and the central mapping data point.

[0022] Furthermore, the method for determining whether a mapping data point is an outlier based on the local statistical characteristic parameters of each mapping data point is as follows:

[0023] For any given mapping data point, calculate the mean distance between that data point and all its final neighboring points. And the distance threshold between the mapping data point and all its final neighboring points:

[0024] ;in, This is the distance threshold between the mapping data point and all its final neighboring points. This is the mean distance between all survey data points and their respective final neighboring points. Let be the standard deviation of the distances between all survey data points and their respective final neighboring points. This is a preset multiplier factor;

[0025] like Greater than The mapping data point is an outlier. Not greater than The data point in the survey is not an outlier.

[0026] Furthermore, the overall similarity between each preliminary neighborhood point and the mapping data point is determined based on the following method:

[0027] The similarity in spatial distribution between each preliminary neighborhood point and the mapping data point is obtained and denoted as . And the similarity in geometric features between the preliminary neighborhood point and the mapping data point, denoted as ;

[0028] Will and The arithmetic square root of the sum of squares is used as the comprehensive similarity between the initial neighborhood point and the mapping data point.

[0029] The above technical solution integrates spatial distribution similarity and geometric feature similarity into a comprehensive index. This balances the weights of the two similarity indices to a certain extent, preventing either index from excessively influencing the comprehensive similarity result, thus more comprehensively and reasonably reflecting the overall similarity between preliminary neighborhood points and mapping data points.

[0030] Furthermore, the spatial similarity between each preliminary neighborhood point and the mapping data point is determined based on the following formula:

[0031] ;

[0032] In the formula, For the first The first mapping data point The similarity in spatial distribution between the initial neighboring points and the mapping data points. For the first The first mapping data point Local density of a preliminary neighborhood point For the first The local density of each mapping data point, and and The calculation method is consistent. For the first The first survey data and the first The distance between the initial neighboring points It is a natural exponential function.

[0033] The above technical solution comprehensively considers the local density differences and spatial distance relationships of data points, thereby more accurately measuring the similarity of spatial distribution. This comprehensive consideration helps to effectively distinguish the degree of spatial distribution between different data points in complex urban environmental data, and provides an accurate basis for subsequent screening of truly related neighborhood points in spatial distribution.

[0034] Furthermore, the similarity in geometric features between each preliminary neighborhood point and the mapping data point is determined based on the following formula:

[0035] In the formula, For the first The first mapping data point Similarity in geometric features between preliminary neighboring points and the mapping data point For the first The mapping data point and its first The angle between the initial neighboring points For the first Slope of each survey data point For the first The first mapping data point The slope of the initial neighborhood points, and and The calculation method is consistent. It is the absolute value symbol. This is used to prevent positive numbers with a denominator of 0.

[0036] The above technical solution, by comprehensively considering differences in angle and slope, quantifies the similarity of survey data points in terms of geometric shape. In the complex urban environmental data processing, it can more accurately identify data points with similar geometric features, providing a basis for selecting neighboring points that match the geometric features of the central survey data point.

[0037] Furthermore, the formula for calculating the angle is:

[0038] In the formula, For the first The mapping data point and its first The angle between the initial neighboring points It is an inverse cosine function. For the first Normal vector of each survey data point For the first The first mapping data point The normal vector of the initial neighboring points To calculate the modulus.

[0039] Furthermore, the formula for calculating the slope is:

[0040] In the formula, For the first Slope of each survey data point , , The first Each mapping data point is located at axis, axis, The value of the axis, , , They are respectively with the first The nearest mapping data point to each mapping data point is in axis, axis, The value of the axis.

[0041] The present invention has the following effects:

[0042] This invention performs secondary filtering of neighboring points by comprehensively considering the spatial distribution and geometric features of survey data points. This comprehensive and flexible filtering mechanism can adapt to complex and ever-changing urban environments and different data characteristics, effectively eliminate interference points, calculate reliable local statistical feature parameters, accurately identify and remove outliers, ensure the effectiveness of statistical filtering, and improve the accuracy and reliability of urban data collection. Attached Figure Description

[0043] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0044] Figure 2 This is a schematic diagram of the method flow for step S2 of the present invention. Detailed Implementation

[0045] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0046] Reference Figure 1 The present invention provides a method for urban data collection for the Internet of Things, comprising steps S1-S3:

[0047] S1: Initial collection of surveying data points within the city limits.

[0048] Urban surveying data points are the fundamental information units for building digital models of smart cities, and their accuracy and completeness directly affect the reliability of subsequent urban planning, management, and analysis. In the initial data collection phase, rigorous and scientific methods must be employed to acquire this crucial data.

[0049] The Internet of Things (IoT) technology has established an efficient information exchange and collaboration platform for urban surveying data collection. High-performance lidar equipment is strategically deployed in key urban areas and connected into a cohesive whole via the IoT. Based on the city's size, terrain complexity, and target mapping accuracy requirements, the remote control and data transmission capabilities of the IoT allow for precise setting of various operating parameters of the lidar, such as scanning angle range, scanning frequency, and laser emission power. IoT technology enables real-time feedback and adjustment of these parameters, ensuring comprehensive and detailed radar coverage of every corner of the city and the acquisition of massive amounts of surveying data points.

[0050] The collected mapping data points are all three-dimensional coordinates, which can accurately reflect the spatial position information of urban objects. The laser mapping radar is activated and performs a full-range scan of the city according to a preset scanning mode. The laser mapping radar emits high-frequency laser beams, and the emission angle of the laser beams is recorded. Using this data, the three-dimensional coordinate information of each reflection point can be obtained, and these three-dimensional coordinate information constitutes a series of mapping data points.

[0051] As the scanning continues, a large number of surveying data points are collected, which in turn construct a three-dimensional spatial model of the city. These initially collected surveying data points provide raw materials for subsequent data processing and analysis, and are the data foundation for realizing digital management of the city and the construction of a smart city.

[0052] S2: Preprocess the survey data points using improved statistical filtering.

[0053] After the initial collection of survey data points within the city, the data points often contain noise and interference caused by the complex urban environment. Preprocessing is required to improve data quality and provide reliable data support for the subsequent construction of smart cities.

[0054] The core of this step lies in performing a secondary screening of the initially set neighborhood points for each surveying data point using traditional statistical filtering methods. This improves the accuracy of the neighborhood points and thus enhances the filtering effect on the surveying data points. Specific steps are as follows: Figure 2 As shown:

[0055] S21: Determine multiple preliminary neighborhood points for each mapping data point.

[0056] For each acquired mapping data point, a neighborhood radius is set with that point as the center. (Empirical value) All mapping data points within the radius of this neighborhood are identified as the initial neighborhood points of this mapping data point. This is the same as the initial screening method of traditional statistical filtering, which defines the initial range for further processing and is a continuation of traditional methods. Through this initial screening process, local data sets that may be closely related to each mapping data point can be quickly extracted from a massive amount of mapping data points, excluding a large number of obviously irrelevant mapping data points, so that the subsequent secondary screening can focus more on the truly key neighborhood information.

[0057] S22: Analyze the overall similarity between each preliminary neighborhood point and the mapping data point.

[0058] This step involves in-depth analysis of the spatial distribution similarity between each preliminary neighbor point and the mapping data point, as well as their similarity in geometric features. This yields a comprehensive similarity between each preliminary neighbor point and the mapping data point, providing a foundation for the subsequent accurate selection of the final neighbor points of the mapping data point.

[0059] S221: Similarity analysis in spatial distribution.

[0060] Spatial similarity focuses on the spatial relationship between preliminary neighboring points and mapping data points. Local density and distance each play an irreplaceable role in judging spatial similarity. Local density reflects the degree of clustering of mapping data points within a certain area, providing a macroscopic perspective for judging spatial similarity; while the distance between mapping data points and their preliminary neighboring points precisely describes the specific spatial relationship between the preliminary neighboring points and the mapping data points, further refining the judgment of similarity at a micro level. Combining the two constructs a comprehensive and accurate analytical system for spatial distribution.

[0061] In different areas of the city, mapping data points exhibit drastically different distribution characteristics. For example, in commercial districts, the dense arrangement of high-rise buildings results in a close distribution of mapping data points, with a relatively high local density. In contrast, in open areas such as parks, the mapping data points corresponding to natural landscapes such as trees and lawns are more dispersed.

[0062] In one embodiment, the formula for calculating the local density of any mapping data point is:

[0063]

[0064] In this formula, For the first Local density of a mapping data point For the first The total number of preliminary neighborhood points for each mapping data point The initial neighboring point number, Take all All integers within the range, For the first The mapping data point and its first The distance between the initial neighboring points.

[0065] This formula accurately describes the spatial density of survey data points by comprehensively considering the number and distance of initial neighboring points. If a survey data point has a larger number of initial neighboring points and a smaller sum of distances between the survey data point and all its initial neighboring points, it indicates that there are more initial neighboring points distributed around the survey data point. Furthermore, the closer the survey data point is to its surrounding initial neighboring points, the greater the local density of the survey data point will be, and vice versa.

[0066] The local density of each mapping data point and its initial neighboring points is obtained according to the above formula (since each initial neighboring point is also a mapping data point), and then the following analysis is performed:

[0067] For the initial neighborhood points of a mapping data point:

[0068] The smaller the difference between the local density of the initial neighboring points and the local density of the surveyed data point, the more similar the clustering of the initial neighboring points and the surrounding data points is, indicating a higher probability that they belong to the same spatial distribution area and have a higher similarity in spatial distribution. The smaller the spatial distance between the initial neighboring points and the surveyed data point, the closer they are in spatial location and the higher their similarity in spatial distribution. Conversely, points that are far apart, such as the surveyed data points of streetlights in a square and buildings, have a larger distance, indicating a greater difference in their spatial distribution.

[0069] Therefore, based on the above analysis, the spatial similarity between each preliminary neighborhood point and the mapping data point is calculated.

[0070] In one embodiment, the spatial similarity between each initial neighborhood point and the mapping data point is determined based on the following formula:

[0071]

[0072] In this formula, For the first The first mapping data point The similarity in spatial distribution between the initial neighboring points and the mapping data points. For the first Local density of a mapping data point For the first The first mapping data point The local density of the initial neighborhood points, and and The calculation method is consistent (because the first... Each initial neighborhood point is essentially a mapping data point itself. For the first The first survey data and the first The distance between the initial neighboring points It is a natural exponential function.

[0073] The smaller the difference between the local density of the initial neighboring points and the local density of the mapping data points, and the smaller the distance, the more closely connected the initial neighboring points and the mapping data points are in space, and they are likely to belong to the same object or the same region, with a high degree of similarity in spatial distribution. The larger the difference between the local density of the initial neighboring points and the local density of the mapping data points, and the larger the distance, the more significant the difference in spatial distribution between the initial neighboring points and the mapping data points, with a lower degree of similarity in spatial distribution.

[0074] This comprehensive analysis can effectively determine the degree of similarity in spatial distribution between preliminary neighborhood points and survey data points, providing a reliable basis for subsequent selection of final neighborhood points.

[0075] S222: Similarity analysis based on geometric features.

[0076] Geometric feature similarity primarily focuses on the morphological characteristics between surveying data points. Since different terrains and building structures possess unique geometric features, analyzing these features can effectively distinguish different types of surveying data points, thus laying a solid foundation for accurate data processing and application. Geometric feature similarity can be specifically demonstrated through the analysis of slope differences and angular relationships.

[0077] In this analysis, the angle specifically refers to the angle between the normal vector of the initial neighboring point and the normal vector of the surveyed data point. As a vector perpendicular to the surface of an object, the direction of the normal vector contains rich spatial information. By calculating this angle, the directional differences between the initial neighboring point and the surveyed data point in space can be revealed intuitively and accurately. Taking urban buildings as an example, which contain walls with different orientations, when the initial neighboring point and the surveyed data point are located on these walls with different orientations, their normal vector directions will inevitably be different. By rigorously calculating the angle between the normal vectors, the degree of difference between them in spatial direction can be clearly determined. If the angle is 0°, it indicates that their normal vector directions are consistent, and they are highly similar in spatial direction; if the angle is 180°, it means that their normal vector directions are completely opposite, and their spatial direction differences are extremely large. This angle analysis provides an important basis for judging the similarity of geometric features.

[0078] In one embodiment, the formula for calculating the angle between each mapping data point and each of its initial neighboring points is:

[0079]

[0080] In this formula, For the first The mapping data point and its first The angle between the initial neighboring points It is an inverse cosine function. For the first Normal vector of each survey data point For the first The first mapping data point The normal vector of the initial neighboring points To calculate the modulus.

[0081] Slope describes the degree of inclination of a surface relative to the horizontal plane, and it is a key indicator reflecting the characteristics of terrain and building structures. In urban environments, various objects have unique slopes; roads, ramps, rooftops, etc., all possess unique slope values. By carefully comparing the slope differences between preliminary neighboring points and survey data points, we can gain a deeper understanding of their spatial inclination and relative height changes, which plays a decisive role in accurately determining the similarity of terrain and building structures. For example, on a continuous sloping road, because its overall slope is relatively consistent, the slope difference between preliminary neighboring points and survey data points located on the same road is usually small. Conversely, if one point is located on a sloping road and the other is located on a horizontal plaza, their slope difference will be very significant. By analyzing slope differences, we can quickly determine whether preliminary neighboring points and survey data points belong to the same terrain or building structure, thereby filtering out preliminary neighboring points with similar geometric characteristics to the survey data points.

[0082] In one embodiment, the formula for calculating the slope of any survey data point is:

[0083]

[0084] In this formula, For the first Slope of each survey data point , , The first Each mapping data point is located at axis, axis, The value of the axis, , , They are respectively with the first The nearest mapping data point to each mapping data point is in axis, axis, The value of the axis.

[0085] Therefore, by comprehensively analyzing the differences in slope and the relationship between angles, a comprehensive and accurate system for judging the similarity of geometric features can be constructed.

[0086] In one embodiment, the similarity in geometric features between each initial neighborhood point and the mapped data point is determined based on the following formula:

[0087]

[0088] In this formula, For the first The first mapping data point The similarity in geometric features between the initial neighboring points and the mapping data point; the larger the value, the stronger the similarity in geometric features. The first mapping data point The higher the similarity in geometric features between a preliminary neighborhood point and the mapping data point, the lower the similarity; conversely, the smaller the value, the lower the similarity. For the first The mapping data point and its first The angle between the initial neighboring points For the first Slope of each survey data point For the first The first mapping data point The slope of the initial neighborhood points, and and The calculation method is consistent (because the first...) Each initial neighborhood point is essentially a mapping data point. It is the absolute value symbol. To prevent positive numbers with a denominator of 0, the value is... .

[0089] In this formula, Part of the calculation comprehensively considers two key factors: angle and slope differences. By multiplying the angle and slope differences, the formula comprehensively reflects the combined differences between the preliminary neighboring points and the surveyed data points in two important geometric feature dimensions: spatial orientation and inclination. Smaller angle and slope differences indicate higher similarity in geometric features between the preliminary neighboring points and the surveyed data points; conversely, larger angle and slope differences mean lower similarity in geometric features. This comprehensive calculation method avoids the one-sided influence of a single factor on similarity judgment, making the quantification of geometric feature similarity more comprehensive and accurate.

[0090] S223: Obtain the comprehensive similarity between each preliminary neighborhood point and the mapping data point.

[0091] Calculate the spatial similarity between each preliminary neighbor point and the mapping data point, as well as the geometric similarity between each preliminary neighbor point and the mapping data point. The arithmetic square root of the sum of their squares is taken as the comprehensive similarity, expressed by the formula:

[0092]

[0093] In the formula, For the first The first mapping data point The overall similarity between the initial neighboring points and the mapping data point. For the first The first mapping data point The similarity in spatial distribution between the initial neighboring points and the mapping data points. For the first The first mapping data point The similarity in geometric features between the initial neighborhood points and the mapping data points.

[0094] This formula comprehensively considers the similarity between each preliminary neighborhood point and the survey data point in two key dimensions: spatial location and geometric shape. Spatial distribution similarity reflects the positional relationship of points in space, while geometric feature similarity reflects the morphological characteristics between points. Both are indispensable for accurately determining the correlation between neighborhood points and survey data points. By using the square root of the sum of squares calculation method, the excessive influence of single-dimensional similarity on the comprehensive result is avoided, ensuring the comprehensiveness and objectivity of the evaluation.

[0095] In urban surveying data processing, this formula allows for a more accurate measurement of the similarity between preliminary neighboring points and surveyed data points. For example, in a commercial area, when determining whether a preliminary neighboring point belongs to the same building as a surveyed data point, relying solely on spatial distribution similarity might misclassify points that are geographically close but have significantly different geometric features as related points; conversely, relying solely on geometric feature similarity might overlook points that are spatially distant but have similar geometric features. A comprehensive similarity formula, however, considers both factors, effectively improving the accuracy of the judgment and filtering out truly relevant neighboring points.

[0096] To further explain, the overall similarity was normalized using max-min normalization, which maps the overall similarity values ​​to the range of 0 to 1. This eliminates discrepancies in the range of overall similarity values ​​across different datasets, ensuring comparability of calculated overall similarities under various conditions. In urban mapping data, the original overall similarity values ​​for different regions and object types can vary significantly. Normalization standardizes these values ​​to a uniform scale, facilitating subsequent data comparison and analysis.

[0097] S23: Based on comprehensive similarity, a second screening is performed on the initial neighborhood points to obtain the final neighborhood points.

[0098] This section describes an improvement upon traditional methods, which lack a secondary screening mechanism. This new method, based on comprehensive similarity, employs a preset threshold or dynamically calculated average as the threshold to perform a secondary screening of initial neighboring points. Among all initial neighboring points for each mapping data point, those with a comprehensive similarity lower than the threshold are removed, and the remaining initial neighboring points are used as the final neighboring points for that mapping data point. This effectively eliminates interfering points, ensuring that the final neighboring points more accurately reflect the true local features of the central mapping data point.

[0099] Secondary filtering, through further analysis of these neighboring points, can eliminate those that do not meet expectations, thereby improving the noise identification accuracy of urban surveying data and thus enhancing the filtering effect.

[0100] In one embodiment, a method for determining the final neighboring points of a mapping data point by performing a secondary screening of the initial neighboring points based on the comprehensive similarity is as follows:

[0101] The preset threshold for overall similarity is 0.62 (empirical value).

[0102] If, among all the preliminary neighboring points of each mapping data point, the normalized value of the comprehensive similarity between a preliminary neighboring point and the mapping data point is less than 0.62, the preliminary neighboring point is removed, and the remaining preliminary neighboring points are taken as the final neighboring points of the mapping data point.

[0103] This preset fixed threshold provides a clear standard for the screening process, ensuring that the screening criteria remain consistent regardless of data changes. This guarantees the stability of the screening results when processing large amounts of data.

[0104] In one embodiment, another method for determining the final neighboring points of a mapping data point by performing a secondary screening of the initial neighboring points based on the comprehensive similarity is as follows:

[0105] Calculate the mean of the overall similarity between all preliminary neighboring points and the mapping data point, and use this mean as the threshold for overall similarity;

[0106] If the normalized value of the comprehensive similarity between a preliminary neighbor point and the survey data point is less than the threshold among all preliminary neighbor points of each survey data point, the preliminary neighbor point is removed, and the remaining preliminary neighbor points are taken as the final neighbor points of the survey data point.

[0107] Urban mapping data from different regions may exhibit varying distribution characteristics. Fixed thresholds cannot adapt to changes in data distribution and characteristics. If urban mapping data differs significantly due to regional variations, different measuring equipment, or environmental factors, fixed thresholds may fail to accurately select suitable neighboring points. Therefore, by calculating dynamic thresholds, a secondary screening can be performed based on the characteristics of the mapping data points themselves, ensuring that the selected neighboring points more accurately reflect the local features of the data point. For example, in areas where mapping data points are concentrated, the threshold will be relatively high, eliminating initial neighboring points that differ significantly from the central mapping data point; conversely, in areas where mapping data points are more dispersed, the threshold will be lowered accordingly to avoid excessive elimination of initial neighboring points.

[0108] In summary, this step accurately screened the initial neighborhood points of each survey data point, resulting in final neighborhood points that better reflect the local characteristics of the survey data point.

[0109] S24: Calculate local statistical feature parameters based on the final neighborhood points, and remove outliers to complete statistical filtering.

[0110] This section presents an improvement. Traditional methods suffer from inaccurate neighbor point selection, leading to unreliable calculations of local statistical characteristic parameters. For each mapping data point, its local statistical characteristic parameters include: the mean distance between the mapping data point and its final neighbor points, denoted as... .

[0111] Calculate the distance threshold between each mapping data point and its final neighboring points:

[0112]

[0113] In this formula, This is the distance threshold between the survey data point and its final neighboring points. This is the mean distance between all survey data points and their respective final neighboring points. Let be the standard deviation of the distances between all survey data points and their respective final neighboring points. This is a preset multiplier factor, and here it is set to 0.5 (an empirical value).

[0114] Determine whether each survey data point is an outlier: If Greater than The mapping data point is an outlier. Not greater than The data point in the survey is not an outlier.

[0115] This operation, based on accurately selected final neighboring points, makes the calculation of local statistical feature parameters more reliable. This is because the final neighboring points are highly similar to the survey data points in terms of spatial distribution and geometric features, which can truly reflect the actual characteristics of the survey data points in the local area, making outlier removal more accurate, optimizing the effect of statistical filtering and denoising, and completing the preprocessing of the survey data points.

[0116] S3: Use the preprocessed survey data points as the final survey data points to complete the urban data collection.

[0117] The preprocessed surveying data points, used as the final surveying data points, underwent rigorous screening and optimization to remove noise and outliers, accurately reflecting the city's actual geographic information. This completed the city data collection process, providing solid and reliable data support for subsequent urban planning, construction, management, and various smart city applications, thus contributing to the city's digital development and efficient operation.

Claims

1. A method for urban data collection for the Internet of Things, characterized in that, include: A comprehensive scan of the city is conducted using a laser mapping radar to acquire mapping data points within the city limits; these mapping data points are in three-dimensional coordinates. For any acquired survey data point, a neighborhood radius is set with the survey data point as the center, and all survey data points within the neighborhood radius are determined as the initial neighborhood points of the survey data point. Based on the spatial similarity between each preliminary neighbor point and the mapping data point, and the geometric similarity between each preliminary neighbor point and the mapping data point, the comprehensive similarity between each preliminary neighbor point and the mapping data point is determined. This includes: obtaining the spatial similarity between each preliminary neighbor point and the mapping data point, denoted as... And the similarity in geometric features between the preliminary neighborhood point and the mapping data point, denoted as ;Will and The arithmetic square root of the sum of squares is used as the comprehensive similarity between the preliminary neighborhood point and the mapping data point; The spatial similarity between each initial neighborhood point and the mapping data point is determined based on the following formula: In the formula, For the first The first mapping data point The similarity in spatial distribution between the initial neighboring points and the mapping data points. For the first The first mapping data point Local density of a preliminary neighborhood point For the first The local density of each mapping data point, and and The calculation method is consistent. For the first The first survey data and the first The distance between the initial neighboring points It is a natural exponential function; The similarity in geometric features between each initial neighborhood point and the mapped data point is determined based on the following formula: In the formula, For the first The first mapping data point Similarity in geometric features between preliminary neighboring points and the mapping data point For the first The mapping data point and its first The angle between the initial neighboring points For the first Slope of each survey data point For the first The first mapping data point The slope of the initial neighborhood points, and and The calculation method is consistent. It is the absolute value symbol. This is to prevent positive numbers with a denominator of 0; Based on the comprehensive similarity, the initial neighboring points are further filtered to determine the final neighboring points of the mapping data point; The local statistical feature parameters of each survey data point are calculated based on all the final neighboring points of each survey data point. Based on the local statistical feature parameters of each survey data point, it is determined whether the survey data point is an outlier and the outlier is removed to achieve statistical filtering. All survey data points that have undergone statistical filtering are used as the final survey data points to complete the urban data collection.

2. The urban data collection method for the Internet of Things according to claim 1, characterized in that, The local statistical characteristic parameters of each mapping data point include: the mean distance between each mapping data point and all its final neighboring points.

3. The urban data collection method for the Internet of Things according to claim 1, characterized in that, One method for determining the final neighboring points of a mapping data point by performing a secondary screening of the initial neighboring points based on the comprehensive similarity is as follows: Set a threshold for overall similarity; If the normalized value of the comprehensive similarity between a preliminary neighbor point and the survey data point is less than the threshold among all preliminary neighbor points of each survey data point, the preliminary neighbor point is removed, and the remaining preliminary neighbor points are taken as the final neighbor points of the survey data point.

4. The urban data collection method for the Internet of Things according to claim 1, characterized in that, Another method for determining the final neighboring points of a mapping data point by performing a secondary screening of the initial neighboring points based on the comprehensive similarity is as follows: Calculate the mean of the overall similarity between all preliminary neighboring points and the mapping data point, and use this mean as the threshold for overall similarity; If the normalized value of the comprehensive similarity between a preliminary neighbor point and the survey data point is less than the threshold among all preliminary neighbor points of each survey data point, the preliminary neighbor point is removed, and the remaining preliminary neighbor points are taken as the final neighbor points of the survey data point.

5. The urban data collection method for the Internet of Things according to claim 2, characterized in that, The method for determining whether a mapping data point is an outlier based on the local statistical characteristic parameters of each mapping data point is as follows: For any given mapping data point, calculate the mean distance between that data point and all its final neighboring points. And the distance threshold between the mapping data point and all its final neighboring points: ;in, This is the distance threshold between the mapping data point and all its final neighboring points. This is the mean distance between all survey data points and their respective final neighboring points. Let be the standard deviation of the distances between all survey data points and their respective final neighboring points. This is a preset multiplier factor; like Greater than The mapping data point is an outlier. Not greater than The data point in the survey is not an outlier.

6. The urban data collection method for the Internet of Things according to claim 3, characterized in that, The formula for calculating the angle is: In the formula, For the first The mapping data point and its first The angle between the initial neighboring points It is an inverse cosine function. For the first Normal vector of a survey data point For the first The first mapping data point The normal vector of the initial neighboring points To calculate the modulus.

7. The urban data collection method for the Internet of Things according to claim 6, characterized in that, The formula for calculating slope is: In the formula, For the first Slope of each survey data point , , The first Each mapping data point is located at axis, axis, The value of the axis, , , They are respectively with the first The nearest mapping data point to each mapping data point is in axis, axis, The value of the axis.

Citation Information

Patent Citations

  • Point cloud mapping method based on statistical filtering and KD tree index

    CN118447163A

  • Power transmission line three-dimensional reconstruction method based on unmanned aerial vehicle and laser radar

    CN119810311A

  • Multi-source laser point cloud data adaptive registration method and system for power grid equipment

    CN120014006A