Method and system for monitoring abnormality of geographic surveying and mapping data based on the Internet of Things
Through the method of kernel functions and density attracting points, the problem of abnormal detection in geographic surveying and mapping data is solved, efficient and accurate abnormal monitoring is achieved, and data quality is improved.
Patent Information
- Application Number
- CN202510796178.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-06-16
AI Technical Summary
The existing geographic surveying and mapping data anomaly detection methods are difficult to effectively identify potential anomaly patterns in high-dimensional data, and the traditional methods are not effective in the face of environmental noise and electromagnetic interference, which affects the accuracy of subsequent analysis.
By obtaining the geographic surveying and mapping data collected by IoT devices, using the kernel function to calculate the overall density function, combining the data gradient and local density curvature to determine the density attraction point, and filtering based on the fuzzy membership and the overlap of the influence domain to calculate the degree of abnormality of the data points.
It improves the accuracy and accuracy of abnormal monitoring of geographic surveying and mapping data, reduces complexity, and ensures data quality.
Smart Images

Figure CN120316689B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of geographic surveying and mapping, and in particular to a method and system for monitoring geographic surveying and mapping data anomalies based on the Internet of Things. Background Art
[0002] Geographic surveying and mapping plays a vital role in smart cities, environmental monitoring, disaster warning, and other fields. With the widespread adoption of the Internet of Things (IoT), a large number of sensor devices are being deployed in various geographical environments, collecting real-time data such as topography, meteorology, and geology. However, due to environmental noise and electromagnetic interference, geographic surveying and mapping data often contain noise. If these anomalies are not addressed, they will impact subsequent geospatial analysis and applications. For example, when constructing digital elevation models, anomalous elevation points can cause spikes or dips on the terrain surface, misleading flood simulations, hydrological analysis, or earthwork calculations. In remote sensing image classification, anomalous pixel values can be misidentified as specific landform types, reducing classification accuracy. While many methods exist for detecting anomalies, such as statistical methods, k-means clustering, and local anomaly factors, these methods struggle to effectively identify underlying anomalous patterns. For example, statistical methods with fixed thresholds cannot adapt to the dynamics of data streams; clustering-based algorithms are sensitive to initial parameters and struggle to handle local anomalies in high-dimensional data; and distance-based methods tend to fail in densely populated areas. The DENCLUE algorithm has certain advantages in anomaly detection due to its density-based nature, but it also has significant limitations. For example, fixed kernel function parameters are difficult to adapt to the heterogeneity of geographic data, and hill climbing algorithms are prone to falling into local optimality. Therefore, there is an urgent need for an anomaly detection method that can be applied to geographic surveying and mapping data. Summary of the Invention
[0003] To address the above issues, this application proposes a method for monitoring geographic surveying and mapping data anomalies based on the Internet of Things, including:
[0004] Obtaining geographic surveying and mapping data collected by an IoT device, obtaining a kernel function based on local density and time attributes of the geographic surveying and mapping data, and calculating an overall density function of the geographic surveying and mapping data using the kernel function;
[0005] An initial point is determined based on the prior characteristics of the geographic surveying and mapping data. Starting from the initial point, at each movement, a moving step is obtained based on the data gradient and local density curvature in the neighborhood of the current search position. The moving step is used to search for the local maximum of the overall density function and determine the density attraction point; the influence domain overlap of the density attraction point is calculated, and the density attraction point is filtered based on the influence domain overlap;
[0006] The fuzzy membership of each geographic surveying and mapping data point to the density attraction point is calculated, and the abnormality degree of the data point is calculated based on the fuzzy membership of the data point and the distance to each density attraction point. The geographic surveying and mapping data point whose abnormality degree exceeds the threshold is determined as abnormal data.
[0007] Optionally, obtaining a kernel function according to the local density and time attributes of the geographic surveying and mapping data includes:
[0008] Calculating the local density of the geographic surveying and mapping data point within a preset neighborhood, if the data point is located in a data point dense area, using a first bandwidth value as the bandwidth of the Gaussian kernel function; otherwise, using a second bandwidth value as the bandwidth of the Gaussian kernel function, where the first bandwidth value is smaller than the second bandwidth value;
[0009] The time distance between the acquisition time of the data point and the current time is calculated, a time decay factor is obtained according to the time distance, and the time decay factor is used as the weight of the Gaussian kernel function.
[0010] Optionally, the calculating the overall density function of the geographic surveying and mapping data using the kernel function includes:
[0011] For each target data point in the geographic surveying and mapping data set, the remaining data points are substituted into the kernel function centered on the target data point to obtain the kernel function contribution value of each remaining data point to the target data point;
[0012] Summing all kernel function contribution values for the target data point to obtain the overall density function value at the target data point;
[0013] Traverse all geographic mapping data points to obtain the overall density function of the entire geographic mapping data set.
[0014] Optionally, obtaining the moving step size according to the data gradient and local density curvature in the neighborhood of the current search position includes:
[0015] Calculate the gradient vector of the current search position on the overall density function, where the direction of the gradient vector points to the direction in which the density increases fastest;
[0016] Calculating the Hessian matrix of the overall density function in the neighborhood of the current search position, taking the eigenvalue of the Hessian matrix with the largest absolute value as the local density curvature, and inputting the absolute value of the local density curvature into a monotonically decreasing function to obtain a step size adjustment factor;
[0017] The modulus of the gradient vector is multiplied by the basic step size, and the product of the multiplication result and the step size adjustment factor is used as the moving step size, and the direction of the moving step size is consistent with the direction of the gradient vector.
[0018] Optionally, the adopting a moving step to search for a local maximum of the overall density function and determining a density attraction point includes:
[0019] Starting from the initial point, the current search position point is moved a corresponding distance in the direction determined by the current moving step to reach a new search position point; the moving step of the new search position point is recalculated based on the data gradient and local density curvature in the neighborhood of the current search position, and the process is repeated until: the moving step is lower than the convergence threshold, or the change in the search position point in multiple consecutive iterations is less than the minimum displacement value, then the current search position point is taken as a local maximum point, and the local maximum point is determined as a density attraction point.
[0020] Optionally, the calculating the influence domain overlap of the density attraction points and filtering the density attraction points based on the influence domain overlap includes:
[0021] The circular area with the density attraction point as the center and a radius equal to the attraction point density value multiplied by the influence domain coefficient is used as the influence domain of the density attraction point;
[0022] The area intersection between the influence domains of any two density attraction points is calculated, and the ratio of the area intersection to the area of the smallest of the two influence domains is taken as the influence domain overlap. When the influence domain overlap between any two density attraction points exceeds the overlap threshold, the density attraction point with the largest density value is retained, and the density attraction point with the smallest density value is removed.
[0023] Optionally, the calculating of the fuzzy membership of each geographic surveying and mapping data point to the density attraction point includes:
[0024] For each geographic mapping data point, calculate the Euclidean distance between it and each determined density attractor;
[0025] Based on the Euclidean distance, the following formula is used to calculate the fuzzy membership of the data point i to the kth density attraction point: :
[0026]
[0027] in is the distance between data point i and density attraction point k, C is the total number of density attraction points, m is the fuzzy factor and m>1.
[0028] Optionally, calculating the abnormality degree of the data point based on the fuzzy membership of the data point and the distance from each density attraction point includes:
[0029] Calculating a membership weighted distance of the geographic surveying and mapping data point to each of the density attraction points, wherein the membership weighted distance is the fuzzy membership degree of the data point to the density attraction point multiplied by the Euclidean distance from the data point to the density attraction point;
[0030] The membership weighted distances of the data point to all the density attraction points are summed, and the sum is used as the abnormality degree of the data point.
[0031] This application also proposes a geographic surveying and mapping data anomaly monitoring system based on the Internet of Things, including:
[0032] a density estimation unit, configured to obtain geographic surveying and mapping data collected by an IoT device, obtain a kernel function based on the local density and time attributes of the geographic surveying and mapping data, and calculate an overall density function of the geographic surveying and mapping data using the kernel function;
[0033] A search and filtering unit is configured to determine an initial point based on the prior characteristics of the geographic surveying and mapping data, and to determine a moving step size at each movement based on the data gradient and local density curvature in the neighborhood of the current search position, and to search for a local maximum of the overall density function using the moving step size, and to determine a density attraction point; calculate an influence domain overlap of the density attraction point, and to filter the density attraction point based on the influence domain overlap;
[0034] The anomaly detection unit is used to calculate the fuzzy membership of each geographic surveying and mapping data point to the density attraction point, calculate the abnormality of the data point based on the fuzzy membership of the data point and the distance to each density attraction point, and determine the geographic surveying and mapping data point whose abnormality exceeds a threshold as abnormal data.
[0035] Optionally, obtaining a kernel function according to the local density and time attributes of the geographic surveying and mapping data includes:
[0036] Calculating the local density of the geographic surveying and mapping data point within a preset neighborhood, if the data point is located in a data point dense area, using a first bandwidth value as the bandwidth of the Gaussian kernel function; otherwise, using a second bandwidth value as the bandwidth of the Gaussian kernel function, where the first bandwidth value is smaller than the second bandwidth value;
[0037] The time distance between the acquisition time of the data point and the current time is calculated, a time decay factor is obtained according to the time distance, and the time decay factor is used as the weight of the Gaussian kernel function.
[0038] Optionally, the calculating the overall density function of the geographic surveying and mapping data using the kernel function includes:
[0039] For each target data point in the geographic surveying and mapping data set, the remaining data points are substituted into the kernel function centered on the target data point to obtain the kernel function contribution value of each remaining data point to the target data point;
[0040] Summing all kernel function contribution values for the target data point to obtain the overall density function value at the target data point;
[0041] Traverse all geographic mapping data points to obtain the overall density function of the entire geographic mapping data set.
[0042] Optionally, obtaining the moving step size according to the data gradient and local density curvature in the neighborhood of the current search position includes:
[0043] Calculate the gradient vector of the current search position on the overall density function, where the direction of the gradient vector points to the direction in which the density increases fastest;
[0044] Calculating the Hessian matrix of the overall density function in the neighborhood of the current search position, taking the eigenvalue of the Hessian matrix with the largest absolute value as the local density curvature, and inputting the absolute value of the local density curvature into a monotonically decreasing function to obtain a step size adjustment factor;
[0045] The modulus of the gradient vector is multiplied by the basic step size, and the product of the multiplication result and the step size adjustment factor is used as the moving step size, and the direction of the moving step size is consistent with the direction of the gradient vector.
[0046] Optionally, the adopting a moving step to search for a local maximum of the overall density function and determining a density attraction point includes:
[0047] Starting from the initial point, the current search position point is moved a corresponding distance in the direction determined by the current moving step to reach a new search position point; the moving step of the new search position point is recalculated based on the data gradient and local density curvature in the neighborhood of the current search position, and the process is repeated until: the moving step is lower than the convergence threshold, or the change in the search position point in multiple consecutive iterations is less than the minimum displacement value, then the current search position point is taken as a local maximum point, and the local maximum point is determined as a density attraction point.
[0048] Optionally, the calculating the influence domain overlap of the density attraction points and filtering the density attraction points based on the influence domain overlap includes:
[0049] The circular area with the density attraction point as the center and a radius equal to the attraction point density value multiplied by the influence domain coefficient is used as the influence domain of the density attraction point;
[0050] The area intersection between the influence domains of any two density attraction points is calculated, and the ratio of the area intersection to the area of the smallest of the two influence domains is taken as the influence domain overlap. When the influence domain overlap between any two density attraction points exceeds the overlap threshold, the density attraction point with the largest density value is retained, and the density attraction point with the smallest density value is removed.
[0051] Optionally, the calculating of the fuzzy membership of each geographic surveying and mapping data point to the density attraction point includes:
[0052] For each geographic mapping data point, calculate the Euclidean distance between it and each determined density attractor;
[0053] Based on the Euclidean distance, the following formula is used to calculate the fuzzy membership of the data point i to the kth density attraction point: :
[0054]
[0055] in is the distance between data point i and density attraction point k, C is the total number of density attraction points, m is the fuzzy factor and m>1.
[0056] Optionally, calculating the abnormality degree of the data point based on the fuzzy membership of the data point and the distance from each density attraction point includes:
[0057] Calculating a membership weighted distance of the geographic surveying and mapping data point to each of the density attraction points, wherein the membership weighted distance is the fuzzy membership degree of the data point to the density attraction point multiplied by the Euclidean distance from the data point to the density attraction point;
[0058] The membership weighted distances of the data point to all the density attraction points are summed, and the sum is used as the abnormality degree of the data point.
[0059] This application obtains a kernel function through the local density and time attributes of the geographic surveying and mapping data, thereby improving the accuracy of density estimation; obtains a moving step size based on the data gradient and local density curvature in the neighborhood of the current search position to avoid falling into the local optimum; filters density attraction points based on the overlap of the influence domain, thereby reducing complexity; calculates the degree of abnormality of the data point based on the fuzzy membership of the data point and the distance to each density attraction point, thereby improving the accuracy of abnormality monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 This is a flow chart of Example 1;
[0061] Figure 2 Schematic diagram of the bandwidth of data points with different local densities;
[0062] Figure 3 is the heat map of the overall density function;
[0063] Figure 4 Schematic diagram of the fuzzy membership of a single data point. DETAILED DESCRIPTION
[0064] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0065] In a specific embodiment, the present application proposes a method for monitoring geographic surveying and mapping data anomalies based on the Internet of Things, such as Figure 1 Shown, including:
[0066] Step 1: Obtain geographic surveying and mapping data collected by an IoT device, obtain a kernel function based on the local density and time attributes of the geographic surveying and mapping data, and use the kernel function to calculate the overall density function of the geographic surveying and mapping data;
[0067] Geographical mapping data is collected through IoT devices deployed in various locations, including but not limited to environmental sensors, remote sensing devices, or mobile mapping terminals. The collected data includes but is not limited to latitude and longitude coordinates, timestamps, and various environmental indicators such as temperature, humidity, or terrain features. A kernel function is derived based on the local density and temporal attributes of the geographic mapping data. Specifically, in data-dense areas, the bandwidth of the kernel function is automatically reduced; in sparse areas, the bandwidth is expanded to ensure smoothness, such as Figure 2 As shown in Figure 2, a time decay factor is introduced to give more weight to recent data and gradually reduce the contribution of old data. Preferably, a sliding window statistics or local density estimation method is used to dynamically adjust the kernel parameters.
[0068] Step 2: determining an initial point based on the prior characteristics of the geographic surveying and mapping data; starting from the initial point, at each movement, obtaining a moving step length based on the data gradient and local density curvature in the neighborhood of the current search position; using the moving step length to search for the local maximum of the overall density function and determine the density attraction point; calculating the influence domain overlap of the density attraction point, and filtering the density attraction point based on the influence domain overlap;
[0069] When selecting the initial point from the geographic surveying and mapping data in combination with prior knowledge or data features, preferably, by analyzing the gradient changes of historical abnormal data or current data, high-density areas are preferentially selected as the initial points. More specifically, the high-density areas of the geographic surveying and mapping data are determined, and for each high-density area, the geometric center point of the area is used as the initial point, or each high-density area is clustered, preferably using k-means clustering, and the centroid of each cluster or the point closest to the centroid is used as the initial point. Prior features include but are not limited to administrative area divisions, land properties, data types, data stability, etc. Those skilled in the art should know that determining the initial point based on prior features is not limited to the above-mentioned method. In an alternative embodiment, the local density of the geographic surveying and mapping data points is calculated, and the k points with the largest local density are used as the initial points, or the points with a local density greater than the local density threshold are used as the initial points.
[0070] A hill climbing algorithm with a variable step size is used to search for density attractors, also known as density attractors. The step size is determined by the data gradient and / or curvature in the neighborhood of the current search location. In areas of gently varying density, larger step sizes are used for rapid advancement; in areas of drastic density fluctuations, smaller step sizes are used to improve accuracy. After each move, neighborhood features are recalculated and the step size adjusted. After the search is complete, redundant and low-quality attractors are filtered out by calculating the overlap of the density attractor's influence domain, ensuring that each attractor represents an independent cluster structure.
[0071] Step 3: Calculate the fuzzy membership of each geographic surveying and mapping data point to the density attraction point, calculate the abnormality of the data point based on the fuzzy membership of the data point and the distance to each density attraction point, and determine the geographic surveying and mapping data point whose abnormality exceeds the threshold as abnormal data.
[0072] The fuzzy membership of each data point to multiple density attractors is calculated. The membership reflects the strength of the association between the data point and the cluster represented by each attractor, preferably achieved through a weighted kernel function value or distance. For example, a data point may have a high membership to two adjacent attractors, indicating that it is located at the boundary of the cluster. The degree of anomaly is calculated based on the fuzzy membership distribution and the distance to the attractor. For example, if a data point has a very scattered membership distribution or is far away from all attractors, its anomaly is high.
[0073] In a specific embodiment, obtaining a kernel function based on the local density and time attributes of the geographic surveying and mapping data includes:
[0074] Calculating the local density of the geographic surveying and mapping data point within a preset neighborhood, if the data point is located in a data point dense area, using a first bandwidth value as the bandwidth of the Gaussian kernel function; otherwise, using a second bandwidth value as the bandwidth of the Gaussian kernel function, where the first bandwidth value is smaller than the second bandwidth value;
[0075] The time distance between the acquisition time of the data point and the current time is calculated, a time decay factor is obtained according to the time distance, and the time decay factor is used as the weight of the Gaussian kernel function.
[0076] Specifically, the local density distribution of geographic mapping data points is calculated. For example, in urban centers, where sensors are densely packed and data points are concentrated, a smaller bandwidth is used as the parameter for the Gaussian kernel function to capture local details. In data-sparse areas such as suburban or rural areas, a larger bandwidth is used. This avoids overfitting or underfitting caused by a fixed bandwidth, making the kernel function more closely aligned with the actual data distribution. Because time affects data, the kernel function is further adjusted based on the temporal properties of the data. For example, suppose a meteorological sensor collects temperature data 30 days ago. Based on the preset time decay function, its time decay factor is calculated to be 0.2, while the decay factor for the most recently collected data is 0.9. This results in a greater contribution of recent data to the density estimate, while the influence of historical data gradually decreases.
[0077] In a specific embodiment, the calculating the overall density function of the geographic surveying and mapping data using the kernel function includes:
[0078] For each target data point in the geographic surveying and mapping data set, the remaining data points are substituted into the kernel function centered on the target data point to obtain the kernel function contribution value of each remaining data point to the target data point;
[0079] Summing all kernel function contribution values for the target data point to obtain the overall density function value at the target data point;
[0080] Traverse all geographic mapping data points to obtain the overall density function of the entire geographic mapping data set.
[0081] Specifically, assume that the geographic mapping dataset contains three data points A, B, and C, each of which represents a geographic location and its measurement value. Take point A as the target data point, substitute point B and point C into the Gaussian kernel function centered on point A, and calculate the kernel function contribution values of point B and point C to point A. Point B is closer to point A and has a higher contribution value; point C is farther away from point A and has a lower contribution value. Add these two contribution values to obtain the overall density function value at point A, which represents the local data density around point A. Similarly, repeat the above process for points B and C, and finally obtain the overall density function of the entire dataset, which represents the spatial distribution characteristics of all data points. The calculation of the overall density function is to draw a smooth influence range for each data point and superimpose all influences. For example, in urban areas, due to the dense data points, the overall density function value after superposition is higher; while in suburban areas, the data points are sparse and the density value after superposition is lower. Figure 3 It is the heat map of the overall density function.
[0082] In a specific embodiment, obtaining the moving step size according to the data gradient and local density curvature in the neighborhood of the current search position includes:
[0083] Calculate the gradient vector of the current search position on the overall density function, where the direction of the gradient vector points to the direction in which the density increases fastest;
[0084] Calculating the Hessian matrix of the overall density function in the neighborhood of the current search position, taking the eigenvalue of the Hessian matrix with the largest absolute value as the local density curvature, and inputting the absolute value of the local density curvature into a monotonically decreasing function to obtain a step size adjustment factor;
[0085] The modulus of the gradient vector is multiplied by the basic step size, and the product of the multiplication result and the step size adjustment factor is used as the moving step size, and the direction of the moving step size is consistent with the direction of the gradient vector.
[0086] First, calculate the gradient vector of the current search position on the overall density function, and its direction indicates the path with the fastest density growth. For example, if the current search position is in a valley area, the gradient vector will point to the direction of the nearby ridge, thereby guiding the movement to the area with higher density, avoiding the inefficiency of random search. By calculating the Hessian matrix of the overall density function in the neighborhood of the current search position, the local density curvature information is extracted. If the absolute value of the maximum eigenvalue of the Hessian matrix is large, it indicates that the density in the area changes dramatically. At this time, the step size adjustment factor will be reduced to avoid the algorithm skipping the real density attraction point due to excessive step size. On the contrary, in areas with gentle density changes, the step size adjustment factor will be increased to accelerate convergence. The moving step size is determined by the modulus of the gradient vector, the basic step size, and the step size adjustment factor to ensure that the search process is both efficient and accurate. The direction of the gradient vector of the current search position on the overall density function points to the path with the fastest density growth, and the modulus reflects the intensity of the change. For example, the gradient vector The modulus value is . At the same time, calculate the maximum absolute eigenvalue of the Hessian matrix H Characterize the local curvature and input it into a monotonically decreasing function Get the adjustment factor The final moving step length s is determined by the basic step length , gradient modulus and adjustment factors Joint decision, the formula is , the direction is consistent with the gradient. For example, in a flat area, Smaller Close to 1, the step size is mainly determined by the gradient modulus, which enables fast movement; in areas with large curvature, Larger lead As it approaches 0, the step size automatically decreases to avoid oscillation.
[0087] In a specific embodiment, the step of searching for the local maximum of the overall density function using a moving step and determining a density attraction point includes:
[0088] Starting from the initial point, the current search position point is moved a corresponding distance in the direction determined by the current moving step to reach a new search position point; the moving step of the new search position point is recalculated based on the data gradient and local density curvature in the neighborhood of the current search position, and the process is repeated until: the moving step is lower than the convergence threshold, or the change in the search position point in multiple consecutive iterations is less than the minimum displacement value, then the current search position point is taken as a local maximum point, and the local maximum point is determined as a density attraction point.
[0089] Specifically, when searching for the local maximum of the overall density function, the algorithm starts from the initial point and moves in the direction determined by the current moving step. For example, if the current search position is P, the moving step is s, and the direction is the gradient vector direction, then the new search position is It can be expressed as . Each time it moves to a new position, the gradient and local density curvature of the point are recalculated to obtain the next moving step. This process is repeated until the moving step is less than the preset convergence threshold, or the position change is extremely small in multiple consecutive iterations, indicating that it is close to the local maximum. At this time, the current search position point is the density attraction point. For example, in geographic surveying and mapping data, if the initial point is located in a valley area, it will gradually move in the direction of density growth, and the step size will be dynamically adjusted with the terrain curvature. The step size is larger in flat areas and quickly approaches high-density areas; the step size is automatically reduced in steep areas to ensure convergence to the top of the mountain. The top of the mountain is the density attraction point, representing the data aggregation center of the area.
[0090] In a specific embodiment, the calculating the influence domain overlap of the density attraction points and filtering the density attraction points based on the influence domain overlap includes:
[0091] The circular area with the density attraction point as the center and a radius equal to the attraction point density value multiplied by the influence domain coefficient is used as the influence domain of the density attraction point;
[0092] The area intersection between the influence domains of any two density attraction points is calculated, and the ratio of the area intersection to the area of the smallest of the two influence domains is taken as the influence domain overlap. When the influence domain overlap between any two density attraction points exceeds the overlap threshold, the density attraction point with the largest density value is retained, and the density attraction point with the smallest density value is removed.
[0093] To avoid excessively similar or redundant attractors, density attractors are further filtered. The circular area with a radius equal to the density value multiplied by the influence coefficient, centered at the attractor, is used as the influence area of the attractor. For example, if a density attractor has a density of 50 and an influence coefficient of 0.1, its influence radius is 5. The degree of overlap between the influence areas of any two density attractors is calculated. Specifically, the area of the intersection of the two circular influence areas is calculated, then this intersection area is divided by the area of the smaller of the two influence areas to obtain a ratio, which is the influence area overlap. If the influence area overlap of a pair of density attractors exceeds a pre-defined threshold, such as 0.6, the two attractors may represent the same or very similar data clusters. Since the density attractor with the larger density value is more likely to represent the center of the data cluster, it is retained, while the smaller density attractor is removed. For example, in urban hotspot analysis, density attractors may be identified in two adjacent blocks. If their influence areas highly overlap, the one with the larger pedestrian flow is retained as the primary hotspot.
[0094] In a specific embodiment, the calculating of the fuzzy membership of each geographic surveying and mapping data point to the density attraction point includes:
[0095] For each geographic mapping data point, calculate the Euclidean distance between it and each determined density attractor;
[0096] Based on the Euclidean distance, the following formula is used to calculate the fuzzy membership of the data point i to the kth density attraction point: :
[0097]
[0098] in is the distance between data point i and density attraction point k, C is the total number of density attraction points, m is the fuzzy factor and m>1.
[0099] When calculating the fuzzy membership of a data point and a density attraction point, the Euclidean distance between the data point and each density attraction point is calculated. For example, in an environmental monitoring application, the distances between an air quality monitoring point and the centers of three pollution sources are 3km, 5km, and 7km respectively. These distance values will be substituted into the fuzzy membership formula. In the formula, the square inverse of the distance It reflects the proximity between the data point and the attraction point. The closer the distance, the larger the value. The fuzzy factor m controls the fuzziness of the membership. When m=2, the membership is inversely proportional to the square of the distance. Suppose there are three density attraction points, and the distances between a data point and them are 2, 4, and 6 respectively. When m=2, the membership of the data point to the first attraction point is calculated to be 0.72, indicating that it mainly belongs to the cluster represented by the first attraction point. Figure 4 Schematic diagram of fuzzy membership of a single data point.
[0100] In a specific embodiment, the calculating the abnormality degree of the data point based on the fuzzy membership of the data point and the distance from each density attraction point includes:
[0101] Calculating a membership weighted distance of the geographic surveying and mapping data point to each of the density attraction points, wherein the membership weighted distance is the fuzzy membership degree of the data point to the density attraction point multiplied by the Euclidean distance from the data point to the density attraction point;
[0102] The membership weighted distances of the data point to all the density attraction points are summed, and the sum is used as the abnormality degree of the data point.
[0103] In the process of calculating the degree of anomaly, the relationship between each data point and each density attraction point is determined by fuzzy membership and Euclidean distance. For example, if a data point P has a fuzzy membership of 0.8 to density attraction point A and a Euclidean distance of 5 to A, its weighted membership distance is 4; if P has a membership of 0.2 to attraction point B and a distance of 10 units, the weighted distance is 2 units. Summing the weighted distances of all attraction points, the degree of anomaly of P is 6. In farmland monitoring applications, normal crop data points have high membership and short distances to the nearest attraction point, while anomalies have low membership or long distances. For example, data points in pest and disease-infested areas may have low membership to all attraction points, and the sum of weighted distances is significantly higher than the threshold, thus being accurately identified as anomalies.
[0104] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the above embodiments, it should be understood by those skilled in the art that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. In addition, the various different implementations of the embodiments of the present invention can also be arbitrarily combined, as long as they do not violate the ideas of the embodiments of the present invention, and they should also be regarded as the contents disclosed in the embodiments of the present invention.
Claims
1. A method for monitoring geographic surveying and mapping data anomalies based on the Internet of Things, characterized in that: include: Obtaining geographic surveying and mapping data collected by an IoT device, obtaining a kernel function based on local density and time attributes of the geographic surveying and mapping data, and calculating an overall density function of the geographic surveying and mapping data using the kernel function; An initial point is determined based on the prior characteristics of the geographic surveying and mapping data. Starting from the initial point, at each movement, a moving step is obtained based on the data gradient and local density curvature in the neighborhood of the current search position. The moving step is used to search for a local maximum of the overall density function and determine a density attraction point. Calculate the influence domain overlap of the density attraction points, and filter the density attraction points based on the influence domain overlap; The fuzzy membership of each geographic surveying and mapping data point to the density attraction point is calculated, and the abnormality degree of the data point is calculated based on the fuzzy membership of the data point and the distance to each density attraction point. The geographic surveying and mapping data point whose abnormality degree exceeds the threshold is determined as abnormal data.
2. The method according to claim 1, characterized in that The obtaining of a kernel function according to the local density and time attributes of the geographic surveying and mapping data includes: Calculating the local density of the geographic surveying and mapping data point within a preset neighborhood, if the data point is located in a data point dense area, using a first bandwidth value as the bandwidth of the Gaussian kernel function; otherwise, using a second bandwidth value as the bandwidth of the Gaussian kernel function, where the first bandwidth value is smaller than the second bandwidth value; The time distance between the acquisition time of the data point and the current time is calculated, a time decay factor is obtained according to the time distance, and the time decay factor is used as the weight of the Gaussian kernel function.
3. The method according to claim 1, characterized in that The method of calculating the overall density function of the geographic surveying and mapping data by using the kernel function includes: For each target data point in the geographic surveying and mapping data set, the remaining data points are substituted into the kernel function centered on the target data point to obtain the kernel function contribution value of each remaining data point to the target data point; Summing all kernel function contribution values for the target data point to obtain the overall density function value at the target data point; Traverse all geographic mapping data points to obtain the overall density function of the entire geographic mapping data set.
4. The method according to claim 1, wherein The step size is obtained based on the data gradient and local density curvature in the neighborhood of the current search position, including: Calculate the gradient vector of the current search position on the overall density function, where the direction of the gradient vector points to the direction in which the density increases fastest; Calculating the Hessian matrix of the overall density function in the neighborhood of the current search position, taking the eigenvalue of the Hessian matrix with the largest absolute value as the local density curvature, and inputting the absolute value of the local density curvature into a monotonically decreasing function to obtain a step size adjustment factor; The modulus of the gradient vector is multiplied by the basic step size, and the product of the multiplication result and the step size adjustment factor is used as the moving step size, and the direction of the moving step size is consistent with the direction of the gradient vector.
5. The method according to claim 1, characterized in that The step of searching for the local maximum of the overall density function by using a moving step and determining a density attraction point includes: Starting from the initial point, the current search position point is moved a corresponding distance in the direction determined by the current moving step to reach a new search position point; the moving step of the new search position point is recalculated based on the data gradient and local density curvature in the neighborhood of the current search position, and the process is repeated until: the moving step is lower than the convergence threshold, or the change in the search position point in multiple consecutive iterations is less than the minimum displacement value, then the current search position point is taken as a local maximum point, and the local maximum point is determined as a density attraction point.
6. The method according to claim 1, characterized in that The calculating the influence domain overlap of the density attraction point and filtering the density attraction points based on the influence domain overlap includes: The circular area with the density attraction point as the center and a radius equal to the attraction point density value multiplied by the influence domain coefficient is used as the influence domain of the density attraction point; The area intersection between the influence domains of any two density attraction points is calculated, and the ratio of the area intersection to the area of the smallest of the two influence domains is taken as the influence domain overlap. When the influence domain overlap between any two density attraction points exceeds the overlap threshold, the density attraction point with the largest density value is retained, and the density attraction point with the smallest density value is removed.
7. The method according to claim 1, characterized in that The calculating of the fuzzy membership of each geographic surveying and mapping data point to the density attraction point includes: For each geographic mapping data point, calculate the Euclidean distance between it and each determined density attractor; Based on the Euclidean distance, the following formula is used to calculate the fuzzy membership of the data point i to the kth density attraction point: : in is the distance between data point i and density attraction point k, C is the total number of density attraction points, m is the fuzzy factor and m>
1.
8. The method according to claim 1, characterized in that The calculating the abnormality degree of the data point based on the fuzzy membership of the data point and the distance from each density attraction point includes: Calculating a membership weighted distance of the geographic surveying and mapping data point to each of the density attraction points, wherein the membership weighted distance is the fuzzy membership degree of the data point to the density attraction point multiplied by the Euclidean distance from the data point to the density attraction point; The membership weighted distances of the data point to all the density attraction points are summed, and the sum is used as the abnormality degree of the data point.
9. A geographic surveying and mapping data anomaly monitoring system based on the Internet of Things, characterized in that: include: a density estimation unit, configured to obtain geographic surveying and mapping data collected by an IoT device, obtain a kernel function based on the local density and time attributes of the geographic surveying and mapping data, and calculate an overall density function of the geographic surveying and mapping data using the kernel function; a search and filtering unit configured to determine an initial point based on the prior characteristics of the geographic surveying and mapping data, and to determine a moving step size based on the data gradient and local density curvature in the neighborhood of the current search position at each movement, and to search for a local maximum of the overall density function using the moving step size, and to determine a density attraction point; Calculate the influence domain overlap of the density attraction points, and filter the density attraction points based on the influence domain overlap; The anomaly detection unit is used to calculate the fuzzy membership of each geographic surveying and mapping data point to the density attraction point, calculate the abnormality of the data point based on the fuzzy membership of the data point and the distance to each density attraction point, and determine the geographic surveying and mapping data point whose abnormality exceeds a threshold as abnormal data.
10. The system according to claim 9, characterized in that The obtaining of a kernel function according to the local density and time attributes of the geographic surveying and mapping data includes: Calculating the local density of the geographic surveying and mapping data point within a preset neighborhood, if the data point is located in a data point dense area, using a first bandwidth value as the bandwidth of the Gaussian kernel function; otherwise, using a second bandwidth value as the bandwidth of the Gaussian kernel function, where the first bandwidth value is smaller than the second bandwidth value; The time distance between the acquisition time of the data point and the current time is calculated, a time decay factor is obtained according to the time distance, and the time decay factor is used as the weight of the Gaussian kernel function.
Citation Information
Patent Citations
Mining method for abnormal data of basic geographic information
CN104035985A
Surveying and mapping data processing method
CN118093673A