Geographic information data processing method and system
Through self-activating data nodes and edge-fog-cloud collaborative processing architecture, the transmission reliability and resource utilization problems of geographic information systems in unstable network environments are solved, efficient data compression and dynamic resource allocation are achieved, and the reliability and efficiency of data processing are improved.
Patent Information
- Application Number
- CN202510655731.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing geographic information systems have low transmission reliability, low compression efficiency, unreasonable resource utilization, and insufficient collaborative processing capabilities in unstable network environments, and are unable to effectively respond to changes in network status and data hotspot processing delays.
Through self-activating data nodes, network environment perception, multi-scale spatiotemporal correlation analysis and edge-fog-cloud collaborative processing architecture, adaptive data transmission and dynamic resource allocation are achieved. Combined with neuron activation mechanism and distributed collaborative retransmission, transmission strategy and resource utilization are optimized.
It improves data transmission reliability, reduces transmission and storage overhead, optimizes resource utilization efficiency, enhances system adaptability and data processing quality, and ensures efficient data processing in complex network environments.
Smart Images

Figure CN120596584A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of geographic information systems, and more particularly, to a geographic information data processing method and system. Background Art
[0002] With the rapid development of smart city construction and geographic information systems, IoT devices, sensor networks, and mobile terminals continue to generate massive amounts of geographic data. This data often has strong temporal and spatial correlations and requires reliable transmission and efficient processing. However, existing technologies have obvious shortcomings in the following aspects:
[0003] Traditional geographic data transmission methods lack the ability to perceive and adapt to network environments. In emergency scenarios with large network fluctuations or in resource-constrained areas, transmission reliability is low and they cannot effectively respond to changes in network status.
[0004] Common data compression methods fail to fully recognize and exploit the spatiotemporal correlations in geographic data streams, resulting in high transmission and storage overheads and low data compression efficiency;
[0005] Existing geospatial data stream processing architectures typically adopt a strategy of evenly distributing computing resources, but lack the ability to dynamically identify data hotspots and schedule resources. This results in increased processing delays in data hotspot areas and low resource utilization in non-hotspot areas.
[0006] The traditional geographic information system architecture fails to effectively integrate edge computing, fog computing, and cloud computing resources, and cannot adaptively select the optimal processing nodes based on data characteristics and network conditions, affecting the overall efficiency and reliability of the system.
[0007] The above problems seriously restrict the application effect of geographic information systems in unstable network environments, especially in emergency scenarios, resource-constrained areas or environments with changeable network conditions. A geographic information data processing method is needed that can perceive the network environment, adaptively make decisions and transmission strategies, efficiently utilize spatiotemporal correlations, and perform collaborative processing. Summary of the Invention
[0008] The present invention provides a geographic information data processing method and system to solve the technical problems in related technologies such as low transmission reliability, low compression efficiency, unreasonable resource utilization and insufficient collaborative processing capability of geographic information data in an unstable network environment.
[0009] The present invention provides a method for processing geographic information data, comprising the following steps:
[0010] Encapsulate geographic data as self-activating data nodes;
[0011] Monitor network status parameters in real time through the network environment perception function to obtain the current network environment status score;
[0012] Based on the neuron activation principle, an activation threshold function is constructed to determine whether a data node should transmit and its transmission strategy according to the environment state and node state information.
[0013] Apply wavelet transform to the geographic data stream to conduct multi-scale spatiotemporal correlation analysis, build a spatiotemporal prediction model, and achieve adaptive data compression;
[0014] Perform spatiotemporal analysis based on compressed data streams to predict data hotspot distribution and dynamically adjust computing resource allocation based on hotspot prediction results.
[0015] Based on the resource allocation results, a three-level edge-fog-cloud collaborative processing architecture is constructed to achieve distributed processing and reliable transmission of data.
[0016] In a preferred embodiment, the self-activation data node includes three components: geographic data content, node status information, and a self-activation behavior set.
[0017] In a preferred embodiment, the network environment perception function evaluates the current network environment status in real time by monitoring network bandwidth, network delay, network jitter, network packet loss rate and other parameters.
[0018] In a preferred embodiment, the activation threshold function is based on the environmental status, node status information and priority parameters, and determines whether the data node is activated for transmission by calculating the comparison result of the weighted feature value and the preset threshold, wherein the weighted feature value is calculated by the sum of the products of multiple decision-related features and their corresponding weight coefficients.
[0019] In a preferred embodiment, the activation threshold is dynamically updated through an adaptive adjustment mechanism. The adaptive adjustment mechanism adjusts the threshold size according to a certain learning rate based on the gap between the target transmission quality and the actual transmission quality. When the actual quality is lower than the target quality, the threshold increases, and vice versa, the threshold decreases, thereby realizing the system's adaptive adjustment to changes in the network environment.
[0020] In a preferred embodiment, the multi-scale spatiotemporal correlation analysis includes:
[0021] Applying wavelet transform to geographic data series to obtain multi-scale coefficients;
[0022] Build a spatiotemporal prediction model that uses historical data sequences and spatial context information to predict future data values;
[0023] Calculate the error between the predicted value and the actual value, and select the optimal encoding strategy based on the error distribution characteristics;
[0024] The quality assessment function is used to ensure that the reconstructed data meets the application accuracy requirements.
[0025] In a preferred embodiment, dynamically adjusting computing resource allocation according to hotspot prediction results includes:
[0026] Predict hotspot distribution in different regions at different times based on historical data, contextual factors, and event information;
[0027] Calculate the computing resources required for each area based on hotspot distribution, computational complexity, and service quality requirements;
[0028] Dynamic allocation of resources is achieved by minimizing the weighted square difference between resource requirements and actual allocation;
[0029] Configure multi-level cache for different areas based on hotspot intensity and implement differentiated data processing strategies.
[0030] In a preferred embodiment, the data processing in the constructed edge-fog-cloud three-level collaborative processing architecture includes:
[0031] Select the optimal processing node for each data based on data characteristics and system status;
[0032] Use data sharding strategies to divide large datasets into multiple shards that can be processed in parallel;
[0033] Allocate multiple processing nodes for key data shards to achieve redundant computing and improve system reliability;
[0034] Determine the optimal data transmission plan based on network conditions and data characteristics.
[0035] In a preferred embodiment, a geographic information data processing method further includes data fusion and quality assurance steps:
[0036] Convert heterogeneous data from different sources into a unified format through standardization;
[0037] Adopt semantic-level fusion methods to achieve effective integration of multi-source data;
[0038] Perform quality assessment and uncertainty quantification on fused data;
[0039] Enables automated anomaly detection and correction.
[0040] In a preferred embodiment, a geographic information data processing system is used to execute a geographic information data processing method, comprising:
[0041] A self-activating data node building module for encapsulating geographic data into self-activating data nodes;
[0042] Network environment perception module, used to monitor network status parameters in real time;
[0043] The transmission decision module is used to determine the transmission strategy of the data node based on the neuron activation principle;
[0044] Spatiotemporal correlation analysis module for multi-scale analysis and adaptive compression of geographic data streams;
[0045] Hotspot prediction and resource scheduling module, used to predict data hotspot distribution and dynamically allocate computing resources;
[0046] Collaborative processing module, used to build a three-level collaborative processing architecture of edge-fog-cloud to achieve distributed data processing;
[0047] The data fusion module is used to achieve the fusion and quality assurance of multi-source geographic information data.
[0048] The beneficial effects of the present invention are:
[0049] Improve data transmission reliability: The self-activating data node of the present invention can adaptively adjust the transmission strategy according to the network environment status. Combined with the neuron activation mechanism and the distributed collaborative retransmission mechanism, it significantly improves the transmission success rate in an unstable network environment.
[0050] Reduce transmission and storage overhead: Through multi-scale spatiotemporal correlation analysis and adaptive compression, the present invention fully utilizes the spatiotemporal correlation patterns in geographic data streams and achieves more efficient data compression while ensuring data quality.
[0051] Optimizing resource utilization efficiency: The hotspot prediction and computing resource dynamic scheduling mechanism of the present invention realizes on-demand allocation of computing resources and improves resource utilization efficiency.
[0052] Enhanced System Adaptability: This invention utilizes multiple adaptive mechanisms (including dynamic threshold adjustment, adaptive error coding, and dynamic resource scheduling) to enable the system to adapt to various complex network environments and data load conditions, maintaining stable performance. This adaptive capability makes geographic information systems more reliable and efficient in a variety of application scenarios, such as smart cities and emergency response.
[0053] Improve data processing quality: The multi-source data fusion and quality assurance mechanism of the present invention realizes the standardized processing, semantic fusion, quality assessment and anomaly detection and correction of heterogeneous geographic data, improves the accuracy and reliability of data processing results, and provides a more solid data foundation for decision-making based on geographic information. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a flow chart of a geographic information data processing method of the present invention. DETAILED DESCRIPTION
[0055] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.
[0056] At least one embodiment of the present invention discloses a method for processing geographic information data, such as Figure 1 As shown, the following steps are included:
[0057] Step 1: Encapsulate geographic data into self-activating data nodes;
[0058] It includes the following sub-steps:
[0059] Step 1.1, self-activation data node construction;
[0060] Encapsulate geographic data as self-activating data nodes:
[0061] N=(D,S node ,A);
[0062] Where N represents the self-activated data node; D represents the geographic data content, such as location coordinates, geographic attributes, timestamp, etc.; S node Indicates node status information, including data priority, validity period, data quality score, etc.; A represents the self-activation behavior set, including optional behaviors such as transmission, waiting, compression, and discarding;
[0063] The construction process of self-activating data nodes includes three steps: data analysis, attribute extraction, and behavior definition, laying the foundation for subsequent autonomous decision-making.
[0064] Step 1.2: Network environment perception;
[0065] The network status parameters are monitored in real time through the network environment perception function E(t) = {b(t), l(t), j(t), p(t)}, where E(t) represents the network environment perception function, t represents the current time point, b(t) represents the network bandwidth at time t, l(t) represents the network delay, j(t) represents the network jitter, and p(t) represents the network packet loss rate.
[0066] Each self-activated data node obtains the current network environment parameters by actively detecting or passively receiving network status information, and stores them in the node status information S node middle.
[0067] Step 1.3, environmental status assessment;
[0068] Calculate the network environment status score based on the obtained network environment parameters:
[0069] Q net =∑ i w net,i ·q i ;
[0070] Among them, Q net represents the network environment status score, i represents the index of the network parameter, w net,i Indicates the degree of influence of different network parameters on the overall network status score. For example, the bandwidth parameter weight may be higher than the jitter parameter. i is the normalized quality score of each parameter.
[0071] In addition, according to Q net The value of is used to divide the network environment into four levels: "good" (indicating a stable network status, sufficient bandwidth, and low latency), "general" (indicating a basically stable network status that can meet regular transmission needs), "poor" (indicating significant network fluctuations that may affect data transmission), and "bad" (indicating an extremely unstable network with highly restricted transmission). This serves as the basis for subsequent transmission decisions.
[0072] Step 2: Monitor network status parameters in real time through the network environment perception function to obtain the current network environment status score;
[0073] The specific steps include:
[0074] Step 2.1, construction of activation threshold function;
[0075] Construct the activation threshold function:
[0076] σ(E,S node ,P)=exp[∑ i w act,i ·f i (E,S node ,P)>θ act |;
[0077] Among them, σ(E, S node , P) represents the activation threshold function, E is the environment state, which represents the set of various parameters of the current network environment, including bandwidth, delay, jitter and packet loss rate, etc.; i represents the index of the network parameter; S node is the node status information, including the data node’s priority, validity period, data quality score and other attributes; P is the priority parameter, which is used to indicate the importance and urgency of data transmission; f i is a characteristic function used to evaluate the necessity of data transmission from different dimensions, including data importance score, timeliness score, network status matching degree, etc.; w act,iθ is the activation weight coefficient, which is used to adjust the importance of different features in decision making. The higher the weight value, the greater the influence of the feature on the decision. act is the activation threshold, which is the critical value that determines whether the data node activates the transmission; exp[·] is the indicator function, which outputs 1 (activation) when the condition in the square brackets is met (that is, the sum of the weighted features exceeds the activation threshold), otherwise it outputs 0 (inactivation).
[0078] In some embodiments, the characteristic function may also include additional dimensions such as data dependency score (reflecting the degree of association between data and other data) and energy efficiency score (reflecting the energy consumption efficiency of transmitting the data) to adapt to the needs of different application scenarios.
[0079] Step 2.2, dynamic threshold adjustment;
[0080] Dynamically adjust the activation threshold θ based on the historical transmission success rate and the current network environment status act .
[0081] When the network status is good, lower the threshold to increase the transmission rate;
[0082] When the network status is poor, increase the threshold to reduce transmission conflicts.
[0083] Adaptive adjustment of the threshold is achieved through the function:
[0084] θ act (t+1)=θ act (t)·(1+β·(Q target -Q actual ));
[0085] Among them, θ act (t+1) represents the activation threshold at the next moment; θ act (t) represents the activation threshold at the current moment; β is the learning rate, which controls the speed and amplitude of threshold adjustment; Q target Q is the target transmission quality, which indicates the transmission performance level that the system expects to achieve; actual The actual transmission quality indicates the transmission performance actually measured by the system.
[0086] The threshold adjustment can also adopt a reinforcement learning-based method to balance the transmission success rate S through the reward function rate and energy consumption E cost :
[0087] R reward =α1·S rate -α2·E cost ;
[0088] Among them, R rewardRepresents the total reward value obtained by the system; α1 is the reward weight coefficient of the transmission success rate, which indicates the importance the system attaches to the transmission success rate; S rate is the data transmission success rate, which indicates the proportion of successfully transmitted data packets to the total transmitted data packets; α2 is the penalty weight coefficient of energy consumption, which indicates the sensitivity of the system to energy consumption; E cost is the energy consumption of the transmission process.
[0089] Step 2.3, the excitability propagation process is realized;
[0090] When a key data node is activated and decides to transmit, it notifies other semantically related nodes through the excitability propagation process to increase their activation probability, forming a propagation chain.
[0091] This process is implemented through a recursive function:
[0092]
[0093] Among them, p j (t) is the activation probability of node j at time t, indicating the possibility that node j is activated for data transmission at time t; p j (t+1) is the updated activation probability of node j at the next time point t+1; N a is the activated node set, which includes all activated data nodes at the current time point; a i (t) is the activation intensity of node i, which indicates the degree of activation of the activated node i at time t. The higher the value, the greater the influence. sem,ij is the semantic association between node i and node j, which quantifies the correlation between two data nodes. The higher the value, the closer the association. prop The propagation attenuation coefficient controls the attenuation degree of the activation signal during the propagation process. The value range is between 0 and 1. The smaller the value, the faster the attenuation.
[0094] This approach ensures that relevant data can be transmitted collaboratively at the right time, improving data integrity and effectiveness.
[0095] Step 3: Based on the neuron activation principle, an activation threshold function is constructed to determine whether the data node should transmit and its transmission strategy according to the environment state and node state information;
[0096] It includes the following sub-steps:
[0097] Step 3.1, multi-scale analysis of wavelet transform;
[0098] Apply wavelet transforms to geographic data streams to decompose the data into patterns of variation at different spatial and temporal scales.
[0099] For a geographic data sequence X = {x1, x2, ..., x n}, and obtain the multi-scale coefficient W={a J , d J , d J-1 ,...,d1}, where X represents the original geographic data sequence, x1, x2, x n Represent the values of the 1st, 2nd, and nth data points respectively, and n represents the total length of the sequence; W represents the multi-scale coefficient set obtained after wavelet transform; a J is the approximation coefficient, which indicates the low-frequency approximation of the data at the coarsest scale and reflects the overall trend characteristics of the data; d J d J-1 , d1 represent the detail coefficients of the Jth, J-1th and 1st layers respectively, reflecting the local variation characteristics of the data at this scale; J represents the maximum number of layers of wavelet decomposition, which determines the coarsest scale level of the analysis.
[0100] By analyzing the energy distribution of detail coefficients at each layer, we can identify the main characteristics of data changes at different spatiotemporal scales. The energy distribution here refers to the distribution of the sum of squares of detail coefficients at each layer. Layers with higher energy values indicate more significant data changes at that scale.
[0101] Step 3.2, spatiotemporal prediction model construction;
[0102] Build a spatiotemporal prediction model based on historical data sequences and spatial context:
[0103]
[0104] in, represents the predicted value of the data at time point t+1; f pred Represents a prediction function, which is used to generate a prediction value based on historical data and spatial context; x t-k:t is a historical data sequence of length k, containing continuous data points from time point tk to t; S space,t is the spatial context information at time t, including the data values of the neighboring areas, spatial correlation features, etc.
[0105] This model can be implemented in a variety of ways, including spatiotemporal autoregressive models, recurrent neural networks, etc., and the most suitable model structure is selected according to the data characteristics.
[0106] In some embodiments, a spatiotemporal graph convolutional network (ST-GCN) can be used to process geographic data with complex spatial topological relationships by constructing G = (V, E, A adj) space-time graph, where G represents the constructed space-time graph, which is used to represent the space-time relationship of geographic data; V is the node set, which represents each observation point or area in the geographic space; E is the edge set, which represents the connection relationship between nodes, reflecting the proximity or correlation of geographic locations; A adj It is an adjacency matrix, which is used to quantitatively represent the connection strength or relationship weight between nodes.
[0107] By combining the convolution operation in the time dimension, ST-GCN can effectively capture the spatiotemporal dependencies in geographic data and improve prediction accuracy.
[0108] Step 3.3, adaptive error coding;
[0109] Calculate the error between the predicted value and the actual value Select the optimal coding strategy based on the distribution characteristics of the error. t represents the prediction error at time point t, that is, the difference between the actual value and the predicted value; x t Represents the actual data value at time point t; Represents the predicted data value at time point t.
[0110] Therefore, through the objective function Selecting the optimal encoder in, represents the optimal encoder selected for efficient error encoding; argmin represents finding the parameters that minimize the objective function; E i represents the i-th encoder among the candidate encoders; ε represents the set of all optional candidate encoders; L eval represents a comprehensive evaluation function, which is used to evaluate the coding effect, taking into account both compression rate and reconstruction error; E i (e t ) indicates the use of encoder E i Error e t The result of encoding; B t represents the current bandwidth constraint at time t, which limits the available transmission resources.
[0111] Differentiated compression parameters are applied to regions with different spatiotemporal correlation patterns to ensure data quality in important areas.
[0112] Step 3.4, quality control methods;
[0113] Through the quality evaluation function Ensure that the reconstructed data meets the accuracy requirements of the application, including: represents the quality assessment function, x is the original data, To reconstruct the data, θ qis the quality threshold. When the reconstruction quality does not meet the requirements, the encoding parameters are automatically adjusted or an alternative encoding strategy is selected to ensure a balance between data quality and compression efficiency.
[0114] Step 4: Apply wavelet transform to the geographic data stream to be transmitted to perform multi-scale spatiotemporal correlation analysis, build a spatiotemporal prediction model, and achieve adaptive data compression;
[0115] The specific steps include:
[0116] Step 4.1, spatiotemporal hotspot prediction;
[0117] Based on historical data traffic and current trends, predict the data intensity of each geographical area in the future period.
[0118] By the function H(r, t) = f hot (D hist , F contextual , I event ) modeling the hotspot distribution of region r at time t, where H(r, t) represents the hotspot intensity value of region r at time t. A higher value indicates a higher data flow in the region. r represents a specific geographical area, which can be a spatial unit such as a city area, a road section, or a monitoring point. t represents a time point, which can be accurate to different time granularities such as minutes, hours, or days. f hot represents the hotspot prediction function, which is used to calculate the hotspot intensity based on the input parameters; D hist The historical data traffic pattern includes the data traffic statistics of the area in the past period of time, such as average traffic, peak traffic, periodic changes, etc.; F contextual Contextual features include external factors that affect data flow, such as event calendar, weather conditions, and time characteristics; I event Real-time event information refers to current emergencies that may affect data traffic, such as traffic accidents, temporary road closures, and public emergencies.
[0119] The data intensity exceeds the threshold θ h The area is marked as a hot spot, forming a hot spot map M hotspot .
[0120] Step 4.2, resource requirement estimation;
[0121] Based on the hotspot prediction results and data complexity factors, estimate the computing resources required for each region.
[0122] The resource demand function is used to calculate the resource demand of region r at time t:
[0123] R res (r,t)=H(r,t)·C comp (r,t)·Sqos (r,t);
[0124] Among them, R res (r, t) represents the computing resource demand of region r at time t, which is a comprehensive indicator that reflects the computing power required by the region at a specific time point; H(r, t) represents the hotspot intensity of region r at time t. A higher value indicates a larger data flow in the region and requires more computing resources; C comp (r, t) represents the data complexity coefficient, which takes into account factors such as data type and processing algorithm complexity. Data processing with high complexity requires more computing resources; S qos (r, t) represents the service quality requirement coefficient, which reflects the service quality requirement for region r at time t. The higher the service quality requirement, the more resources need to be allocated to ensure processing speed and accuracy.
[0125] Summarize the resource requirements of each region to form a resource requirement matrix:
[0126] RM={R res (r,t)|r∈Regions,t∈T future};
[0127] Among them, RM represents the resource demand matrix, which contains the resource demand forecast of all regions in the future time period; R res (r, t) represents the computing resource demand of region r at time t; Regions represents the set of all geographical regions covered by the system; T future Represents a collection of future time periods, usually a series of discrete time points, used for resource planning and scheduling.
[0128] Step 4.3, dynamic resource scheduling;
[0129] Based on the resource demand matrix, dynamic allocation of computing resources is achieved.
[0130] Resource allocation by optimizing the objective function:
[0131] min∑ r,t w res (r,t)·(R res (r,t)-A alloc (r,t)) 2 ;
[0132] Among them, min represents minimizing the objective function, that is, finding a resource allocation solution that makes the objective function achieve the minimum value; ∑ r,t represents the sum of all combinations of regions r and time points t; w res (r, t) is the resource weight coefficient, reflecting the priority of region r at time t. The higher the value, the more important the region is.res (r, t) represents the computing resource demand of region r at time t; A alloc (r, t) is the actual amount of resources allocated to region r at time t; (R res (r, t)-A alloc (r, t)) 2 It represents the square of the difference between resource demand and actual allocation, and is used to quantify the degree of mismatch in resource allocation.
[0133] Ensure that high-priority and hotspot areas receive sufficient resources while avoiding resource waste and processing delays.
[0134] Step 4.4, multi-level cache configuration;
[0135] Configure multi-level cache for hot spots to improve data access efficiency.
[0136] Determine the cache level based on the hotspot intensity H(r, t):
[0137]
[0138] Where H(r, t) represents the hotspot intensity value of region r at time t. A higher value indicates a greater data flow in the region. h represents the hotspot determination threshold, which is used to standardize the hotspot intensity; log represents the natural logarithm function; L cache (r, t) represents the cache level allocated to region r at time t, where a higher level means more cache resources are allocated; Represents the rounding up function to ensure that the cache level is a positive integer;
[0139] And allocate cache capacity accordingly:
[0140]
[0141] Among them, C cache (r, t) represents the actual cache capacity allocated to region r at time t; C base represents the basic cache capacity, which is the cache capacity benchmark value allocated to the lowest level hotspot area; α cache Represents the level capacity coefficient, a constant greater than 1, which determines the capacity multiplication relationship between different cache levels; L cache (r, t) represents the cache level allocated to region r at time t, where a higher level means more cache resources are allocated.
[0142] Differentiated data prefetching and cache update strategies are implemented for areas with different hotspot intensities to further improve data access response speed.
[0143] Step 5: Perform spatiotemporal analysis based on the compressed data stream to predict data hotspot distribution and dynamically adjust computing resource allocation based on the hotspot prediction results.
[0144] It includes the following sub-steps:
[0145] Step 5.1, multi-level node collaborative processing;
[0146] Build an edge-fog-cloud three-level collaborative processing architecture to achieve distributed and efficient processing of geographic data.
[0147] Select the optimal processing node for data d through the collaborative processing function:
[0148] P proc (d)=select(P edge , P fog , P cloud |d,τ);
[0149] Among them, P proc (d) represents the processing node selection result of data d, that is, the final selected processing node; d represents geographic information data, including characteristic information such as data type, scale, and complexity; select() represents the node selection function, which selects the optimal processing node according to the input parameters; P edge represents a collection of edge processing nodes, located near the data collection point, suitable for processing low-latency, small-scale tasks; P fog represents a set of fog processing nodes, located between the edge and core of the network, with medium computing power; P cloud It represents a collection of cloud processing nodes, located in data centers, with powerful computing and storage capabilities. τ represents the real-time requirement for data processing. The lower the value, the lower the tolerance for processing delays.
[0150] Small-scale, high-real-time tasks suitable for edge processing are completed at edge nodes; medium-scale tasks with moderate timeliness requirements are handled by fog nodes; large-scale, complex analysis tasks are completed by the cloud.
[0151] Step 5.2, adaptive sharding and task scheduling;
[0152] Based on data characteristics and node capabilities, data adaptive sharding and intelligent task scheduling are achieved.
[0153] Divide the data d into k shards through the sharding strategy function:
[0154] S split (d,N nodes )={s1,s2,…,s k};
[0155] Among them, S splitrepresents the sharding strategy function, which is used to split the data into multiple parts that can be processed in parallel; d represents geographic information data; N nodes Represents the set of processing nodes available in the system; s1, s2, s k They represent the 1st, 2nd, and kth data shard sets obtained after sharding, respectively; k represents the number of shards, which is dynamically determined based on the data size and the number of available nodes.
[0156] Each shard i Assign to the most suitable processing node n j ∈N nodes , fitness calculation:
[0157] F fit (s i , n j )=w fit,1 ·C match (s i , n j )+w fit,2 ·T cost (s i , n j )+w fit,3 ·E energy (s i , n j )
[0158] Among them, F fit (s i , n j ) represents shard s i With node n j The higher the value, the more suitable it is for processing at this node; C match (s i , n j ) represents the calculation of matching degree, measuring node n j Computing power and shards i The degree of matching of processing requirements; cost (s i , n j ) represents the transmission overhead, which measures the fragmentation of s i Transmit to node n j The time and network resource consumption required; E energy (s i , n j ) represents the energy consumption index, which measures the node n i Processing shards i Energy consumed; w fit,1 、w fit,2 、w fit,3 Respectively represent the fitness weight coefficients of the calculation matching degree, transmission overhead and energy consumption indicators, and satisfy wfit,1 +w fit,2 +w fit,3 =1.
[0159] Step 5.3, fault tolerance processing and result merging;
[0160] Realize system fault-tolerant processing and merging of multi-source processing results.
[0161] By redundant calculation strategy R redund (s i )={n j1 , n j2 ,...,n jr} is the key data shard i Allocate r processing nodes for parallel processing to improve system reliability. redund (s i ) represents data shard s i Designed redundant computing strategy; i Indicates key data fragments, which are data fragments that need to be processed reliably; n j1 、n j2 、n jr They represent the 1st, 2nd, and rth processing nodes assigned to process the same data shard, respectively; j represents the index number of the processing node; and r represents the redundancy, that is, the number of processing nodes assigned to the same data shard.
[0162] The processing results are passed to the merge function Integrate the output results of each node O i , forming the final output O final , the merging process adopts a confidence-based weighted method to ensure the accuracy of the results. merge Indicates the result merging function; O1, O2, Represent the output result sets from the 1st, 2nd, and k1th processing nodes respectively; k1 represents the total number of output results; O i Represents the output result of the i-th processing node; O final Indicates the final output result after merging.
[0163] When processing high-value data, a voting mechanism can be used Select the result with the highest frequency, or use a screening method based on confidence threshold. final ={O i |C conf (O i )>θ c}, where V vote Represents the voting mechanism function; mode represents the majority function, which is used to select the result with the highest frequency; O1, O2, O kRepresent the output result sets from the 1st, 2nd, and k1th processing nodes respectively; k1 represents the total number of output results; C conf (O i ) represents the result O i The confidence score is used to quantify the reliability of the results; i represents the output result of the i-th processing node; θ c Represents the confidence threshold. Only results with a confidence exceeding this threshold will be included in the final output. final In the confidence threshold based method, it represents the set of all results that meet the confidence requirement.
[0164] Step 5.4, adaptive transmission optimization;
[0165] Optimize transmission strategies based on network conditions and data importance.
[0166] Through the transmission strategy function from the source node n s To the target node n d The data d determines the optimal transmission solution:
[0167] T trans (d,n s ,n d )={path, protocol, priority, QoS};
[0168] Among them, T trans represents the transmission strategy function, which is used to determine the optimal data transmission solution; d represents geographic information data, including attributes such as data type, size, and importance; n s Represents the source node, that is, the node that sends the data; n d It represents the target node, that is, the receiving node of the data; path represents the selected transmission path, which may include multiple relay nodes; protocol represents the selected transmission protocol, such as TCP, UDP, MQTT, etc.; priority represents the priority assigned to data transmission, which affects the order in which data is processed in the network; QoS represents a set of service quality parameters, including bandwidth allocation, delay requirements, reliability level, etc.
[0169] For high-value data, reliable transmission methods are used and higher bandwidth resources are allocated; for time-sensitive data, low-latency paths are selected; for ordinary data, reliability and resource consumption are balanced.
[0170] Step 6: Based on the resource allocation results, a three-level edge-fog-cloud collaborative processing architecture is constructed to achieve distributed processing and reliable transmission of data.
[0171] It includes the following sub-steps:
[0172] Step 6.1: Standardization and unified expression of heterogeneous data;
[0173] For geographic information data from different sources and in different formats, standardized processing and unified expression are achieved.
[0174] The original data d is normalized by the normalization function raw According to the standard frame F frame Convert to standardized data d std :
[0175] S stand (d raw ,F frame )=d std ;
[0176] Among them, S stand Represents a standardization function, which is used to convert the original data into data that conforms to the standard format; d raw Represents the raw geographic information data to be processed, which may come from different sources and in different formats; frame Represents a standardized reference framework that defines the standard specifications and structures that data should follow; d std Indicates data that has been standardized and conforms to unified expression formats and specifications.
[0177] The standardization process includes coordinate system conversion, unit unification, time correction, and format normalization, ensuring that data from different sources are comparable and fusible within the same reference system. Specific standardization processes are designed for different data types (vector, raster, point cloud, etc.).
[0178] Step 6.2, semantic fusion of multi-source data;
[0179] Multi-source data fusion is performed based on the semantic information of the data.
[0180] Multiple data sources are fused into a unified dataset through semantic fusion functions:
[0181]
[0182] Among them, F sem represents the semantic fusion function, which is used to integrate multiple heterogeneous data sources into a fused dataset with unified semantics; D1, D2, They represent the 1st, 2nd, and n1th datasets from different sources, respectively. Each dataset may have different formats, structures, and semantics. n1 represents the number of data sources involved in the fusion. D fused It represents a unified dataset generated after semantic fusion, which contains comprehensive information from various data sources.
[0183] The fusion process takes into account the spatial, temporal and semantic relationships of the data, adopts a graph-based data representation method, maps entities from different data sources into graph nodes, and maps relationships between entities into graph edges, and realizes data fusion through graph operations.
[0184] For example, in traffic monitoring scenarios, video data from cameras, traffic flow data from sensors, and location data from mobile devices can be integrated to construct a comprehensive traffic status map, providing a more comprehensive decision-making basis for traffic management.
[0185] Step 6.3, quality assessment and uncertainty quantification;
[0186] Perform quality assessment and uncertainty quantification on the fused geographic information data.
[0187] The multi-dimensional quality indicators of the fused data are calculated through the quality assessment function:
[0188] Q assess (D fused )={q1,q2…,q m};
[0189] Among them, Q assess Denotes the quality assessment function, which is used to evaluate the quality of fused data; D fused Represents the geographic information dataset after fusion processing; q1, q2, q m They represent the quality scores of the 1st, 2nd, and mth specific dimensions respectively; m represents the total number of quality assessment dimensions.
[0190] At the same time, the uncertainty of the data in different dimensions is estimated through the uncertainty quantification function:
[0191]
[0192] Among them, U uncert Represents the uncertainty quantification function, which is used to quantify the degree of uncertainty in the data; each u1, u2, They represent the uncertainty scores of the 1st, 2nd, and k2th specific dimensions respectively; k2 represents the total number of dimensions for uncertainty assessment.
[0193] The system uses an entropy-based method to quantify the information uncertainty in the data. The higher the entropy value, the greater the uncertainty of the data and the lower the reliability, which requires special attention or additional processing in subsequent analysis.
[0194] Step 6.4, automated anomaly detection and correction;
[0195] Implement automated anomaly detection and correction processes for geographic information data.
[0196] Identify the set of outliers in data d through the anomaly detection function:
[0197]
[0198] Among them, A anom represents anomaly detection function, which is used to identify abnormal points from the data; d represents the geographic information data to be detected - M model Represents the model used for anomaly detection, which can be a statistical model, a machine learning model, or a rule model; a1, a2, They represent the 1st, 2nd, and j1th abnormal data points respectively; j1 represents the total number of abnormal points detected.
[0199] For the detected outliers, the corrected data d′ is generated by the correction function:
[0200] C correct (a i ,d,K context ) = d′;
[0201] Among them, C correct represents an anomaly correction function, which is used to repair abnormal data points; a i Indicates the abnormal data point to be corrected; d represents geographic information data, providing reference information required for correction; K context represents contextual information, including spatiotemporal correlation, domain knowledge, and historical data; d′ represents the corrected dataset, in which the outliers have been fixed.
[0202] Anomaly detection uses a combination of techniques, including statistical methods, machine learning methods, and expert rule methods, to improve the accuracy and robustness of detection.
[0203] Application examples of this implementation:
[0204] Smart City Traffic Management System:
[0205] The geographic information data processing method of the present invention was deployed in a large city's intelligent traffic management system, achieving efficient geographic data processing in a complex network environment. The system includes multiple data source devices, including traffic cameras, roadside units, and vehicle-mounted terminals, distributed throughout the city. These devices need to transmit and process large amounts of geographic location data in real time under fluctuating network conditions.
[0206] Implementation steps:
[0207] Self-activation data node construction: The system encapsulates vehicle trajectory data from traffic cameras, traffic flow data from roadside radars, and GPS location data from vehicle terminals into self-activation data nodes. For example, for vehicle trajectory data captured by a traffic camera, its structure is:
[0208] N vehicle =(D trajectory , S quality , A behaviors );
[0209] Among them, N vehicle represents the vehicle trajectory data, D trajectory Contains vehicle ID, timestamp, geographic coordinates, speed and other information; S quality Contains status information such as data integrity, accuracy score and timeliness; A behaviors Contains behavioral functions such as "automatic compression", "priority adjustment", and "autonomous transmission decision".
[0210] Network environment perception: The system deploys network quality monitoring nodes in different areas of the city to detect the network status of each area in real time.
[0211] During peak hours, the network environment parameters for a commercial area are: bandwidth b(t) = 2.8 Mbps, latency l(t) = 120 ms, jitter j(t) = 45 ms, and packet loss rate p(t) = 3.2%. Using the environmental status scoring function, the system calculates the network quality score for this area as 72 out of 100.
[0212] Data transmission decision: Based on the network status score of the area and the importance of the data captured by the traffic camera, the system applies the activation threshold function to calculate the transmission decision. For example, for a piece of data that detects a traffic accident, its priority P is set to 0.9 (out of 1.0), and the final calculated activation score is 0.86, which exceeds the current threshold of 0.65, so the system decides to transmit the data immediately. For ordinary traffic statistics, the system calculates an activation score of 0.59, which is below the threshold, so the system decides to temporarily store it and wait for network conditions to improve or reduce data accuracy before transmitting it.
[0213] Multi-scale spatiotemporal analysis and compression: The system applies wavelet transforms to multi-scale analysis of the large amount of vehicle trajectory data that needs to be transmitted. By analyzing the spatiotemporal correlations of traffic flow data in the area over the past 30 minutes, the system constructs a traffic flow prediction model. This model predicts traffic flow distribution within the next five minutes based on current and historical traffic flow data, with a prediction error of only 8.3%. Based on this prediction model, the system adaptively compresses the traffic flow data, reducing the original data volume by 67% while maintaining the accuracy of the reconstructed data above 95%.
[0214] Hotspot Prediction and Resource Scheduling: The system predicted that the city center would become a data processing hotspot within the next two hours due to an upcoming large-scale event. Based on this prediction, the system preemptively dispatched computing resources from the suburbs to edge servers and fog nodes in the central area. Specifically, the system increased the computing resources of the central edge nodes from 12 CPU cores to 28, expanded the memory from 32GB to 80GB, and pre-cached historical data and models for the relevant area. This predictive resource scheduling reduced system response time by 43% during the event and increased service quality satisfaction by 35%.
[0215] Three-level collaborative processing: When processing traffic congestion warning tasks, the system automatically selects the optimal processing node combination based on data characteristics and processing requirements:
[0216] Assign low-latency tasks such as real-time traffic flow detection to roadside edge nodes for processing;
[0217] Assign regional congestion situation analysis tasks to regional fog computing nodes for processing;
[0218] Assign complex tasks such as city-wide traffic pattern mining and congestion prediction to the cloud for processing.
[0219] The system also implements a redundant computing strategy for key data. For example, for traffic accident detection tasks, the system simultaneously assigns them to two different edge nodes for parallel processing, and then merges their results through confidence weighting, improving the accuracy and reliability of detection.
[0220] Multi-source data fusion: The system successfully integrates visual data from traffic cameras, traffic flow data from radar, and GPS trajectory data from onboard terminals to generate a highly accurate real-time traffic situation map. During a major event, the system accurately identified five potential traffic congestion points by integrating multi-source data and issued a 10-minute advance warning, enabling traffic management to implement preemptive diversion measures, effectively alleviating traffic congestion and reducing average vehicle travel time by 23%.
[0221] Application effect:
[0222] By deploying the geographic information data processing method of the present invention, the intelligent traffic management system achieves the following significant effects:
[0223] In an environment with large fluctuations in network bandwidth (bandwidth changes up to 60%), the system can still maintain stable data transmission efficiency, and the transmission success rate of key data reaches 99.2%.
[0224] Data processing delay has been significantly reduced from an average of 2.8 seconds to 0.9 seconds, meeting the needs of real-time traffic management.
[0225] System resource utilization efficiency has increased by 42%, and the amount of data that can be processed with the same hardware configuration has increased by 2.3 times.
[0226] The warning accuracy rate has increased to 91.5%, 18 percentage points higher than traditional methods, significantly reducing traffic accidents and congestion.
[0227] System maintenance costs were reduced by 31%, primarily due to the adaptive processing mechanism that reduced the need for manual intervention.
[0228] Natural Disaster Monitoring and Emergency Response System:
[0229] In an earthquake-prone region, the geographic information data processing method of the present invention was deployed in an earthquake monitoring and emergency response system. This system integrates multi-source geographic information data, including a distributed seismic sensor network, drone inspection systems, and satellite remote sensing data, enabling efficient data processing and decision support during natural disasters such as earthquakes.
[0230] Implementation steps:
[0231] Self-activation data node construction: The system encapsulates earthquake waveform data collected by seismic sensors, disaster area images taken by drones, and surface deformation data acquired by satellites into self-activation data nodes. Taking seismic sensors as an example, the self-activation data node structure is as follows:
[0232] N seismic =(D waveform , S sensor , A emergency );
[0233] Among them, N seismic represents the self-activation data of the seismic sensor, D waveform Contains seismic waveform data, timestamps, and sensor locations; S sensor Contains information such as sensor status, battery level, accuracy rating, etc.; A emergency Contains behavioral functions such as "emergency detection", "emergency transmission", and "survival mode switching".
[0234] Network environment perception: After the earthquake, the system detected that communication networks in some areas were severely affected. Network parameters in one disaster area were: bandwidth b(t) = 0.6 Mbps, latency l(t) = 780 ms, jitter j(t) = 230 ms, and packet loss rate p(t) = 21.5%. The system calculated a network quality score of only 28 out of 100 for this area, indicating an extremely poor network environment.
[0235] Data transmission decisions: The system prioritizes different types of disaster data. For data showing clear earthquake waveforms, its priority (P) is set at 0.98, resulting in an activation score of 0.92, significantly exceeding the emergency threshold of 0.45. The system immediately initiates emergency transmission mode. Simultaneously, the system increases the sensor sampling rate to its maximum to capture more detailed earthquake waveform data.
[0236] Multi-scale spatiotemporal analysis and compression: The system applies multi-scale wavelet analysis to high-resolution drone-captured images of disaster areas, extracting key features (such as damaged buildings, blocked roads, and concentrated areas of people). After identifying key areas, the system applies a low compression rate (retaining over 90% of the information) to these areas, while using a high compression rate (retaining only 40% of the information) for non-critical areas. This differentiated compression strategy reduces the overall data volume by 76% while preserving the critical information required for disaster assessment.
[0237] Hotspot Prediction and Resource Scheduling: Based on earthquake intensity maps and population density data, the system predicted three potential disaster hotspots. For Area A, predicted to be the hardest hit, the system automatically increased edge computing resources, deploying four mobile computing units and two emergency communication base stations. It also preloaded building information, terrain data, and road network data from the cloud. This pre-allocation of resources enabled the system to complete a preliminary disaster assessment within 30 seconds after receiving first-hand data from the disaster area, three times faster than conventional methods.
[0238] Three-level collaborative processing: In disaster data processing, the system adopts the following collaborative strategies:
[0239] Edge layer: Mobile computing units deployed in disaster areas are responsible for processing real-time sensor data and low-altitude drone imagery, performing tasks such as preliminary disaster assessment and vital sign detection.
[0240] Fog layer: The fog nodes in the regional command center are responsible for integrating data from multiple edge nodes, generating a regional disaster situation map, and coordinating rescue resources within the region.
[0241] Cloud layer: The remote data center is responsible for global disaster analysis, historical data comparison, and large-scale rescue resource scheduling optimization calculations.
[0242] For critical life detection data, the system adopts a triple redundant computing strategy, arranging processing at the edge, fog, and cloud layers simultaneously, and determining the final result through a majority voting mechanism, ensuring data processing reliability under harsh conditions.
[0243] Multi-source data fusion: The system successfully integrates earthquake waveform data, drone imagery, satellite remote sensing, and social media data to generate a comprehensive disaster situation map. Through semantic-level data fusion, the system automatically identifies key information such as the extent of building damage, the location of road disruptions, and the possible locations of trapped people. During a response to a magnitude 5.8 earthquake, the system accurately identified 24 high-risk locations by integrating multi-source data, prioritized rescue efforts, and successfully rescued 89 trapped people.
[0244] Application effect:
[0245] By deploying the geographic information data processing method of the present invention, the earthquake monitoring and emergency response system achieves the following significant effects:
[0246] Under extreme conditions of partial interruption of the communication network, the system can still maintain a key data transmission success rate of more than 85%, ensuring the timely transmission of disaster information.
[0247] The disaster assessment time has been shortened from the traditional 2 to 3 hours to 15 to 30 minutes, greatly improving the emergency response speed.
[0248] The false alarm rate was reduced from 22% to 4.5%, significantly improving the accuracy of early warning and disaster assessment.
[0249] The system's adaptive resource scheduling maximizes the effectiveness of limited computing and communication resources, supporting an efficiency improvement of approximately 65%.
[0250] The comprehensive situation map generated by the fusion of multi-source data enables decision makers to fully understand the disaster situation, improves the efficiency of rescue resource utilization by about 47%, and saves an estimated 23% additional lives.
[0251] These two application examples fully demonstrate the present invention's ability to process geographic information data in complex network environments, validating the effectiveness, practicality, and advancement of the method. Whether in everyday smart city management or extreme disaster situations, the present invention can provide an efficient and reliable geographic information data processing solution.
[0252] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.
Claims
1. A geographic information data processing method, characterized in that: The following steps are involved: Encapsulate geographic data as self-activating data nodes; Monitor network status parameters in real time through the network environment perception function to obtain the current network environment status score; Based on the neuron activation principle, an activation threshold function is constructed to determine whether a data node should transmit and its transmission strategy according to the environment state and node state information. Apply wavelet transform to the geographic data stream to conduct multi-scale spatiotemporal correlation analysis, build a spatiotemporal prediction model, and achieve adaptive data compression; Perform spatiotemporal analysis based on compressed data streams to predict data hotspot distribution and dynamically adjust computing resource allocation based on hotspot prediction results. Based on the resource allocation results, a three-level edge-fog-cloud collaborative processing architecture is constructed to achieve distributed processing and reliable transmission of data.
2. A geographic information data processing method according to claim 1, characterized in that: The self-activation data node includes three components: geographic data content, node status information and self-activation behavior set.
3. A geographic information data processing method according to claim 1, characterized in that: The network environment perception function evaluates the current network environment status in real time by monitoring network bandwidth, network delay, network jitter, network packet loss rate and other parameters.
4. A geographic information data processing method according to claim 1, characterized in that: The activation threshold function is based on the environmental status, node status information and priority parameters. It determines whether the data node is activated for transmission by calculating the comparison result between the weighted feature value and the preset threshold. The weighted feature value is calculated by the sum of the products of multiple decision-related features and their corresponding weight coefficients.
5. A geographic information data processing method according to claim 4, characterized in that: The activation threshold is dynamically updated through an adaptive adjustment mechanism. The adaptive adjustment mechanism adjusts the threshold size according to a certain learning rate based on the gap between the target transmission quality and the actual transmission quality. When the actual quality is lower than the target quality, the threshold increases, and vice versa, the threshold decreases, thereby realizing the system's adaptive adjustment to changes in the network environment.
6. A geographic information data processing method according to claim 1, characterized in that: The multi-scale spatiotemporal correlation analysis includes: Applying wavelet transform to geographic data series to obtain multi-scale coefficients; Build a spatiotemporal prediction model that uses historical data sequences and spatial context information to predict future data values; Calculate the error between the predicted value and the actual value, and select the optimal encoding strategy based on the error distribution characteristics; The quality assessment function is used to ensure that the reconstructed data meets the application accuracy requirements.
7. A geographic information data processing method according to claim 1, characterized in that: The dynamically adjusting computing resource allocation according to the hotspot prediction result includes: Predict hotspot distribution in different regions at different times based on historical data, contextual factors, and event information; Calculate the computing resources required for each area based on hotspot distribution, computational complexity, and service quality requirements; Dynamic allocation of resources is achieved by minimizing the weighted square difference between resource requirements and actual allocation; Configure multi-level cache for different areas based on hotspot intensity and implement differentiated data processing strategies.
8. A geographic information data processing method according to claim 1, characterized in that: The data processing in the constructed edge fog cloud three-level collaborative processing architecture includes: Select the optimal processing node for each data based on data characteristics and system status; Use data sharding strategies to divide large datasets into multiple shards that can be processed in parallel; Allocate multiple processing nodes for key data shards to achieve redundant computing and improve system reliability; Determine the optimal data transmission plan based on network conditions and data characteristics.
9. A geographic information data processing method according to claim 1, characterized in that: It also includes data fusion and quality assurance steps: Convert heterogeneous data from different sources into a unified format through standardization; Adopt semantic-level fusion methods to achieve effective integration of multi-source data; Perform quality assessment and uncertainty quantification on fused data; Enables automated anomaly detection and correction.
10. A geographic information data processing system, configured to execute a geographic information data processing method according to any one of claims 1 to 9, characterized in that: include: A self-activating data node building module for encapsulating geographic data into self-activating data nodes; Network environment perception module, used to monitor network status parameters in real time; The transmission decision module is used to determine the transmission strategy of the data node based on the neuron activation principle; Spatiotemporal correlation analysis module for multi-scale analysis and adaptive compression of geographic data streams; Hotspot prediction and resource scheduling module, used to predict data hotspot distribution and dynamically allocate computing resources; Collaborative processing module, used to build a three-level collaborative processing architecture of edge-fog-cloud to achieve distributed data processing; The data fusion module is used to achieve the fusion and quality assurance of multi-source geographic information data.
Citation Information
Cited By
Electronic equipment network data detection optimization control method based on edge computing
CN121547456A