A GIS spatial large-scale data parallel processing and analysis method and system
By performing spatial and spatial segmentation and indexing of large-scale data in GIS space, data sharding and task scheduling are optimized, and the problems of inefficiency in the analysis of complex scenarios by traditional GIS processing frameworks are solved, and efficient parallel processing analysis is achieved.
Patent Information
- Application Number
- CN202510369233.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-27
AI Technical Summary
When dealing with complex scenarios such as high-resolution remote sensing image analysis, urban-level three-dimensional modeling and climate change simulation, the traditional GIS processing framework faces the problems of storage bottlenecks, inefficient computing efficiency and insufficient algorithm scalability, especially in spatiotemporal data indexing and parallel computing.
A method of parallel processing and analysis of large-scale data in GIS space is proposed. By segmenting and indexing spatiotemporal data, using the minimum enclosing rectangle and feature value to build an index structure, optimizing data sharding and task scheduling, reducing data skew and transmission, and improving parallel computing efficiency.
By improving the spatiotemporal data index and parallel computing methods, the speed of extracting data to be analyzed from large-scale data in GIS space and the efficiency of parallel computing are significantly improved, which can more effectively support the analysis needs of complex scenarios.
Smart Images

Figure CN119884174B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of parallel computing, and in particular to a method and system for parallel processing and analyzing large-scale spatial data in a GIS. Background Art
[0002] The scale of spatial data in geographic information systems (GIS) has grown exponentially with the rapid development of remote sensing technology, the Internet of Things, and mobile Internet. The traditional GIS processing framework of single machines or small-scale clusters faces problems such as storage bottlenecks, low computing efficiency, and insufficient algorithm scalability, making it difficult to support the needs of complex scenarios such as high-resolution remote sensing image analysis, city-level three-dimensional modeling, and climate change simulation. Parallel processing technology based on distributed computing frameworks (such as Hadoop and Spark) has become an important way to solve the large-scale data computing of GIS. Through distributed storage, parallel task scheduling, and memory computing optimization, it can effectively break through the resource limitations of a single node and realize the full process parallelization of data sharding, task splitting, and result aggregation, thereby significantly improving the processing efficiency of massive spatial data. Distributed computing frameworks show higher performance advantages than small-scale clusters in intensive spatial analysis such as terrain analysis, urban population flow trends, and pollution diffusion analysis. However, as a general big data computing framework, on the one hand, it does not support GIS spatiotemporal data enough, especially in terms of indexing. On the other hand, when performing parallel computing, data skew and data transmission limit the speed of parallel computing. Summary of the invention
[0003] In the analysis and processing of large-scale GIS spatial data, in order to solve the problems of insufficient support for spatiotemporal data index and low parallel efficiency caused by data skew, a first aspect of the present invention provides a method for parallel processing and analysis of large-scale GIS spatial data, the method comprising the following steps:
[0004] The large-scale GIS spatial data to be analyzed is divided into data blocks according to time and space, the characteristic values of the spatiotemporal data are calculated according to the spatial information and time information of the spatiotemporal data, and the index structure of the spatiotemporal data is constructed using the minimum enclosing rectangle and the characteristic value;
[0005] The data to be analyzed is extracted from the data block in parallel using the index structure, the keyword frequency is obtained by sampling the data to be analyzed, the data skewness is calculated under different reduction parallelisms using the keyword frequency, the reduction parallelism with the smallest data skewness is used as the parallelism of the reduction stage, a corresponding relationship between each reduction task and a keyword is established, and the keyword with the largest keyword probability among the keywords corresponding to the reduction task is used as the label of the reduction task;
[0006] The data to be analyzed is divided into slices according to the parallelism and the keyword frequency and a target mapping task is obtained; the reduction task and the target mapping task having the same label are scheduled to the same node;
[0007] The parallelism and the data sharding are used to perform parallel processing and analysis on the data to be analyzed.
[0008] Preferably, the characteristic value of the spatiotemporal data is calculated based on the spatial information and time information of the spatiotemporal data, specifically:
[0009] Obtain the spatial information of the spatiotemporal data, and convert each dimension of the spatial information of all the spatiotemporal data in the data block into binary data of the same length; for each spatiotemporal data, starting from the highest bit, alternately take one bit from the binary data of each dimension to form a spatial binary;
[0010] Encode the time information of spatiotemporal data into time bins of the same length;
[0011] The spatial binary and temporal binary are concatenated, and the concatenated binary is converted into decimal as the characteristic value of the spatiotemporal data.
[0012] Preferably, the index structure of spatiotemporal data is constructed by using the minimum enclosing rectangle and the eigenvalue, specifically:
[0013] Determine the maximum and minimum number of storage entries based on the disk page and spatiotemporal data size; sort the spatiotemporal data according to the characteristic value to obtain the spatiotemporal data sequence;
[0014] Add the spatiotemporal data whose sequence number in the spatiotemporal data sequence is greater than the minimum number of stored entries and not greater than the maximum number of stored entries to the set, add the spatiotemporal data in the set whose spatial binary difference with the previous spatiotemporal data is not zero to the subset, sequentially extract the spatiotemporal data from the subset, calculate the minimum enclosing rectangle of the spatiotemporal data before the extracted spatiotemporal data in the spatiotemporal data sequence, and take the smallest of all the minimum enclosing rectangles of the subset as the leaf of the index structure; delete the leaf object from the spatiotemporal data sequence, and after clearing the set and the subset, repeat the leaf acquisition process until the spatiotemporal data sequence is empty;
[0015] The index structure is constructed based on the leaves in a recursive manner.
[0016] Preferably, the method of calculating the degree of data skewness under different reduction parallelisms by using the keyword frequency is specifically as follows:
[0017] For each preset reduction parallelism, the keywords are combined into combinations of the number of reduction parallelisms so that the sum of the keyword probabilities of each combination is closest, the sum of the probabilities of the keywords in the combination is taken as the probability of the combination, and the variance or standard deviation of the probabilities of all combinations is calculated to obtain the degree of data skew corresponding to the reduction parallelism.
[0018] Preferably, the data to be analyzed is segmented according to the parallelism and the keyword frequency and the target mapping task is obtained, specifically:
[0019] Sort the keywords in descending order of their frequency;
[0020] Associating the keywords of the number of reduction tasks mentioned above with a target mapping task respectively, and using the associated keywords as labels of the target mapping task;
[0021] When sharding, the data to be analyzed that contains the target mapping task label is preferentially allocated to the shard corresponding to the target mapping task.
[0022] In a second aspect of the present invention, a GIS spatial large-scale data parallel processing and analysis system is provided, the system comprising the following modules:
[0023] The index module is used to divide the large-scale GIS spatial data to be analyzed into data blocks according to time and space, calculate the characteristic value of the spatiotemporal data according to the spatial information and time information of the spatiotemporal data, and construct the index structure of the spatiotemporal data using the minimum enclosing rectangle and the characteristic value;
[0024] A reduction parallelism determination module is used to extract the data to be analyzed from the data block in parallel using the index structure, sample the data to be analyzed to obtain the keyword frequency, calculate the data skewness under different reduction parallelisms using the keyword frequency, use the reduction parallelism with the smallest data skewness as the parallelism of the reduction stage, establish a corresponding relationship between each reduction task and the keyword, and use the keyword with the largest keyword probability among the keywords corresponding to the reduction task as the label of the reduction task;
[0025] A scheduling module, used to partition the data to be analyzed according to the parallelism and keyword frequency and obtain target mapping tasks; and schedule the reduction tasks and target mapping tasks with the same label to the same node;
[0026] The parallel analysis module is used to perform parallel processing and analysis on the data to be analyzed by using the parallelism and the data sharding.
[0027] Preferably, the characteristic value of the spatiotemporal data is calculated based on the spatial information and time information of the spatiotemporal data, specifically:
[0028] Obtain the spatial information of the spatiotemporal data, and convert each dimension of the spatial information of all the spatiotemporal data in the data block into binary data of the same length; for each spatiotemporal data, starting from the highest bit, alternately take one bit from the binary data of each dimension to form a spatial binary;
[0029] Encode the time information of spatiotemporal data into time bins of the same length;
[0030] The spatial binary and temporal binary are concatenated, and the concatenated binary is converted into decimal as the characteristic value of the spatiotemporal data.
[0031] Preferably, the index structure of spatiotemporal data is constructed by using the minimum enclosing rectangle and the eigenvalue, specifically:
[0032] Determine the maximum and minimum number of storage entries based on the disk page and spatiotemporal data size; sort the spatiotemporal data according to the characteristic value to obtain the spatiotemporal data sequence;
[0033] Add the spatiotemporal data whose sequence number in the spatiotemporal data sequence is greater than the minimum number of stored entries and not greater than the maximum number of stored entries to the set, add the spatiotemporal data in the set whose spatial binary difference with the previous spatiotemporal data is not zero to the subset, sequentially extract the spatiotemporal data from the subset, calculate the minimum enclosing rectangle of the spatiotemporal data before the extracted spatiotemporal data in the spatiotemporal data sequence, and take the smallest of all the minimum enclosing rectangles of the subset as the leaf of the index structure; delete the leaf object from the spatiotemporal data sequence, and after clearing the set and the subset, repeat the leaf acquisition process until the spatiotemporal data sequence is empty;
[0034] The index structure is constructed based on the leaves in a recursive manner.
[0035] Preferably, the method of calculating the degree of data skewness under different reduction parallelisms by using the keyword frequency is specifically as follows:
[0036] For each preset reduction parallelism, the keywords are combined into combinations of the number of reduction parallelisms so that the sum of the keyword probabilities of each combination is closest, the sum of the probabilities of the keywords in the combination is taken as the probability of the combination, and the variance or standard deviation of the probabilities of all combinations is calculated to obtain the degree of data skew corresponding to the reduction parallelism.
[0037] Preferably, the data to be analyzed is segmented according to the parallelism and the keyword frequency and the target mapping task is obtained, specifically:
[0038] Sort the keywords in descending order of their frequency;
[0039] Associating the keywords of the number of reduction tasks mentioned above with a target mapping task respectively, and using the associated keywords as labels of the target mapping task;
[0040] When sharding, the data to be analyzed that contains the target mapping task label is preferentially allocated to the shard corresponding to the target mapping task.
[0041] The present invention improves the speed of extracting data to be analyzed from large-scale GIS spatial data by improving the spatiotemporal data index; and improves the data tilt and data transmission in parallel computing, thereby improving the speed of parallel computing. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 This is a flow chart of Embodiment 1;
[0043] Figure 2 A schematic diagram of the relationship between eigenvalue and position;
[0044] Figure 3 A schematic diagram of an index structure;
[0045] Figure 4 A schematic diagram of keyword combination;
[0046] Figure 5 A schematic diagram of the relationship between mapping tasks, reduction tasks, and slicing;
[0047] Figure 6 This is a structural diagram of the second embodiment. DETAILED DESCRIPTION
[0048] In this article, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0050] Figure 1A first embodiment of the present invention is shown. Figure 1 The GIS spatial large-scale data parallel processing and analysis method shown in includes the following steps:
[0051] S1, the large-scale GIS spatial data to be analyzed is divided into data blocks according to time and space, the characteristic values of the spatiotemporal data are calculated according to the spatial information and time information of the spatiotemporal data, and the index structure of the spatiotemporal data is constructed using the minimum enclosing rectangle and the characteristic value;
[0052] GIS space contains a large amount of data, but these data have spatiotemporal information including space and time, where spatial information is geographic location information, such as coordinates, regions, addresses, etc., and time information is the time when data is generated or collected, or when an event occurs. For example, GPS trajectory data of vehicles and pedestrians, meteorological data that records meteorological elements such as temperature, precipitation, wind force at different time points, and traffic flow data that records the number and speed of vehicles on the road at different time points. A typical spatiotemporal data is [id, (x, y), time, others], where id is the sequence number or unique identifier of spatiotemporal data, (x, y) is the coordinate such as longitude and latitude, time is time, and others is other information of spatiotemporal data, such as temperature, license plate number, etc. Processing and analyzing large-scale GIS spatial data can help people make some decisions, such as analyzing the path and intensity changes of typhoons, predicting their landing time and impact range, and issuing warning information, analyzing which sections of roads are prone to congestion in a specific time period, how long the congestion lasts, and predicting future traffic flow, etc.
[0053] GIS spatial large-scale data contains data from different locations and times. The amount of data is large. Generally, the data is stored according to the spatiotemporal information. The segmented data blocks can be stored and managed independently. Use grid partitioning, quadtree, KD tree and other methods to divide the geographic space into multiple regions, and divide the time series data into multiple time periods according to time intervals such as hours, days, months, etc., so that each data block corresponds to a certain area and time. For example, a city is divided into multiple grids, and the data generated by each grid every day is regarded as a data block.
[0054] Even after being divided into data blocks, each data block still has a lot of spatiotemporal data. For example, a data block is 128MB. If a spatiotemporal data is 100 bytes, then a data block can store about a million spatiotemporal data. This requires the use of an index to quickly determine the spatiotemporal data to be analyzed. In one embodiment, the feature value of the spatiotemporal data is calculated based on the spatial information and time information of the spatiotemporal data, specifically:
[0055] Obtain the spatial information of the spatiotemporal data, and convert each dimension of the spatial information of all the spatiotemporal data in the data block into binary data of the same length; for each spatiotemporal data, starting from the highest bit, alternately take one bit from the binary data of each dimension to form a spatial binary;
[0056] Encode the time information of spatiotemporal data into time bins of the same length;
[0057] The spatial binary and temporal binary are concatenated, and the concatenated binary is converted into decimal as the characteristic value of the spatiotemporal data.
[0058] Each spatiotemporal data has spatial information and time information. The spatial position is two-dimensional or three-dimensional. The maximum value of each dimension of each spatial position is taken out, and the maximum value of all dimensions of all spatial positions of all spatiotemporal data in the data block is taken as the maximum value of the space. The maximum value of the space is converted into binary, and the number of binary bits is obtained and used as the number of binary bits of the space. For example, if the data block has three spatiotemporal data, whose spatial coordinates are (1, 2), (5, 3), and (9, 4), the maximum value of the spatial position is 9, and the binary of 9 is 1001, then the number of binary bits of the space is 4.
[0059] Then, the spatial data of all spatiotemporal data in the data block are converted into binary with the same number of bits of the spatial binary, for example, (1, 2) and (5, 3) are converted into (0001, 0010) and (0101, 0011). For each spatiotemporal data, after the spatial position is converted into binary, starting from the highest bit of the binary, one bit is alternately taken from the binary of each dimension to merge into the spatial binary of this spatiotemporal data. At the initial moment, the spatial binary is empty, the highest bit value is taken from the binary of the first dimension and added to the end of the spatial binary, the highest bit value is taken from the binary of the second dimension and added to the end of the spatial binary, if there are still dimensions, continue to add, if there are no dimensions, proceed to the next round; in the next round, the second highest bit value is taken from the binary of the first dimension and added to the end of the spatial binary, the second highest bit value is taken from the binary of the second dimension and added to the end of the spatial binary, if there are still dimensions, continue to add, if there are no dimensions, proceed to the next round; iterate continuously until all binary bits of all dimensions are added to the end of the spatial binary to obtain the spatial binary. For example (01, 10), the highest bit of the first dimension 01 is 0, the spatial binary becomes 0, the highest bit of the second dimension 10 is 1, and the spatial binary becomes 01; then the second highest bit of the first dimension 01 is 1, the spatial binary becomes 011, the second highest bit of the second dimension 10 is 0, and the spatial binary becomes 0110. Figure 2 The eigenvalues calculated based on the spatial binary before adding the time binary are shown. Figure 2 It can be seen that the closer the spatial positions are, the closer the eigenvalues are.
[0060] The time information in the spatiotemporal data is encoded into time binaries of the same length. The later the time is or the closer the time is to the current moment, the larger the time binaries are. Then the space binaries and time binaries are concatenated. Specifically, the time binaries are added to the space binaries. For example, if the space binaries are 0110 and the time binaries are 010, the concatenated binaries are 0110010. The concatenated binaries are converted to decimal to obtain the eigenvalue. The eigenvalue corresponding to the concatenated binaries is 50. The closer the spatial position and time are, the closer the eigenvalues are.
[0061] After obtaining the eigenvalue of each spatiotemporal data in the data block, the index structure of the spatiotemporal data in the data block is constructed using the minimum enclosing rectangle and the eigenvalue. Specifically:
[0062] Determine the maximum and minimum number of storage entries based on the disk page and spatiotemporal data size; sort the spatiotemporal data according to the characteristic value to obtain the spatiotemporal data sequence;
[0063] Add the spatiotemporal data whose sequence number in the spatiotemporal data sequence is greater than the minimum number of stored entries and not greater than the maximum number of stored entries to the set, add the spatiotemporal data in the set whose spatial binary difference with the previous spatiotemporal data is not zero to the subset, sequentially extract the spatiotemporal data from the subset, calculate the minimum enclosing rectangle of the spatiotemporal data before the extracted spatiotemporal data in the spatiotemporal data sequence, and take the smallest of all the minimum enclosing rectangles of the subset as the leaf of the index structure; delete the leaf object from the spatiotemporal data sequence, and after clearing the set and the subset, repeat the leaf acquisition process until the spatiotemporal data sequence is empty;
[0064] The index structure is constructed based on the leaves in a recursive manner.
[0065] Disk reading is done in pages, that is, at least one page is read during reading. In order to improve disk reading efficiency, the maximum and minimum number of data that can be stored in each R-tree node are determined according to the size of the disk page and the size of the spatiotemporal data. According to the characteristic values of the spatiotemporal data, all spatiotemporal data are sorted to obtain an ordered spatiotemporal data sequence. Similar characteristic values represent spatiotemporal data with similar positions and times. From the sorted spatiotemporal data sequence, a group of data is selected and put into the set. The selection method is to put the spatiotemporal data with sequence numbers in the spatiotemporal data sequence between (minimum number of stored entries, maximum number of stored entries] as a selected group of data into the set, where the sequence number of the spatiotemporal data sequence starts from 1, and the order of the spatiotemporal data in the set remains unchanged. For example, if the minimum number of stored entries is 10 and the maximum number of stored entries is 20, then the 11th to 20th data in the spatiotemporal data sequence are put into the set.
[0066] For each spatiotemporal data in the set, calculate the spatial binary difference with the previous spatiotemporal data in the spatiotemporal data sequence. If the difference is not 0, put this spatiotemporal data into the subset. For each spatiotemporal data in the subset, calculate the minimum bounding rectangle (MBR) of the spatiotemporal data before the spatiotemporal data in the spatiotemporal data sequence. In this way, each spatiotemporal data in the subset will calculate a minimum bounding rectangle, and determine the minimum value of the minimum bounding rectangle corresponding to all spatiotemporal data in the subset. For example, there are 4 minimum bounding rectangles, and the smallest one is found from these 4. Each minimum bounding rectangle contains at least one spatiotemporal data. Then the spatiotemporal data contained in this minimum is taken as a whole as a leaf of the index structure, and the object in the leaf is the spatiotemporal data corresponding to this minimum. The smaller the minimum bounding rectangle, the closer the space and time of the data in the leaf are, which reduces the MBR overlap on the one hand and the search range on the other hand. For example, to count the number of traffic lights newly added in a certain range from 2023 to 2024, when searching in the data block, all MBRs that overlap with the certain range will be searched, and if multiple MBRs overlap, they may be searched repeatedly.
[0067] After obtaining a leaf and its object, delete the leaf object from the spatiotemporal data sequence to obtain a new spatiotemporal data sequence, then clear the set and subset, and use the new spatiotemporal data sequence to obtain the next leaf, and continue until all spatiotemporal data are assigned to a leaf.
[0068] If the difference of all spatiotemporal data in the collection is 0, the spatiotemporal data between the first and the maximum number of stored entries in the spatiotemporal data sequence will be taken as a leaf object.
[0069] After all leaves are obtained, the index structure is obtained by iteration. Specifically, if there is only one leaf in the input leaf node set, the node is the root node of the index structure tree. After the construction is completed, the root node is returned. If there are multiple leaf nodes in the input leaf node set, a new parent node needs to be created, and the leaf nodes in the leaf node set are grouped. The number of leaf nodes in each group does not exceed a preset number, such as 4, 6, etc. For each group of nodes, a new parent node is created, and the minimum enclosing rectangle of the parent node is calculated. The MBR of the parent node is the union of the MBRs of all its child nodes, and the child nodes are added to the parent node. The newly created parent node set is used as input, and the function of building the index structure is recursively called, and the execution is repeated until only one root node is left. Among them, the index structure preferably adopts R-tree. Figure 3 A schematic diagram of an index structure is shown, where MBR is the minimum bounding rectangle, and R11, R12, etc. are spatiotemporal data.
[0070] S2, extracting the data to be analyzed from the data block in parallel using the index structure, sampling the data to be analyzed to obtain the keyword frequency, using the keyword frequency to calculate the data skewness under different reduction parallelisms, taking the reduction parallelism with the smallest data skewness as the parallelism of the reduction stage, establishing the corresponding relationship between each reduction task and the keyword, and taking the keyword with the highest keyword probability among the keywords corresponding to the reduction task as the label of the reduction task;
[0071] After obtaining the index structure of each data block, when analyzing the data, the data to be analyzed is extracted from the data block according to the command or statement. For example, to count the types and numbers of new street trees in City A from 2023 to 2024, all spatiotemporal data containing new street trees in City A from 2023 to 2024 will be extracted from the data block.
[0072] In order to prevent data skew in the reduction phase in distributed computing frameworks such as Hadoop and Spark, the keyword frequency of the data to be analyzed is sampled and counted. The sampling methods include but are not limited to random sampling, equal interval sampling, hierarchical sampling, etc. Keywords such as sycamore, pagoda tree, camphor, etc., the frequency of each keyword is counted, and the keyword frequency is used to calculate the degree of data skew under different reduction parallelism. In one embodiment, reduction refers to the reduce phase in mapreduce, and different parallelisms correspond to different numbers of reduce tasks. In one embodiment, the degree of data skewness under different reduction parallelism is calculated using the keyword frequency, specifically:
[0073] For each preset reduction parallelism, the keywords are combined into combinations of the number of reduction parallelisms so that the sum of the keyword probabilities of each combination is closest, the sum of the probabilities of the keywords in the combination is taken as the probability of the combination, and the variance or standard deviation of the probabilities of all combinations is calculated to obtain the degree of data skew corresponding to the reduction parallelism.
[0074] Preset several reduction parallelisms, for example, the preset reduction parallelism is 4, 8, 12, etc. For each reduction parallelism, combine the keywords, and the number of combinations is the preset reduction parallelism. At the same time, it is required that the sum of the keyword probabilities of each combination is closest. Here, greedy algorithms or backtracking search methods can be used to determine the keyword combination method corresponding to each reduction parallelism. For example, a combination method such as Figure 4 As shown, Figure 4In the example, keywords 1 and 2 are taken as a combination, and keyword 3 is taken as a combination alone. Each reduction parallelism corresponds to a group of combinations. For this group of combinations, the sum of the probability of keywords in each combination is calculated, and the variance or standard deviation of all combinations is calculated. The standard deviation or variance is used as the data skewness of the reduction parallelism. Then, the reduction parallelism with the smallest data skewness is used as the parallelism of the reduction stage. At this time, the data skewness is the smallest. For a group of combinations with the smallest skewness, a corresponding relationship between the combination and the reduction task is established. For example, if the reduction parallelism is 2, this group of combinations includes combinations of 1, 2, and 3. Then, keywords 1 and 2 are established with the first reduction task, and keyword 3 is established with the second reduction task. The probabilities corresponding to keywords 1-3 are 0.6, 0.3, and 0.1, respectively. Then the label of the first reduction task is keyword 1, and the label of the second reduction task is keyword 3. When performing reduction, the mapping task only maps the keywords in the combination corresponding to this mapping task. For example, in the mapping stage, all the data of keywords 1 and 2 will be reduced by the first reduction task, and all the data of keyword 3 will be reduced by the second reduction task.
[0075] S3, sharding the data to be analyzed according to the parallelism and keyword frequency and obtaining target mapping tasks; scheduling the reduction tasks and target mapping tasks with the same label to the same node;
[0076] MapReduce includes a mapping phase and a reduction phase. After the parallelism of the reduction phase is determined, the reduction task is also determined. However, when the mapping task and the reduction task are not on the same node, the data of the mapping task needs to be sent to the node of the reduction task through the network, which often causes a bottleneck for parallel processing. Generally speaking, the number of mapping tasks, that is, the parallelism, is greater than that of the reduction phase. When the data is segmented according to the parallelism and keyword frequency and the target mapping task is obtained, and the spatiotemporal data mainly processed by the target mapping task is the same as the keyword mainly reduced by the reduction task, network data transmission can be reduced. In one embodiment, the data segmentation according to the parallelism and keyword frequency and the target mapping task is obtained are specifically:
[0077] Sort the keywords in descending order of their frequency;
[0078] Associating the keywords of the number of reduction tasks mentioned above with a target mapping task respectively, and using the associated keywords as labels of the target mapping task;
[0079] When sharding, the data to be analyzed that contains the target mapping task label is preferentially allocated to the shard corresponding to the target mapping task.
[0080] Sort the keywords in descending order of keyword frequency. The larger the keyword ranking, the higher the proportion of these keywords in the spatiotemporal data to be analyzed, and the more such data should be avoided from being transmitted between different nodes. Select the target mapping task from all mapping tasks. The number of target mapping tasks is the same as the number of reduction tasks, and the number of shards is the same as the number of mapping tasks. Associate the keyword of the number of parallelism mentioned above, that is, the number of reduction tasks, with a target mapping task, and use this keyword as the label of the target mapping task. When sharding, the data to be analyzed containing the target mapping task label is preferentially allocated to the shard corresponding to the target mapping task. Among them, the size of each shard is the same. For example, if there are 10 mapping tasks and 1,000 data to be analyzed, each shard contains 100 data to be analyzed. When sharding, find the label of each target mapping task. If the keyword of a data to be analyzed is the same as the label of this target mapping task, the data to be analyzed is preferentially allocated to the shard corresponding to this target mapping task. If it exceeds the maximum number of shards, it is allocated to the shards corresponding to other mapping tasks. The number of target mapping tasks and the number of reduction tasks are the same, both of which are the parallelism. When scheduling tasks, the target mapping tasks and reduction tasks with the same label are scheduled to one node, which avoids a large amount of data being transmitted through the network. Figure 5 As shown, map tasks and reduce tasks with the same label will be scheduled to the same node.
[0081] S4, using the parallelism and the data sharding to perform parallel processing and analysis on the data to be analyzed.
[0082] After obtaining the degree of parallelism and data partitioning, we also get the relationship between the partitioning and the mapping task and the reduction task, and use a distributed parallel computing framework such as MapReduce to perform parallel processing and analysis on the data.
[0083] like Figure 6 The second embodiment of the present invention provides a GIS spatial large-scale data parallel processing and analysis system, the system comprising the following modules:
[0084] The index module is used to divide the large-scale GIS spatial data to be analyzed into data blocks according to time and space, calculate the characteristic value of the spatiotemporal data according to the spatial information and time information of the spatiotemporal data, and construct the index structure of the spatiotemporal data using the minimum enclosing rectangle and the characteristic value;
[0085] A reduction parallelism determination module is used to extract the data to be analyzed from the data block in parallel using the index structure, sample the data to be analyzed to obtain the keyword frequency, calculate the data skewness under different reduction parallelisms using the keyword frequency, use the reduction parallelism with the smallest data skewness as the parallelism of the reduction stage, establish a corresponding relationship between each reduction task and the keyword, and use the keyword with the largest keyword probability among the keywords corresponding to the reduction task as the label of the reduction task;
[0086] A scheduling module, used to partition the data to be analyzed according to the parallelism and keyword frequency and obtain target mapping tasks; and schedule the reduction tasks and target mapping tasks with the same label to the same node;
[0087] The parallel analysis module is used to perform parallel processing and analysis on the data to be analyzed by using the parallelism and the data sharding.
[0088] Preferably, the characteristic value of the spatiotemporal data is calculated based on the spatial information and time information of the spatiotemporal data, specifically:
[0089] Obtain the spatial information of the spatiotemporal data, and convert each dimension of the spatial information of all the spatiotemporal data in the data block into binary data of the same length; for each spatiotemporal data, starting from the highest bit, alternately take one bit from the binary data of each dimension to form a spatial binary;
[0090] Encode the time information of spatiotemporal data into time bins of the same length;
[0091] The spatial binary and temporal binary are concatenated, and the concatenated binary is converted into decimal as the characteristic value of the spatiotemporal data.
[0092] Preferably, the index structure of spatiotemporal data is constructed by using the minimum enclosing rectangle and the eigenvalue, specifically:
[0093] Determine the maximum and minimum number of storage entries based on the disk page and spatiotemporal data size; sort the spatiotemporal data according to the characteristic value to obtain the spatiotemporal data sequence;
[0094] Add the spatiotemporal data whose sequence number in the spatiotemporal data sequence is greater than the minimum number of stored entries and not greater than the maximum number of stored entries to the set, add the spatiotemporal data in the set whose spatial binary difference with the previous spatiotemporal data is not zero to the subset, sequentially extract the spatiotemporal data from the subset, calculate the minimum enclosing rectangle of the spatiotemporal data before the extracted spatiotemporal data in the spatiotemporal data sequence, and take the smallest of all the minimum enclosing rectangles of the subset as the leaf of the index structure; delete the leaf object from the spatiotemporal data sequence, and after clearing the set and the subset, repeat the leaf acquisition process until the spatiotemporal data sequence is empty;
[0095] The index structure is constructed based on the leaves in a recursive manner.
[0096] Preferably, the method of calculating the degree of data skewness under different reduction parallelisms by using the keyword frequency is specifically as follows:
[0097] For each preset reduction parallelism, the keywords are combined into combinations of the number of reduction parallelisms so that the sum of the keyword probabilities of each combination is closest, the sum of the probabilities of the keywords in the combination is taken as the probability of the combination, and the variance or standard deviation of the probabilities of all combinations is calculated to obtain the degree of data skew corresponding to the reduction parallelism.
[0098] Preferably, the data to be analyzed is segmented according to the parallelism and the keyword frequency and the target mapping task is obtained, specifically:
[0099] Sort the keywords in descending order of their frequency;
[0100] Associating the keywords of the number of reduction tasks mentioned above with a target mapping task respectively, and using the associated keywords as labels of the target mapping task;
[0101] When sharding, the data to be analyzed that contains the target mapping task label is preferentially allocated to the shard corresponding to the target mapping task.
[0102] Through the description of the above implementation methods, technicians in this field can clearly understand that each implementation method can be implemented by adding a necessary general hardware platform, and of course can also be implemented by combining hardware and software. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a computer product, and the present invention can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it, and other embodiments may also be used. Although the present invention has been described in detail with reference to the aforementioned embodiments, ordinary technicians in this field should understand that they can still modify the technical solutions recorded in the aforementioned embodiments, or replace some of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A GIS spatial large-scale data parallel processing and analysis method, characterized in that: The method comprises the following steps: The large-scale GIS spatial data to be analyzed is divided into data blocks according to time and space, the characteristic values of the spatiotemporal data are calculated according to the spatial information and time information of the spatiotemporal data, and the index structure of the spatiotemporal data is constructed using the minimum enclosing rectangle and the characteristic value; The data to be analyzed is extracted from the data block in parallel using the index structure, the keyword frequency is obtained by sampling the data to be analyzed, the data skewness is calculated under different reduction parallelisms using the keyword frequency, the reduction parallelism with the smallest data skewness is used as the parallelism of the reduction stage, a corresponding relationship between each reduction task and a keyword is established, and the keyword with the largest keyword probability among the keywords corresponding to the reduction task is used as the label of the reduction task; The data to be analyzed is divided into slices according to the parallelism and the keyword frequency and a target mapping task is obtained; the reduction task and the target mapping task having the same label are scheduled to the same node; The parallelism and the data sharding are used to perform parallel processing and analysis on the data to be analyzed; The characteristic value of the spatiotemporal data is calculated based on the spatial information and time information of the spatiotemporal data, specifically: Obtain the spatial information of the spatiotemporal data, and convert each dimension of the spatial information of all the spatiotemporal data in the data block into binary data of the same length; for each spatiotemporal data, starting from the highest bit, alternately take one bit from the binary data of each dimension to form a spatial binary; Encode the time information of spatiotemporal data into time bins of the same length; The spatial binary and temporal binary are concatenated, and the concatenated binary is converted into decimal as the characteristic value of the spatiotemporal data.
2. The method according to claim 1, characterized in that The index structure of spatiotemporal data is constructed by using the minimum enclosing rectangle and the eigenvalue, specifically: Determine the maximum and minimum number of storage entries based on the disk page and spatiotemporal data size; sort the spatiotemporal data according to the characteristic value to obtain the spatiotemporal data sequence; Add the spatiotemporal data whose sequence number in the spatiotemporal data sequence is greater than the minimum number of stored entries and not greater than the maximum number of stored entries to the set, add the spatiotemporal data in the set whose spatial binary difference with the previous spatiotemporal data is not zero to the subset, sequentially extract the spatiotemporal data from the subset, calculate the minimum enclosing rectangle of the spatiotemporal data before the extracted spatiotemporal data in the spatiotemporal data sequence, and take the smallest of all the minimum enclosing rectangles of the subset as the leaf of the index structure; delete the leaf object from the spatiotemporal data sequence, and after clearing the set and the subset, repeat the leaf acquisition process until the spatiotemporal data sequence is empty; The index structure is constructed based on the leaves in a recursive manner.
3. The method according to claim 1, characterized in that The method of using the keyword frequency to calculate the degree of data skewness under different reduction parallelisms is specifically as follows: For each preset reduction parallelism, the keywords are combined into combinations of the number of reduction parallelisms so that the sum of the keyword probabilities of each combination is closest, the sum of the probabilities of the keywords in the combination is taken as the probability of the combination, and the variance or standard deviation of the probabilities of all combinations is calculated to obtain the degree of data skew corresponding to the reduction parallelism.
4. The method according to claim 1, characterized in that The data sharding to be analyzed and the target mapping task to be obtained according to the parallelism and the keyword frequency are specifically: Sort the keywords in descending order of their frequency; Associating the keywords of the number of reduction tasks mentioned above with a target mapping task respectively, and using the associated keywords as labels of the target mapping task; When sharding, the data to be analyzed that contains the target mapping task label is preferentially allocated to the shard corresponding to the target mapping task.
5. A GIS spatial large-scale data parallel processing and analysis system, characterized in that: The system includes the following modules: The index module is used to divide the large-scale GIS spatial data to be analyzed into data blocks according to time and space, calculate the characteristic value of the spatiotemporal data according to the spatial information and time information of the spatiotemporal data, and construct the index structure of the spatiotemporal data using the minimum enclosing rectangle and the characteristic value; A reduction parallelism determination module is used to extract the data to be analyzed from the data block in parallel using the index structure, sample the data to be analyzed to obtain the keyword frequency, calculate the data skewness under different reduction parallelisms using the keyword frequency, use the reduction parallelism with the smallest data skewness as the parallelism of the reduction stage, establish a corresponding relationship between each reduction task and the keyword, and use the keyword with the largest keyword probability among the keywords corresponding to the reduction task as the label of the reduction task; A scheduling module, used to partition the data to be analyzed according to the parallelism and the keyword frequency and obtain the target mapping task; and schedule the reduction task and the target mapping task with the same label to the same node; A parallel analysis module, used for performing parallel processing and analysis on the data to be analyzed by using the parallelism and the data sharding; The characteristic value of the spatiotemporal data is calculated based on the spatial information and time information of the spatiotemporal data, specifically: Obtain the spatial information of the spatiotemporal data, and convert each dimension of the spatial information of all the spatiotemporal data in the data block into binary data of the same length; for each spatiotemporal data, starting from the highest bit, alternately take one bit from the binary data of each dimension to form a spatial binary; Encode the time information of spatiotemporal data into time bins of the same length; The spatial binary and temporal binary are concatenated, and the concatenated binary is converted into decimal as the characteristic value of the spatiotemporal data.
6. The system according to claim 5, characterized in that The index structure of spatiotemporal data is constructed by using the minimum enclosing rectangle and the eigenvalue, specifically: Determine the maximum and minimum number of storage entries based on the disk page and spatiotemporal data size; sort the spatiotemporal data according to the characteristic value to obtain the spatiotemporal data sequence; Add the spatiotemporal data whose sequence number in the spatiotemporal data sequence is greater than the minimum number of stored entries and not greater than the maximum number of stored entries to the set, add the spatiotemporal data in the set whose spatial binary difference with the previous spatiotemporal data is not zero to the subset, sequentially extract the spatiotemporal data from the subset, calculate the minimum enclosing rectangle of the spatiotemporal data before the extracted spatiotemporal data in the spatiotemporal data sequence, and take the smallest of all the minimum enclosing rectangles of the subset as the leaf of the index structure; delete the leaf object from the spatiotemporal data sequence, and after clearing the set and the subset, repeat the leaf acquisition process until the spatiotemporal data sequence is empty; The index structure is constructed based on the leaves in a recursive manner.
7. The system according to claim 5, characterized in that The method of using the keyword frequency to calculate the degree of data skewness under different reduction parallelisms is specifically as follows: For each preset reduction parallelism, the keywords are combined into combinations of the number of reduction parallelisms so that the sum of the keyword probabilities of each combination is closest, the sum of the probabilities of the keywords in the combination is taken as the probability of the combination, and the variance or standard deviation of the probabilities of all combinations is calculated to obtain the degree of data skew corresponding to the reduction parallelism.
8. The system according to claim 5, characterized in that The data sharding to be analyzed and the target mapping task to be obtained according to the parallelism and the keyword frequency are specifically: Sort the keywords in descending order of their frequency; Associating the keywords of the number of reduction tasks mentioned above with a target mapping task respectively, and using the associated keywords as labels of the target mapping task; When sharding, the data to be analyzed that contains the target mapping task label is preferentially allocated to the shard corresponding to the target mapping task.
Citation Information
Patent Citations
Space-time data processing platform and method based on time and space data models
CN108959352A
Distributed spatio-temporal data indexing method based on R tree
CN111078634A
Spatial-temporal data double-index structure based on LR tree and B + tree
CN111859034A
Data skew greedy optimization method based on heterogeneous Spark
CN117453482A