A Distributed Storage Method and System for Massive InSAR Deformation Monitoring Data
Through distributed storage and calculation methods, InSAR deformation monitoring data files are divided into subtasks, and distributed tasks are distributed among cluster nodes using distributed task schedulers, which solves the complexity, reliability and scalability problems in the process of massive data entry, and realizes efficient and real-time data processing and entry.
Patent Information
- Application Number
- CN202510632368.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The prior art has problems with high complexity, poor reliability, low efficiency, inability to guarantee timeliness and inscalability when processing massive InSAR deformation monitoring data, resulting in complex data entry process, easy interruption, low efficiency and inability to scale.
Using distributed storage, distributed computing and dynamic scheduling methods, InSAR deformation monitoring data files are divided into independent subtasks, and task packages are allocated between cluster nodes through distributed task schedulers, parallel processing and repeated allocation of failed data blocks are realized, and spatiotemporal indexes are constructed for database processing.
It improves the efficiency of incorporating massive InSAR deformation monitoring data, ensures high efficiency, real-time and reliability of data processing, and provides elastic scaling capabilities, solving the scalability problem under the single-unit architecture.
Smart Images

Figure CN120144541B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of InSAR deformation monitoring data processing, and particularly to a method and system for distributed storage of massive InSAR deformation monitoring data. Background Art
[0002] The InSAR deformation monitoring technology has the characteristics of high precision, wide range, large time span, etc., and has now become a new means with great development potential in the field of surface deformation monitoring. With the application of the InSAR deformation monitoring technology, a large number of historical and real-time data files of deformation monitoring have been generated, reaching the TB or even PB level. The InSAR deformation monitoring data itself carries information such as spatial attributes, time attributes, and change amounts, and can be applied in multiple fields such as geological disaster monitoring and urban ground settlement monitoring.
[0003] If the conventional single-body architecture method is used to process massive InSAR deformation monitoring data files, the following problems will exist: (1) High complexity: Since the entire import logic is implemented in a single-body application, the application will become very complex, making it difficult to modify and maintain; (2) Poor reliability: During the data import process, if an abnormal interruption occurs, it is easy to cause the entire data import process to terminate; (3) Low data import efficiency: Due to limitations of physical storage, memory, etc., reading and processing ultra-large files can only be carried out serially, resulting in low overall import efficiency; (4) Data timeliness cannot be guaranteed: Due to the low data import efficiency, by the time the data is stored in the database, the best analysis time may have been missed, and it no longer has analysis value; (5) Non-scalability: As the amount of processed data accumulates, when the performance reaches the limit, effective expansion cannot be carried out. At the same time, if the processing volume decreases, it occupies a high amount of physical resources and cannot be released.
[0004] Therefore, how to efficiently store massive InSAR deformation monitoring data in the database is a technical problem that needs to be solved urgently. Summary of the Invention
[0005] The present invention provides a method and system for distributed storage of massive InSAR deformation monitoring data, aiming to solve at least one of the above technical problems.
[0006] To achieve the above object, the present invention provides a method for distributed storage of massive InSAR deformation monitoring data, including:
[0007] S1: Distributively store the original InSAR deformation monitoring data files, and construct a list of InSAR deformation monitoring data files; wherein, each file data block of each InSAR deformation monitoring data file is stored in the list of InSAR deformation monitoring data files.
[0008] S2: Take each InSAR deformation monitoring data file as the parent task, and take the file data blocks of each InSAR deformation monitoring data file as the subtasks of the parent task to form a task package queue;
[0009] S3: Use the distributed task scheduler to load the import node information in the cluster, extract the node status of each node and establish heartbeat monitoring between the distributed task scheduler and each node in the cluster, allocate task packages to the nodes with idle node status in the cluster, move the task packages confirmed by the nodes to the running task queue, and mark the relationship between the task packages and the nodes;
[0010] S4: Each node constructs a file data block import task after receiving the task package, traverses and reads the file data blocks in the task package in sequence, and parses the obtained InSAR deformation monitoring data for storage in the database;
[0011] S5: When a node detects that the file data block import fails, send the subtask corresponding to the file data block back to the distributed task scheduler, and drive the distributed task scheduler to put the subtask into the parent task to wait for the next round of task package allocation;
[0012] S6: When a node executes the import action of the file data block, send the real-time import process information of the task package to the distributed task scheduler, and drive the distributed task scheduler to move the task package allocated to the corresponding node from the running task queue to the task package queue according to the received real-time import process information and the heartbeat monitoring result;
[0013] S7: When the import of the file data blocks of each parent task is completed, determine the unimported file data blocks according to the received real-time import process information, parse the unimported file data blocks, and parse the obtained InSAR deformation monitoring data for storage in the database;
[0014] S8: Return to execute the task package allocation step until all the task packages in the task package queue are allocated, and generate an import task execution report according to the received real-time import information.
[0015] Optionally, step S1: Distribute and store the original InSAR deformation monitoring data files, and construct a list of InSAR deformation monitoring data files, specifically including:
[0016] S11: Use the distributed storage technology to store the original InSAR deformation monitoring data files;
[0017] S12: Read the list of InSAR deformation monitoring data files in the distributed storage, and based on the technical characteristics of the file data blocks in the distributed storage, obtain all the file data block characteristic data of each InSAR deformation monitoring data file;
[0018] Among them, the file data block feature data includes a block serial number, a starting offset, and a block length.
[0019] Optionally, step S2: Using each InSAR deformation monitoring data file as a parent task and each file data block of each InSAR deformation monitoring data file as a subtask of the parent task, a task package queue is formed, specifically including:
[0020] S21: Using each InSAR deformation monitoring data file as a parent task and each file data block of each InSAR deformation monitoring data file as a subtask of the parent task, an InSAR deformation monitoring task queue is constructed;
[0021] S22: Start the distributed task scheduler to load the InSAR deformation monitoring task queue, and pack all the subtasks according to a preset quantity to form a task package queue.
[0022] Optionally, step S3: Use the distributed task scheduler to load the import node information in the cluster, extract the node status of each node, and establish heartbeat monitoring between the distributed task scheduler and each node in the cluster, and allocate task packages to the nodes in the cluster with an idle node status, and move the task packages confirmed by the nodes to the running task queue, and mark the relationship between the task packages and the nodes, specifically including:
[0023] S31: Use the distributed task scheduler to load the import node information in the cluster, extract the node status of each node, and establish heartbeat monitoring between the distributed task scheduler and the nodes in the cluster; among them, the node status includes idle and occupied;
[0024] S32: The distributed task scheduler allocates task packages to the nodes in the cluster with an idle node status. When the number of task packages is less than the number of idle nodes, all task packages are allocated. When the number of task packages is more than the number of idle nodes, a corresponding part of the task packages equal to the number of idle nodes is allocated;
[0025] S33: After each node receives a task package, first synchronize its own status to occupied, and send a task package reception confirmation message to the distributed task scheduler. The distributed task scheduler moves the task packages confirmed by the reception to the running task queue, and marks the relationship between the task packages and the nodes, and puts the task packages that have not received confirmation back into the task package queue to wait for the next allocation.
[0026] Optionally, step S4: Each node constructs a file data block import task after receiving a task package, sequentially traverses and reads the file data blocks in the task package, and performs warehousing processing on the parsed InSAR deformation monitoring data, specifically including:
[0027] S41: After each node receives the confirmation, it parses the task package, and constructs a file data block import task for the number of subtasks in the task package in the form of a thread pool;
[0028] S42: Each file data block import task in the node sequentially traverses and reads the complete data blocks in the file data block, parses the key information including spatial features, temporal features, and deformation variable features, and constructs a spatio-temporal index for the feature data and then performs the warehousing process.
[0029] Optionally, step S42: Each file data block import task in the node sequentially traverses and reads the complete data blocks in the file data block, parses the key information including spatial features, temporal features, and deformation variable features, and constructs a spatio-temporal index for the feature data and then performs the warehousing process, specifically including:
[0030] S421: For incomplete data blocks, after the node completes the import of the complete information, it synchronizes the incomplete data blocks to the main thread;
[0031] S422: After the main thread of the node completes the import of all task packages, it synchronizes its own state to idle, and at the same time packs the incomplete data blocks and file data block subtasks and sends them as a whole to the distributed task scheduler.
[0032] Optionally, step S5: When the node detects that the file data block import fails, it sends the subtask corresponding to the file data block back to the distributed task scheduler, driving the distributed task scheduler to put the subtask into the parent task to wait for the next round of task package allocation, specifically including:
[0033] S51: If the import of a certain file block fails during the import process in the node, it sends the task block information back to the distributed task scheduler;
[0034] S52: For the file database subtasks that fail to be imported, the distributed task scheduler puts them back into the parent task for the next round of allocation, and the tasks that fail to be allocated and executed after reaching the target number of times will no longer be allocated and are put into the task queue of failed executions.
[0035] Optionally, step S6: When the node executes the import action of the file data block, it sends the real-time import process information of the task package to the distributed task scheduler, driving the distributed task scheduler to move the task packages that have been allocated to the corresponding nodes from the running task queue to the task package queue according to the received real-time import process information and heartbeat monitoring results, specifically including:
[0036] S61: When a node executes the import action of a file data block, send the real-time import process information of the task package to the distributed task scheduler, and the distributed task scheduler reorganizes and merges the import process information according to files and file data blocks; wherein, the real-time import process information includes import process status, logs, and progress information;
[0037] S62: During the import process, when it is detected that a node does not feedback the real-time import process information within the target duration or the heartbeat monitoring is abnormal, the tasks allocated to the corresponding node will be moved from the running task queue to the task package queue to wait for the next reallocation.
[0038] Optionally, step S7: When the import of the file data blocks of each parent task is completed, determine the unimported file data blocks according to the received real-time import process information, parse the unimported file data blocks, and perform warehousing processing on the InSAR deformation monitoring data obtained by the parsing, specifically including:
[0039] S71: After the import of the file data blocks of each parent task is completed, the distributed task scheduler sorts all the received incomplete data blocks based on the data block order, and estimates the first processing duration for merging several data blocks corresponding to each file data block considering the CPU processing speed and the second processing duration for warehousing processing of each file data block considering the IO port transmission speed according to the data volume size of several data blocks corresponding to each file data block and the predicted CPU occupancy ratio and IO port transmission speed of the task scheduler at each subsequent moment. Based on the first processing duration and the second processing duration, generate the best order control strategy for merging and warehousing all the incomplete data blocks according to the principle of parallel CPU processing and IO port transmission with the shortest time consumption, and use this best order control strategy to sequentially select the corresponding data blocks from the sorting result to merge into complete file data blocks;
[0040] Among them, the principle of parallel CPU processing and IO port transmission with the shortest time consumption is configured as: The first group of data blocks first performs the merge process, and each subsequent group of data blocks performs the merge process while the previous group of data blocks performs the warehousing process until all the data blocks complete the merge process and the warehousing process; and the sum of the longer duration between the second processing duration of the previous group of data blocks for warehousing after merging to obtain the file data block and the first processing duration of the next group of data blocks for merging plus the sum of the first processing duration of the first group of data blocks and the second processing duration of the last group of data blocks is the shortest;
[0041] S72: Sequentially traverse and read the complete data blocks in the file data blocks, parse the key information including spatial features, temporal features, and deformation variable features, construct a spatio-temporal index for the feature data, and then perform the warehousing process. Complete the data warehousing operation of the corresponding InSAR deformation monitoring data file, update the status of the parent task to completed, and at the same time move the corresponding data packet task to the completed task queue.
[0042] In addition, to achieve the above object, the present invention also provides a distributed warehousing system for a large amount of InSAR deformation monitoring data, including:
[0043] A construction module, configured to perform distributed storage on the original InSAR deformation monitoring data file and construct a list of InSAR deformation monitoring data files; wherein, the list of InSAR deformation monitoring data files stores several file data blocks of each InSAR deformation monitoring data file.
[0044] A task packet queue formation module, configured to use each InSAR deformation monitoring data file as a parent task and each file data block of each InSAR deformation monitoring data file as a sub-task of the parent task to form a task packet queue.
[0045] A task packet allocation module, configured to use a distributed task scheduler to load the import node information in the cluster, extract the node status of each node, establish heartbeat monitoring between the distributed task scheduler and each node in the cluster, allocate task packets to the nodes with an idle status in the cluster, move the task packets confirmed by the nodes to the running task queue, and mark the relationship between the task packets and the nodes.
[0046] A first warehousing module, configured to construct a file data block import task for each node after receiving the task packet, sequentially traverse and read the file data blocks in the task packet, and perform the warehousing process on the parsed InSAR deformation monitoring data obtained.
[0047] A sending module, configured to send the sub-task corresponding to the file data block back to the distributed task scheduler when the node detects that the file data block import fails, and drive the distributed task scheduler to put the sub-task into the parent task to wait for the next round of task packet allocation.
[0048] A moving module, configured to send the real-time import process information of the task packet to the distributed task scheduler when the node executes the import action of the file data block, and drive the distributed task scheduler to move the task packet allocated to the corresponding node from the running task queue to the task packet queue according to the received real-time import process information and the heartbeat monitoring result.
[0049] The second storage module is used to determine the unimported file data blocks according to the received real-time import process information when the file data blocks of each parent task are imported, parse the unimported file data blocks, and perform storage processing on the InSAR deformation monitoring data obtained by the parsing;
[0050] The repetition module is used to return to execute the task package allocation step until all the task packages in the task package queue are allocated, and generate an import task execution report according to the received real-time import information.
[0051] The beneficial effects of the present invention are as follows: A method and system for distributed storage of massive InSAR deformation monitoring data are proposed, including extraction and list construction of InSAR deformation monitoring data files for distributed storage, formation of a task package queue including parent tasks and sub-tasks, task package allocation of cluster nodes based on a distributed task scheduler, execution and storage processing of file data block import tasks, parsing and storage processing of imported failed file data blocks and unimported file data blocks, and repeated allocation of task packages. Thus, by adopting technical means such as distributed storage, distributed computing, dynamic scheduling, and parallel processing, based on the idea of divide and conquer, the InSAR deformation monitoring data storage task is divided into independent sub-tasks, and the physical cluster resources are reasonably scheduled and allocated based on the amount of data processed, making full use of limited physical resources to provide efficient, real-time, and reliable storage processing for massive InSAR deformation monitoring data. At the same time, based on the characteristics of the cluster, elastic scaling ability is provided, significantly improving the storage efficiency of massive InSAR deformation monitoring data. Description of the Drawings
[0052] Figure 1 It is a schematic flowchart of the method for distributed storage of massive InSAR deformation monitoring data of the present invention;
[0053] Figure 2 It is a schematic structural diagram of the system for distributed storage of massive InSAR deformation monitoring data of the present invention. Detailed Embodiment
[0054] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0055] An embodiment of the present invention provides a method for distributed storage of massive InSAR deformation monitoring data, referring to Figure 1 , Figure 1 It is a schematic flowchart of the method for distributed storage of massive InSAR deformation monitoring data in an embodiment of the present invention. The method for distributed storage of massive InSAR deformation monitoring data includes the following steps:
[0056] S1: Distribute and store the original InSAR deformation monitoring data files, and construct a list of InSAR deformation monitoring data files; among them, several file data blocks of each InSAR deformation monitoring data file are stored in the list of InSAR deformation monitoring data files.
[0057] S2: Take each InSAR deformation monitoring data file as a parent task, and take the file data blocks of each InSAR deformation monitoring data file as subtasks of the parent task to form a task package queue.
[0058] S3: Use the distributed task scheduler to load the import node information in the cluster, extract the node status of each node and establish heartbeat monitoring between the distributed task scheduler and each node in the cluster, allocate task packages to the nodes with idle node status in the cluster, move the task packages confirmed by the nodes to the running task queue, and mark the relationship between the task packages and the nodes.
[0059] S4: Each node constructs a file data block import task after receiving the task package, sequentially traverses and reads the file data blocks in the task package, and performs database storage processing on the parsed InSAR deformation monitoring data obtained.
[0060] S5: When a node detects that the file data block import fails, send the subtask corresponding to the file data block back to the distributed task scheduler, and drive the distributed task scheduler to put the subtask into the parent task to wait for the next round of task package allocation.
[0061] S6: When a node executes the import action of the file data block, send the real-time import process information of the task package to the distributed task scheduler, and drive the distributed task scheduler to move the task package already allocated to the corresponding node from the running task queue to the task package queue according to the received real-time import process information and the heartbeat monitoring result.
[0062] S7: When the import of the file data blocks of each parent task is completed, determine the unimported file data blocks according to the received real-time import process information, parse the unimported file data blocks, and perform database storage processing on the parsed InSAR deformation monitoring data obtained.
[0063] S8: Return to execute the task package allocation step until all the task packages in the task package queue are allocated, and generate an import task execution report according to the received real-time import information.
[0064] It should be noted that when the prior art processes a large amount of InSAR deformation monitoring data files using the conventional monolithic architecture method, the following problems will occur: (1) High complexity: Since the entire import logic is implemented in a single monolithic application, the application will become very complex, making it difficult to modify and maintain; (2) Poor reliability: During the data import process, if an abnormal interruption occurs, it is easy to cause the entire data import process to terminate; (3) Low data import efficiency: Due to limitations such as physical storage and memory, reading and processing extremely large files can only be done serially, resulting in low overall import efficiency; (4) Data timeliness cannot be guaranteed: Due to the low data import efficiency, the best analysis time may have been missed by the time the data is stored in the database, and it no longer has analytical value; (5) Lack of scalability: As the amount of processed data accumulates, when the performance reaches its limit, effective expansion cannot be carried out. At the same time, if the processing volume decreases, it occupies a high amount of physical resources and cannot be released.
[0065] In this embodiment, through steps such as the extraction and list construction of InSAR deformation monitoring data files with distributed storage, the formation of a task package queue including parent tasks and sub-tasks, the distribution of task packages to cluster nodes based on a distributed task scheduler, the execution and storage processing of file data block import tasks, the parsing and storage processing of imported failure file data blocks and unimported file data blocks, and the repeated distribution of task packages, using technical means such as distributed storage, distributed computing, dynamic scheduling, and parallel processing, based on the idea of divide and conquer, the InSAR deformation monitoring data storage task is divided into independent sub-tasks, and the physical cluster resources are reasonably scheduled and allocated based on the amount of processed data, making full use of limited physical resources to achieve efficient, real-time, and reliable storage processing of a large amount of InSAR deformation monitoring data. At the same time, based on the characteristics of the cluster, elastic scalability is provided, significantly improving the storage efficiency of a large amount of InSAR deformation monitoring data.
[0066] To more clearly explain the present invention, the following provides a specific example of the method for distributed storage of a large amount of InSAR deformation monitoring data of the present invention, which includes the following specific execution steps:
[0067] Step 1: Store the original InSAR deformation monitoring data files using distributed storage technology;
[0068] Step 2: Read the list of InSAR deformation monitoring data files in the distributed storage, and based on the technical characteristics of the file data blocks in the distributed storage, obtain all the file data block characteristic data (block serial number, starting offset, block length) of each InSAR deformation monitoring data file;
[0069] Step 3: Build a task queue based on the InSAR deformation monitoring data file list. Import each InSAR deformation monitoring data file as a parent task, and import the file data blocks of each InSAR deformation monitoring data file as subtasks within the parent task;
[0070] Step 4: Start the distributed task scheduler, load the task list, and pack all the subtasks into task packages according to a specific number (e.g., 10) to form a task package queue;
[0071] Step 5: The distributed task scheduler loads the import node information in the cluster, including node status (occupied, idle), and simultaneously establishes a heartbeat monitoring between the distributed task scheduler and the cluster nodes;
[0072] Step 6: The distributed task scheduler allocates task packages to the nodes in the cluster with an idle status. If the number of task packages is less than the number of idle nodes at this time, all task packages are allocated. If the number of task packages is more than the number of idle nodes at this time, only some task packages are allocated;
[0073] Step 7: After receiving a task package, the node first synchronizes its own status to occupied and replies with a task package acceptance confirmation message. The distributed task scheduler moves the task packages with confirmed acceptance to the running task queue, marks the relationship between the task packages and the nodes, and re - puts the task packages that have not received confirmation back into the task package queue for the next allocation;
[0074] Step 8: After receiving the confirmation, the node will parse the task package and construct file data block import tasks through a thread pool according to the number of subtasks in the task package;
[0075] Step 9: Each file data block import task in the node sequentially traverses and reads the complete data blocks in the file data block, and parses key information such as spatial features, temporal features, and deformation amount features. After constructing a spatio - temporal index for the feature data, it is stored in the database. For incomplete data blocks, after importing the complete information, the incomplete data blocks are synchronized to the main thread;
[0076] Step 10: After the node main thread completes the import of all task packages, it synchronizes its own status to idle, and at the same time packs the incomplete data blocks and file data block subtasks and sends them as a whole to the distributed task scheduler;
[0077] Step 11: If the import of a certain file block fails during the import process in the node, the task block information is sent back to the distributed task scheduler;
[0078] Step Twelve: For the file database subtasks that fail to be imported, the distributed task scheduler will put them back into the parent task for the next round of allocation. Tasks that fail to be executed after multiple allocations will no longer be allocated and will be put into the task queue of failed executions.
[0079] Step Thirteen: When a node executes the import task of a file data block, it will simultaneously synchronize the import process status, log, and progress information of the task package to the distributed task scheduler, and the scheduler will reorganize and merge the import process information according to files and file data blocks.
[0080] Step Fourteen: If a node fails to provide feedback on the process information for a long time or has abnormal heartbeat monitoring during the import process, the tasks allocated to the corresponding node will be moved from the running task queue to the task queue for the next reallocation.
[0081] Step Fifteen: After the scheduling of the import parent task for a file is completed, the distributed task scheduler will merge all the incomplete information of the file data blocks received based on the order of the data blocks to form complete file data blocks, sequentially traverse and read the complete data blocks in the file data blocks, and parse the key information such as spatial features, temporal features, and deformation variable features. After constructing a spatio-temporal index for the feature data, it will be stored in the database. At this time, the data entry operation of the corresponding InSAR deformation monitoring data file is completed, the status of the parent task is updated to completed, and the corresponding data packet task is moved to the completed task queue.
[0082] It should be noted that the distributed task scheduler merges all the incomplete information of the received file data blocks based on the order of the data blocks to form complete file data block steps, which specifically include: after the import of the file data blocks of each parent task is completed, the distributed task scheduler sorts all the received incomplete data blocks based on the data block order, and estimates the first processing duration for merging several data blocks corresponding to each file data block considering the CPU processing speed and the second processing duration for storing each file data block considering the transmission speed of the IO port according to the data volume of several data blocks corresponding to each file data block and the predicted value of the CPU occupancy ratio at each subsequent moment of the task scheduler and the transmission speed of the IO port. Based on the first processing duration and the second processing duration, in accordance with the principle of parallel CPU processing and IO port transmission with the shortest time consumption, an optimal order control strategy for the merging and storage of all incomplete data blocks is generated, and using this optimal order control strategy, the corresponding data blocks are sequentially selected from the sorting result to be merged into complete file data blocks; wherein, the principle of parallel CPU processing and IO port transmission with the shortest time consumption is configured as: the first group of data blocks first performs the merging process, and each subsequent group of data blocks performs the merging process while the previous group of data blocks performs the storage process until all data blocks complete the merging process and the storage process; and the sum of the longer duration between the second processing duration for the previous group of data blocks to perform the storage process after merging to obtain the file data block and the first processing duration for the next group of data blocks to perform the merging process for all adjacent two groups of data blocks plus the sum of the first processing duration of the first group of data blocks and the second processing duration of the last group of data blocks is the shortest.
[0083] Thus, by adopting the principle of parallel CPU processing and IO port transmission with the shortest time consumption, an optimal order control strategy for the merging and storage of all incomplete data blocks is generated, and corresponding data blocks are sequentially selected from the sorting result to be merged into complete file data blocks. In scenarios with strict requirements for the import and storage time control of a large number of incomplete data blocks in InSAR deformation monitoring data, the data import and storage efficiency can be significantly improved, and the upper limit of the data processing capacity of the device can be increased.
[0084] Step Sixteen: Continuously repeat Step Six until all task queues are allocated, and perform file grouping and merging on the collected import process monitoring data to form an import task execution report.
[0085] Refer to Figure 2 , Figure 2 , which is a schematic structural diagram of the distributed storage system for a large amount of InSAR deformation monitoring data according to an embodiment of the present invention. The distributed storage system for a large amount of InSAR deformation monitoring data includes:
[0086] Building module 10 is used to store the original InSAR deformation monitoring data files in a distributed manner and construct a list of InSAR deformation monitoring data files. Among them, several file data blocks of each InSAR deformation monitoring data file are stored in the list of InSAR deformation monitoring data files.
[0087] Task package queue formation module 20 is used to form a task package queue with each InSAR deformation monitoring data file as the parent task and each file data block of each InSAR deformation monitoring data file as the subtask of the parent task.
[0088] Task package allocation module 30 is used to load the import node information in the cluster by using a distributed task scheduler, extract the node status of each node and establish heartbeat monitoring between the distributed task scheduler and each node in the cluster, allocate task packages to the nodes with an idle node status in the cluster, move the task packages confirmed by the nodes to the running task queue, and mark the relationship between the task packages and the nodes.
[0089] First warehousing module 40 is used for each node to construct a file data block import task after receiving a task package, sequentially traverse and read the file data blocks in the task package, and perform warehousing processing on the parsed InSAR deformation monitoring data obtained.
[0090] Sending module 50 is used to send the subtask corresponding to the file data block back to the distributed task scheduler when the node detects that the file data block import fails, driving the distributed task scheduler to put the subtask into the parent task to wait for the next round of task package allocation.
[0091] Moving module 60 is used to send the real-time import process information of the task package to the distributed task scheduler when the node executes the import action of the file data block, driving the distributed task scheduler to move the task package allocated to the corresponding node from the running task queue to the task package queue according to the received real-time import process information and the heartbeat monitoring result.
[0092] Second warehousing module 70 is used to determine the unimported file data blocks according to the received real-time import process information when the import of the file data blocks of each parent task is completed, parse the unimported file data blocks, and perform warehousing processing on the parsed InSAR deformation monitoring data obtained.
[0093] Repeating module 80 is used to return to execute the task package allocation step until all the task packages in the task package queue are allocated, and generate an import task execution report according to the received real-time import information.
[0094] Other embodiments or specific implementation manners of the distributed warehousing system for a large amount of InSAR deformation monitoring data of the present invention can refer to the above method embodiments, and will not be elaborated here.
[0095] It should be understood that in the description of this specification, the descriptions referring to terms such as "one embodiment", "another embodiment", "other embodiments", or "the first embodiment to the Nth embodiment", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0096] It should be noted that in this article, the term "comprising", "including", or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article, or system including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or system. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or system including the element.
[0097] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A distributed data storage method for a large amount of InSAR deformation monitoring data, characterized in that Including: S1: Distribute and store the original InSAR deformation monitoring data files to build a list of InSAR deformation monitoring data files. Among them, several file data blocks of each InSAR deformation monitoring data file are stored in the list of InSAR deformation monitoring data files. S2: Use each InSAR deformation monitoring data file as a parent task, and use the file data blocks of each InSAR deformation monitoring data file as subtasks of the parent task to form a task package queue. S3: Use the distributed task scheduler to load the import node information in the cluster, extract the node status of each node and establish heartbeat monitoring between the distributed task scheduler and each node in the cluster, allocate task packages to the nodes with idle node status in the cluster, move the task packages confirmed by the nodes to the running task queue, and mark the relationship between the task packages and the nodes. S4: Each node constructs a file data block import task after receiving the task package, sequentially traverses and reads the file data blocks in the task package, and performs warehousing processing on the parsed InSAR deformation monitoring data, specifically including: S41: After each node receives and confirms, it parses the task package, and constructs a file data block import task in the form of a thread pool for the number of subtasks in the task package. S42: Each file data block import task in the node sequentially traverses and reads the complete data blocks in the file data blocks, parses the key information including spatial features, time features, and deformation amount features, and performs warehousing processing after constructing a spatio-temporal index for the feature data, specifically including: S421: For incomplete data blocks, after the node completes the import of the complete information, it synchronizes the incomplete data blocks to the main thread. S422: After the main thread of the node completes the import of all task packages, it synchronizes its own status to idle, and at the same time packs the incomplete data blocks and file data block subtasks and sends them to the distributed task scheduler as a whole. S5: When a node detects that the file data block import fails, it sends the subtask corresponding to the file data block back to the distributed task scheduler, driving the distributed task scheduler to put the subtask into the parent task to wait for the next round of task package allocation. S6: When a node executes the import action of the file data block, it sends the real-time import process information of the task package to the distributed task scheduler, driving the distributed task scheduler to move the task packages already allocated to the corresponding nodes from the running task queue to the task package queue according to the received real-time import process information and heartbeat monitoring results. S7: When the import of the file data blocks of each parent task is completed, determine the unimported file data blocks according to the received real-time import process information, parse the unimported file data blocks, and perform warehousing processing on the parsed InSAR deformation monitoring data. S8: Return to the task package allocation step until all the task packages in the task package queue are allocated, and generate an import task execution report according to the received real-time import information.
2. The method for distributed storage of massive InSAR deformation monitoring data according to claim 1, characterized in that Step S1: Store the original InSAR deformation monitoring data files in a distributed manner, and construct a list of InSAR deformation monitoring data files, specifically including: S11: Store the original InSAR deformation monitoring data files using distributed storage technology; S12: Read the list of InSAR deformation monitoring data files in the distributed storage, and based on the technical characteristics of the file data blocks in the distributed storage, obtain all the file data block feature data of each InSAR deformation monitoring data file; Among them, the file data block feature data includes block serial number, starting offset, and block length.
3. The method for distributed storage of massive InSAR deformation monitoring data according to claim 1, characterized in that, Step S2: Use each InSAR deformation monitoring data file as a parent task, and use the file data blocks of each InSAR deformation monitoring data file as subtasks of the parent task to form a task package queue, specifically including: S21: Use each InSAR deformation monitoring data file as a parent task, and use the file data blocks of each InSAR deformation monitoring data file as subtasks of the parent task to construct an InSAR deformation monitoring task queue; S22: Start the distributed task scheduler to load the InSAR deformation monitoring task queue, and pack all the subtasks according to a preset quantity to form a task package queue.
4. The method for distributed storage of massive InSAR deformation monitoring data according to claim 1, wherein Step S3: Use the distributed task scheduler to load the import node information in the cluster, extract the node status of each node and establish heartbeat monitoring between the distributed task scheduler and each node in the cluster, allocate task packages to the nodes with an idle status in the cluster, move the task packages confirmed by the nodes to the running task queue, and mark the relationship between the task packages and the nodes, specifically including: S31: Use the distributed task scheduler to load the import node information in the cluster, extract the node status of each node, and establish heartbeat monitoring between the distributed task scheduler and the nodes in the cluster; among them, the node status includes idle and occupied; S32: The distributed task scheduler allocates task packages to the nodes with an idle status in the cluster. When the number of task packages is less than the number of idle nodes, all task packages are allocated. When the number of task packages is more than the number of idle nodes, a corresponding part of the task packages equal to the number of idle nodes is allocated; S33: After each node receives a task package, first synchronize its own status to occupied, and send a task package reception confirmation message to the distributed task scheduler. The distributed task scheduler moves the task packages confirmed to be received to the running task queue, and marks the relationship between the task packages and the nodes, and puts the task packages that have not received confirmation back into the task package queue to wait for the next allocation.
5. The method for distributed storage of massive InSAR deformation monitoring data according to claim 1, characterized in that Step S5: When a node detects that the import of a file data block fails, send the subtask corresponding to the file data block back to the distributed task scheduler, driving the distributed task scheduler to put the subtask into the parent task to wait for the next round of task package allocation, specifically including: S51: If the import of a certain file block fails during the import process by a node, send the task block information back to the distributed task scheduler; S52: For the file database subtasks that fail to be imported, the distributed task scheduler will put them back into the parent task for the next round of allocation. Tasks that fail to be allocated and executed after reaching the target number of times will no longer be allocated and will be put into the task queue of failed executions.
6. The distributed storage method for massive InSAR deformation monitoring data according to claim 1, characterized in that, Step S6: When the node executes the import action of the file data block, send the real-time import process information of the task package to the distributed task scheduler, driving the distributed task scheduler to move the task package allocated to the corresponding node from the running task queue to the task package queue according to the received real-time import process information and heartbeat monitoring results. Specifically, it includes: S61: When the node executes the import action of the file data block, send the real-time import process information of the task package to the distributed task scheduler, and the distributed task scheduler reorganizes and merges the import process information according to files and file data blocks; among them, the real-time import process information includes import process status, logs, and progress information; S62: During the import process, when it is detected that the node does not feedback the real-time import process information within the target duration or the heartbeat monitoring is abnormal, the tasks allocated to the corresponding node will be moved from the running task queue to the task package queue to wait for the next reallocation.
7. The method for distributed storage of massive InSAR deformation monitoring data according to claim 1, wherein, Step S7: When the import of the file data block of each parent task is completed, determine the unimported file data blocks according to the received real-time import process information, parse the unimported file data blocks, and perform the warehousing process on the InSAR deformation monitoring data obtained by the parsing. Specifically, it includes: S71: After the import of the file data block of each parent task is completed, the distributed task scheduler sorts all the incomplete data blocks received based on the data block order, and estimates the first processing duration for merging several data blocks corresponding to each file data block considering the CPU processing speed and the second processing duration for warehousing each file data block considering the IO port transmission speed according to the data volume of several data blocks corresponding to each file data block and the predicted CPU occupancy ratio and IO port transmission speed of the task scheduler at each subsequent moment. Based on the first processing duration and the second processing duration, generate the best order control strategy for the merging and warehousing of all incomplete data blocks according to the principle of parallel CPU processing and IO port transmission with the shortest time consumption, and use this best order control strategy to sequentially select the corresponding data blocks from the sorting result to merge into complete file data blocks; Among them, the principle of parallel CPU processing and IO port transmission with the shortest time consumption is configured as: the first group of data blocks first performs the merging process, and each subsequent group of data blocks performs the merging process while the previous group of data blocks performs the warehousing process until all data blocks complete the merging process and the warehousing process; and the sum of the longer duration between the second processing duration of the previous group of data blocks for warehousing after merging to obtain the file data block and the first processing duration of the next group of data blocks for merging plus the sum of the first processing duration of the first group of data blocks and the second processing duration of the last group of data blocks is the shortest; S72: Sequentially traverse and read the complete data blocks in the file data blocks, parse the key information including spatial features, temporal features, and deformation variable features, build a spatio-temporal index for the feature data and then perform the warehousing process, complete the data warehousing operation of the corresponding InSAR deformation monitoring data file, update the status of the parent task to completed, and at the same time move the corresponding data packet task to the completed task queue.
8. A distributed storage system for a large amount of InSAR deformation monitoring data, characterized in that, Including: A construction module for distributing and storing the original InSAR deformation monitoring data file and building a list of InSAR deformation monitoring data files; wherein, the list of InSAR deformation monitoring data files stores several file data blocks of each InSAR deformation monitoring data file. A task packet queue formation module for using each InSAR deformation monitoring data file as a parent task and each file data block of each InSAR deformation monitoring data file as a subtask of the parent task to form a task packet queue. A task packet allocation module for using a distributed task scheduler to load the import node information in the cluster, extract the node status of each node and establish heartbeat monitoring between the distributed task scheduler and each node in the cluster, allocate task packets to the nodes with an idle status in the cluster, move the task packets confirmed by the nodes to the running task queue, and mark the relationship between the task packets and the nodes. A first warehousing module for each node to build a file data block import task after receiving a task packet, sequentially traverse and read the file data blocks in the task packet, and perform the warehousing process on the parsed InSAR deformation monitoring data obtained; specifically including: After each node receives and confirms, it parses the task packet and builds a file data block import task in the form of a thread pool for the number of subtasks in the task packet. Each file data block import task in the node sequentially traverses and reads the complete data blocks in the file data blocks, parses the key information including spatial features, temporal features, and deformation variable features, and builds a spatio-temporal index for the feature data and then performs the warehousing process, specifically including: For incomplete data blocks, after the node completes the import of the complete information, it synchronizes the incomplete data blocks to the main thread. After the main thread of the node completes the import of all task packets, it synchronizes its own status to idle, and at the same time packs the incomplete data blocks and the file data block subtasks and sends them as a whole to the distributed task scheduler. A sending module for, when the node detects that the file data block import fails, sending the subtask corresponding to the file data block back to the distributed task scheduler, driving the distributed task scheduler to put the subtask into the parent task to wait for the next round of task packet allocation. A moving module for, when the node executes the import action of the file data block, sending the real-time import process information of the task packet to the distributed task scheduler, driving the distributed task scheduler to move the task packets allocated to the corresponding node from the running task queue to the task packet queue according to the received real-time import process information and the heartbeat monitoring result. The second warehousing module is used to determine the unimported file data blocks according to the received real-time import process information when the import of the file data blocks of each parent task is completed, parse the unimported file data blocks, and perform warehousing processing on the InSAR deformation monitoring data obtained by the parsing; The repetition module is used to return to execute the task package allocation step until all the task packages in the task package queue are allocated, and generate an import task execution report according to the received real-time import information.
Citation Information
Patent Citations
Massive geoscientific data parallel processing method based on distributed file system
CN103198097A
Data synchronization method
CN107491565A