Massive InSAR deformation monitoring data distributed storage method and system

Through distributed storage and distributed task scheduler technology, massive InSAR deformation monitoring data storage tasks are divided into sub-tasks, solving the problems of complex, low efficiency and inscalability of data storage in the existing technology, and achieving efficient, real-time and reliable data storage processing.

CN120144541AActive Publication Date: 2025-06-13YUNNAN PROVINCIAL GEOLOGICAL ENVIRONMENT MONITORING INST (YUNNAN PROVINCIAL INST OF ENVIRONMENTAL GEOLOGY)
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510632368.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

The prior art has problems such as high complexity, poor reliability, low data import efficiency, inability to guarantee data timeliness and inscalability when processing massive InSAR deformation monitoring data.

Method used

Distributed storage and distributed computing technology are adopted to store InSAR deformation monitoring data files in distributed storage, and the data storage task is divided into independent subtasks through a distributed task scheduler, and the physical cluster resources are reasonably scheduled and allocated based on the amount of data processed.

Benefits of technology

It realizes efficient, real-time and reliable database processing of massive InSAR deformation monitoring data, and provides elastic scaling capabilities, significantly improving data database database efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144541A_ABST
    Figure CN120144541A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of InSAR deformation monitoring data processing, and discloses a distributed storage method and system for massive InSAR deformation monitoring data. The method comprises the steps of extraction and list construction of distributed storage InSAR deformation monitoring data files, formation of task package queues containing parent tasks and subtasks, cluster node task package distribution based on a distributed task scheduler, execution and storage processing of file data block import tasks, and storage processing of the distributed storage InSAR deformation monitoring data files. Performing analysis and storage processing on the file data blocks which fail to be imported and the file data blocks which are not imported; and repeatedly distributing task packages. According to the method, technical means such as distributed storage, distributed calculation, dynamic scheduling and parallel processing are adopted, an InSAR deformation monitoring data storage task is divided into independent sub-tasks by utilizing a divide-and-conquer thought, efficient, real-time and reliable storage processing is provided for massive InSAR deformation monitoring data by fully utilizing limited physical resources, and the method has the advantages of high efficiency, high reliability and high efficiency. And meanwhile, elastic expansion and contraction capability can be provided based on cluster characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of InSAR deformation monitoring data processing, and particularly to a method and system for distributed storage of a large amount of InSAR deformation monitoring data. Background Art

[0002] The InSAR deformation monitoring technology has the characteristics of high precision, wide range, large time span, etc., and has now become a new means with great development potential in the field of surface deformation monitoring. With the application of the InSAR deformation monitoring technology, a large number of historical and real-time data files of deformation monitoring have been generated, reaching the TB or even PB level. The InSAR deformation monitoring data itself carries information such as spatial attributes, time attributes, and change amounts, and can be applied in multiple fields such as geological disaster monitoring and urban ground settlement monitoring.

[0003] If the conventional single-body architecture method is used to process a large amount of InSAR deformation monitoring data files, the following problems will exist: (1) High complexity: Since the entire import logic is implemented in a single-body application, the application will become very complex, making it difficult to modify and maintain; (2) Poor reliability: During the data import process, if an abnormal interruption occurs, it is easy to cause the entire data import process to terminate; (3) Low data import efficiency: Due to physical storage, memory and other conditions, reading and processing ultra-large files can only be carried out in a serial manner, resulting in low overall import efficiency; (4) Data timeliness cannot be guaranteed: Due to the low data import efficiency, the best analysis time may have been missed by the time the data is stored in the database, and it no longer has analysis value; (5) Non-scalability: As the amount of processed data accumulates, when the performance reaches the limit, effective expansion cannot be carried out. At the same time, if the processing volume decreases, high amounts of physical resources are occupied and cannot be released.

[0004] Therefore, how to efficiently store a large amount of InSAR deformation monitoring data in the database is a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The present invention provides a method and system for distributed storage of a large amount of InSAR deformation monitoring data, aiming to solve at least one of the above technical problems.

[0006] To achieve the above object, the present invention provides a method for distributed storage of a large amount of InSAR deformation monitoring data, including:

[0007] S1: Distributively store the original InSAR deformation monitoring data files, and construct a list of InSAR deformation monitoring data files; wherein, the list of InSAR deformation monitoring data files stores several file data blocks of each InSAR deformation monitoring data file.

[0008] S2: Take each InSAR deformation monitoring data file as a parent task, and take the file data blocks of each InSAR deformation monitoring data file as subtasks of the parent task to form a task package queue;

[0009] S3: Use the distributed task scheduler to load the import node information in the cluster, extract the node status of each node and establish heartbeat monitoring between the distributed task scheduler and each node in the cluster, allocate task packages to the nodes with idle node status in the cluster, move the task packages confirmed by the nodes to the running task queue, and mark the relationship between the task packages and the nodes;

[0010] S4: Each node constructs a file data block import task after receiving a task package, traverses and reads the file data blocks in the task package in sequence, and parses the obtained InSAR deformation monitoring data for storage in the database;

[0011] S5: When a node detects that the file data block import fails, send the subtask corresponding to the file data block back to the distributed task scheduler, driving the distributed task scheduler to put the subtask into the parent task to wait for the next round of task package allocation;

[0012] S6: When a node executes the import action of the file data block, send the real-time import process information of the task package to the distributed task scheduler, driving the distributed task scheduler to move the task package allocated to the corresponding node from the running task queue to the task package queue according to the received real-time import process information and the heartbeat monitoring result;

[0013] S7: When the import of the file data blocks of each parent task is completed, determine the unimported file data blocks according to the received real-time import process information, parse the unimported file data blocks, and parse the obtained InSAR deformation monitoring data for storage in the database;

[0014] S8: Return to execute the task package allocation step until all the task packages in the task package queue are allocated, and generate an import task execution report according to the received real-time import information.

[0015] Optionally, step S1: Distribute and store the original InSAR deformation monitoring data files, and construct a list of InSAR deformation monitoring data files, specifically including:

[0016] S11: Store the original InSAR deformation monitoring data files using distributed storage technology;

[0017] S12: Read the list of InSAR deformation monitoring data files in the distributed storage, and based on the technical characteristics of the file data blocks in the distributed storage, obtain all the file data block characteristic data of each InSAR deformation monitoring data file;

[0018] Among them, the file data block feature data includes a block serial number, a starting offset, and a block length.

[0019] Optionally, step S2: Using each InSAR deformation monitoring data file as a parent task and the file data blocks of each InSAR deformation monitoring data file as subtasks of the parent task, a task package queue is formed, which specifically includes:

[0020] S21: Using each InSAR deformation monitoring data file as a parent task and the file data blocks of each InSAR deformation monitoring data file as subtasks of the parent task, an InSAR deformation monitoring task queue is constructed;

[0021] S22: Start the distributed task scheduler to load the InSAR deformation monitoring task queue, and package all the subtasks according to a preset quantity to form a task package queue.

[0022] Optionally, step S3: Use the distributed task scheduler to load the import node information in the cluster, extract the node status of each node, establish heartbeat monitoring between the distributed task scheduler and each node in the cluster, and allocate task packages to the nodes in the cluster with an idle status. Move the task packages confirmed by the nodes to the running task queue and mark the relationship between the task packages and the nodes, which specifically includes:

[0023] S31: Use the distributed task scheduler to load the import node information in the cluster, extract the node status of each node, and establish heartbeat monitoring between the distributed task scheduler and the nodes in the cluster; among them, the node status includes idle and occupied;

[0024] S32: The distributed task scheduler allocates task packages to the nodes in the cluster with an idle status. When the number of task packages is less than the number of idle nodes, all task packages are allocated. When the number of task packages is more than the number of idle nodes, a corresponding part of the task packages equal to the number of idle nodes is allocated;

[0025] S33: After each node receives a task package, first synchronize its own status to occupied, and send a task package reception confirmation message to the distributed task scheduler. The distributed task scheduler moves the task packages confirmed by the reception to the running task queue, marks the relationship between the task packages and the nodes, and puts the task packages that have not received confirmation back into the task package queue to wait for the next allocation.

[0026] Optionally, step S4: Each node constructs a file data block import task after receiving a task package, sequentially traverses and reads the file data blocks in the task package, and performs database storage processing on the parsed InSAR deformation monitoring data, which specifically includes:

[0027] S41: After each node receives the confirmation, it parses the task package and constructs a file data block import task for the number of subtasks in the task package in the form of a thread pool;

[0028] S42: Each file data block import task in the node sequentially traverses and reads the complete data blocks in the file data block, parses the key information including spatial features, temporal features, and deformation variable features, and constructs a spatio-temporal index for the feature data and then performs the warehousing process.

[0029] Optionally, step S42: Each file data block import task in the node sequentially traverses and reads the complete data blocks in the file data block, parses the key information including spatial features, temporal features, and deformation variable features, and constructs a spatio-temporal index for the feature data and then performs the warehousing process, specifically including:

[0030] S421: For incomplete data blocks, after the node completes the import of the complete information, it synchronizes the incomplete data blocks to the main thread;

[0031] S422: After the main thread of the node completes the import of all task packages, it synchronizes its own state to idle, and at the same time packs the incomplete data blocks and file data block subtasks and sends them as a whole to the distributed task scheduler.

[0032] Optionally, step S5: When the node detects that the file data block import fails, it sends the subtask corresponding to the file data block back to the distributed task scheduler, driving the distributed task scheduler to put the subtask into the parent task to wait for the next round of task package allocation, specifically including:

[0033] S51: If the import of a certain file block fails during the import process in the node, it sends the task block information back to the distributed task scheduler;

[0034] S52: For the file database subtask with import failure, the distributed task scheduler puts it back into the parent task for the next round of allocation, and the tasks that fail to be allocated after reaching the target number of times will no longer be allocated and are put into the task queue of execution failures.

[0035] Optionally, step S6: When the node executes the import action of the file data block, it sends the real-time import process information of the task package to the distributed task scheduler, driving the distributed task scheduler to move the task package that has been allocated to the corresponding node from the running task queue to the task package queue according to the received real-time import process information and heartbeat monitoring results, specifically including:

[0036] S61: When a node executes the import action of file data blocks, send the real-time import process information of the task package to the distributed task scheduler, and the distributed task scheduler reorganizes and merges the import process information according to files and file data blocks; wherein, the real-time import process information includes import process status, logs, and progress information;

[0037] S62: During the import process, when it is detected that a node does not feedback the real-time import process information within the target duration or the heartbeat monitoring is abnormal, the tasks allocated to the corresponding node will be moved from the running task queue to the task package queue to wait for the next reallocation.

[0038] Optionally, step S7: When the import of file data blocks of each parent task is completed, determine the unimported file data blocks according to the received real-time import process information, parse the unimported file data blocks, and perform warehousing processing on the InSAR deformation monitoring data obtained by parsing, specifically including:

[0039] S71: After the import of file data blocks of each parent task is completed, the distributed task scheduler sorts all the incomplete data blocks received based on the data block order, and estimates the first processing duration for merging several data blocks corresponding to each file data block considering the CPU processing speed and the second processing duration for warehousing processing of each file data block considering the IO port transmission speed according to the data volume sizes of several data blocks corresponding to each file data block and the predicted values of the CPU occupancy ratio and the IO port transmission speed at each subsequent moment of the task scheduler. Based on the first processing duration and the second processing duration, generate the best order control strategy for the merging and warehousing of all incomplete data blocks according to the principle of parallel CPU processing and IO port transmission with the shortest time consumption, and use this best order control strategy to sequentially select the corresponding data blocks from the sorting result to merge into complete file data blocks;

[0040] Among them, the principle of parallel CPU processing and IO port transmission with the shortest time consumption is configured as follows: The first group of data blocks first performs the merging process, and each subsequent group of data blocks performs the merging process while the previous group of data blocks performs the warehousing process until all data blocks complete the merging process and the warehousing process; and the sum of the longer duration between the second processing duration for the previous group of data blocks to perform the warehousing process after merging to obtain the file data block and the first processing duration for the next group of data blocks to perform the merging process among all adjacent two groups of data blocks plus the sum of the first processing duration of the first group of data blocks and the second processing duration of the last group of data blocks is the shortest;

[0041] S72: Sequentially traverse and read the complete data blocks in the file data blocks, parse the key information including spatial features, temporal features, and deformation variable features, build a spatio-temporal index for the feature data, and then perform the warehousing process. Complete the data warehousing operation of the corresponding InSAR deformation monitoring data file, update the status of the parent task to completed, and at the same time move the corresponding data packet task to the completed task queue.

[0042] In addition, to achieve the above object, the present invention also provides a distributed warehousing system for a large amount of InSAR deformation monitoring data, including:

[0043] A construction module for distributing and storing the original InSAR deformation monitoring data files and building a list of InSAR deformation monitoring data files; wherein, the list of InSAR deformation monitoring data files stores several file data blocks of each InSAR deformation monitoring data file.

[0044] A task packet queue formation module for taking each InSAR deformation monitoring data file as a parent task and each file data block of each InSAR deformation monitoring data file as a sub-task of the parent task to form a task packet queue.

[0045] A task packet allocation module for using a distributed task scheduler to load the import node information in the cluster, extracting the node status of each node and establishing heartbeat monitoring between the distributed task scheduler and each node in the cluster, allocating task packets to the nodes with an idle status in the cluster, moving the task packets confirmed by the nodes to the running task queue, and marking the relationship between the task packets and the nodes.

[0046] A first warehousing module for each node to build a file data block import task after receiving the task packet, sequentially traverse and read the file data blocks in the task packet, and perform the warehousing process on the parsed InSAR deformation monitoring data.

[0047] A sending module for, when a node detects that the file data block import fails, sending the sub-task corresponding to the file data block back to the distributed task scheduler, driving the distributed task scheduler to put the sub-task into the parent task to wait for the next round of task packet allocation.

[0048] A moving module for, when a node executes the import action of the file data block, sending the real-time import process information of the task packet to the distributed task scheduler, driving the distributed task scheduler to move the task packet allocated to the corresponding node from the running task queue to the task packet queue according to the received real-time import process information and the heartbeat monitoring result.

[0049] The second warehousing module is used to determine unimported file data blocks according to the received real-time import process information when the file data blocks of each parent task are imported, parse the unimported file data blocks, and perform warehousing processing on the InSAR deformation monitoring data obtained by the parsing;

[0050] The repetition module is used to return to execute the task package allocation step until all the task packages in the task package queue are allocated, and generate an import task execution report according to the received real-time import information.

[0051] The beneficial effects of the present invention are as follows: A method and system for distributed warehousing of massive InSAR deformation monitoring data are proposed, including extraction and list construction of InSAR deformation monitoring data files in distributed storage, formation of a task package queue including parent tasks and sub-tasks, task package allocation of cluster nodes based on a distributed task scheduler, execution and warehousing processing of file data block import tasks, parsing and warehousing processing of imported failed file data blocks and unimported file data blocks, and repeated allocation of task packages, etc. Thus, by adopting technical means such as distributed storage, distributed computing, dynamic scheduling, and parallel processing, based on the idea of divide and conquer, the InSAR deformation monitoring data warehousing task is divided into independent sub-tasks, and the physical cluster resources are reasonably scheduled and allocated based on the amount of data processed, making full use of limited physical resources to provide efficient, real-time, and reliable warehousing processing for massive InSAR deformation monitoring data, and at the same time providing elastic scaling ability based on the characteristics of the cluster, significantly improving the warehousing efficiency of massive InSAR deformation monitoring data. Description of the Drawings

[0052] Figure 1 It is a schematic flowchart of the method for distributed warehousing of massive InSAR deformation monitoring data of the present invention;

[0053] Figure 2 It is a schematic structural diagram of the system for distributed warehousing of massive InSAR deformation monitoring data of the present invention. Detailed Embodiment

[0054] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0055] The embodiment of the present invention provides a method for distributed warehousing of massive InSAR deformation monitoring data, referring to Figure 1 , Figure 1 It is a schematic flowchart of the method for distributed warehousing of massive InSAR deformation monitoring data in the embodiment of the present invention. The method for distributed warehousing of massive InSAR deformation monitoring data includes the following steps:

[0056] S1: Distribute and store the original InSAR deformation monitoring data files, and construct a list of InSAR deformation monitoring data files; wherein, several file data blocks of each InSAR deformation monitoring data file are stored in the list of InSAR deformation monitoring data files.

[0057] S2: Take each InSAR deformation monitoring data file as a parent task, and take the file data blocks of each InSAR deformation monitoring data file as subtasks of the parent task to form a task package queue.

[0058] S3: Use the distributed task scheduler to load the import node information in the cluster, extract the node status of each node and establish heartbeat monitoring between the distributed task scheduler and each node in the cluster, allocate task packages to the nodes with an idle node status in the cluster, move the task packages confirmed by the nodes to the running task queue, and mark the relationship between the task packages and the nodes.

[0059] S4: Each node constructs a file data block import task after receiving the task package, sequentially traverses and reads the file data blocks in the task package, and performs storage processing on the parsed InSAR deformation monitoring data obtained.

[0060] S5: When a node detects that the file data block import fails, send the subtask corresponding to the file data block back to the distributed task scheduler, and drive the distributed task scheduler to put the subtask into the parent task to wait for the next round of task package allocation.

[0061] S6: When a node executes the import action of the file data block, send the real-time import process information of the task package to the distributed task scheduler, and drive the distributed task scheduler to move the task package already allocated to the corresponding node from the running task queue to the task package queue according to the received real-time import process information and the heartbeat monitoring result.

[0062] S7: When the import of the file data blocks of each parent task is completed, determine the unimported file data blocks according to the received real-time import process information, parse the unimported file data blocks, and perform storage processing on the parsed InSAR deformation monitoring data obtained.

[0063] S8: Return to execute the task package allocation step until all the task packages in the task package queue are allocated, and generate an import task execution report according to the received real-time import information.

[0064] It should be noted that when the prior art processes a large amount of InSAR deformation monitoring data files using a conventional monolithic architecture method, the following problems will occur: (1) High complexity: Since the entire import logic is implemented in a single monolithic application, the application will become very complex, making it difficult to modify and maintain; (2) Poor reliability: During the data import process, if an abnormal interruption occurs, it is easy to cause the entire data import process to terminate; (3) Low data import efficiency: Due to limitations such as physical storage and memory, reading and processing extremely large files can only be done serially, resulting in low overall import efficiency; (4) Data timeliness cannot be guaranteed: Due to low data import efficiency, by the time the data is stored in the database, the optimal analysis time may have passed, and the data may no longer be valuable for analysis; (5) Lack of scalability: As the amount of processed data accumulates, when the performance reaches its limit, effective expansion cannot be carried out. At the same time, if the processing volume decreases, high amounts of physical resources are occupied and cannot be released.

[0065] In this embodiment, through steps such as the extraction and list construction of InSAR deformation monitoring data files in distributed storage, the formation of a task package queue containing parent tasks and child tasks, the allocation of task packages to cluster nodes based on a distributed task scheduler, the execution and storage processing of file data block import tasks, the parsing and storage processing of imported failed file data blocks and unimported file data blocks, and the repeated allocation of task packages, using technical means such as distributed storage, distributed computing, dynamic scheduling, and parallel processing, based on the divide-and-conquer idea, the task of storing InSAR deformation monitoring data in the database is divided into independent sub-tasks. Based on the amount of processed data, the physical cluster resources are reasonably scheduled and allocated, making full use of limited physical resources to provide efficient, real-time, and reliable storage processing for a large amount of InSAR deformation monitoring data. At the same time, based on the characteristics of the cluster, elastic scalability is provided, significantly improving the storage efficiency of a large amount of InSAR deformation monitoring data.

[0066] To more clearly explain the present invention, the following provides a specific example of the method for distributed storage of a large amount of InSAR deformation monitoring data of the present invention, which includes the following specific execution steps:

[0067] Step 1: Store the original InSAR deformation monitoring data files using distributed storage technology;

[0068] Step 2: Read the list of InSAR deformation monitoring data files in distributed storage, and based on the technical characteristics of file data blocks in distributed storage, obtain all file data block characteristic data (block number, starting offset, block length) of each InSAR deformation monitoring data file;

[0069] Step 3: Build a task queue based on the list of InSAR deformation monitoring data files. Import each InSAR deformation monitoring data file as a parent task, and import the file data blocks of each InSAR deformation monitoring data file as subtasks within the parent task;

[0070] Step 4: Start the distributed task scheduler, load the task list, and pack all the subtasks into task packages according to a specific number (e.g., 10) to form a task package queue;

[0071] Step 5: The distributed task scheduler loads the import node information in the cluster, including node status (occupied, idle), and simultaneously establishes a heartbeat monitoring between the distributed task scheduler and the cluster nodes;

[0072] Step 6: The distributed task scheduler allocates task packages to the nodes in the cluster with an idle status. If the number of task packages is less than the number of idle nodes at this time, all task packages are allocated. If the number of task packages is more than the number of idle nodes at this time, only some task packages are allocated;

[0073] Step 7: After receiving a task package, the node first synchronizes its own status to occupied and replies with a task package acceptance confirmation message. The distributed task scheduler moves the task packages with accepted confirmations to the running task queue, marks the relationship between the task packages and the nodes, and reinserts the task packages that have not received confirmations into the task package queue to wait for the next allocation;

[0074] Step 8: After receiving the confirmation, the node will parse the task package and construct file data block import tasks through a thread pool according to the number of subtasks in the task package;

[0075] Step 9: Each file data block import task in the node sequentially traverses and reads the complete data blocks in the file data block, and parses key information such as spatial features, temporal features, and deformation amount features. After constructing a spatio-temporal index for the feature data, it is stored in the database. For incomplete data blocks, after importing the complete information, the incomplete data blocks are synchronized to the main thread;

[0076] Step 10: After the main thread of the node completes the import of all task packages, it synchronizes its own status to idle, and at the same time packs the incomplete data blocks and file data block subtasks and sends them to the distributed task scheduler as a whole;

[0077] Step 11: If the import of a certain file block fails during the import process in the node, the task block information is sent back to the distributed task scheduler;

[0078] Step Twelve: For the file database subtasks that failed to be imported, the distributed task scheduler will put them back into the parent task for the next round of allocation. Tasks that failed to be executed after multiple allocations will no longer be allocated and will be put into the task queue of failed executions.

[0079] Step Thirteen: When a node executes the import task of a file data block, it will simultaneously synchronize the import process status, logs, and progress information of the task package to the distributed task scheduler. The scheduler will reorganize and merge the import process information according to files and file data blocks.

[0080] Step Fourteen: If a node does not feedback process information for a long time or the heartbeat monitoring is abnormal during the import process, the tasks allocated to the corresponding node will be moved from the running task queue to the task queue to wait for the next reallocation.

[0081] Step Fifteen: After the scheduling of the import parent task for a file is completed, the distributed task scheduler will merge all the incomplete information of the file data blocks received based on the order of the data blocks to form complete file data blocks. It will sequentially traverse and read the complete data blocks in the file data blocks, parse the key information such as spatial features, temporal features, and deformation variable features, and build a spatio-temporal index for the feature data and then perform the warehousing process. At this time, the data warehousing operation of the corresponding InSAR deformation monitoring data file is completed, the status of the parent task is updated to completed, and the corresponding data packet task is moved to the completed task queue.

[0082] It should be noted that the distributed task scheduler merges all the incomplete file data block information received based on the order of the data blocks to form complete file data block steps, which specifically include: after the file data blocks of each parent task are imported, the distributed task scheduler sorts all the received incomplete data blocks based on the data block order, and estimates the first processing duration for merging several data blocks corresponding to each file data block considering the CPU processing speed and the second processing duration for storing each file data block considering the IO port transmission speed according to the data volume of several data blocks corresponding to each file data block and the predicted CPU occupancy ratio and IO port transmission speed of the task scheduler at each subsequent moment. Based on the first processing duration and the second processing duration, according to the principle of parallel CPU processing and IO port transmission with the shortest time consumption, an optimal order control strategy for merging and storing all incomplete data blocks is generated, and using this optimal order control strategy, the corresponding data blocks are sequentially selected from the sorting result to be merged into complete file data blocks; wherein, the principle of parallel CPU processing and IO port transmission with the shortest time consumption is configured as: the first group of data blocks first performs merging processing, and each subsequent group of data blocks performs merging processing while the previous group of data blocks performs storing processing until all data blocks complete merging processing and storing processing; and the sum of the longer duration between the second processing duration of storing the previous group of data blocks after merging to obtain the file data block and the first processing duration of merging the next group of data blocks for all adjacent two groups of data blocks plus the total duration of the first processing duration of the first group of data blocks and the second processing duration of the last group of data blocks is the shortest.

[0083] Thus, by adopting the principle of parallel CPU processing and IO port transmission with the shortest time consumption, an optimal order control strategy for merging and storing all incomplete data blocks is generated, and based on this, the corresponding data blocks are sequentially selected from the sorting result to be merged into complete file data blocks. In the scenario where there are strict requirements for the import and storage time control of a large number of incomplete data blocks in InSAR deformation monitoring data, it can significantly improve the data import and storage efficiency and increase the upper limit of the device's data processing capacity.

[0084] Step Sixteen: Continuously repeat Step Six until all task queues are allocated, and perform file grouping and merging on the collected import process monitoring data to form an import task execution report.

[0085] Refer to Figure 2 , Figure 2 which is a schematic structural diagram of the distributed storage system for a large amount of InSAR deformation monitoring data in an embodiment of the present invention. The distributed storage system for a large amount of InSAR deformation monitoring data includes:

[0086] Building module 10 is used to store the original InSAR deformation monitoring data files in a distributed manner and construct a list of InSAR deformation monitoring data files; wherein, the list of InSAR deformation monitoring data files stores several file data blocks of each InSAR deformation monitoring data file.

[0087] Task package queue forming module 20 is used to form a task package queue with each InSAR deformation monitoring data file as the parent task and each file data block of each InSAR deformation monitoring data file as the subtask of the parent task.

[0088] Task package allocation module 30 is used to load the import node information in the cluster by using a distributed task scheduler, extract the node status of each node and establish heartbeat monitoring between the distributed task scheduler and each node in the cluster, allocate task packages to the nodes with an idle node status in the cluster, move the task packages confirmed by the nodes to the running task queue, and mark the relationship between the task packages and the nodes.

[0089] First warehousing module 40 is used for each node to construct a file data block import task after receiving a task package, sequentially traverse and read the file data blocks in the task package, and perform warehousing processing on the parsed InSAR deformation monitoring data obtained.

[0090] Sending module 50 is used to send the subtask corresponding to the file data block back to the distributed task scheduler when the node detects that the file data block import fails, driving the distributed task scheduler to put the subtask into the parent task to wait for the next round of task package allocation.

[0091] Moving module 60 is used to send the real-time import process information of the task package to the distributed task scheduler when the node executes the import action of the file data block, driving the distributed task scheduler to move the task package allocated to the corresponding node from the running task queue to the task package queue according to the received real-time import process information and the heartbeat monitoring result.

[0092] Second warehousing module 70 is used to determine the unimported file data blocks according to the received real-time import process information when the import of the file data blocks of each parent task is completed, parse the unimported file data blocks, and perform warehousing processing on the parsed InSAR deformation monitoring data obtained.

[0093] Repeating module 80 is used to return to execute the task package allocation step until all the task packages in the task package queue are allocated, and generate an import task execution report according to the received real-time import information.

[0094] Other embodiments or specific implementation manners of the distributed warehousing system for a large amount of InSAR deformation monitoring data of the present invention may refer to the above method embodiments, and will not be elaborated here.

[0095] It is understood that in the description of this specification, the descriptions referring to terms such as "one embodiment", "another embodiment", "other embodiments", or "the first embodiment to the Nth embodiment" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0096] It should be noted that in this article, the term "comprising", "including", or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or system including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or system including that element.

[0097] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A distributed storage method for massive InSAR deformation monitoring data, characterized in that: include: S1: Distributed storage of original InSAR deformation monitoring data files to construct an InSAR deformation monitoring data file list; wherein the InSAR deformation monitoring data file list stores a plurality of file data blocks of each InSAR deformation monitoring data file; S2: Take each InSAR deformation monitoring data file as a parent task, and take the file data block of each InSAR deformation monitoring data file as a child task of the parent task to form a task package queue; S3: Use the distributed task scheduler to load the imported node information in the cluster, extract the node status of each node and establish heartbeat monitoring between the distributed task scheduler and each node in the cluster, assign task packages to nodes in the cluster whose node status is idle, move the task packages received and confirmed by the nodes to the running task queue, and mark the relationship between the task packages and the nodes; S4: After receiving the task package, each node constructs a file data block import task, traverses and reads the file data blocks in the task package in sequence, and parses the obtained InSAR deformation monitoring data for storage processing; S5: When the node detects that the file data block import fails, the node sends the subtask corresponding to the file data block back to the distributed task scheduler, driving the distributed task scheduler to put the subtask into the parent task and wait for the next round of task package allocation; S6: When the node executes the import action of the file data block, the real-time import process information of the task package is sent to the distributed task scheduler, driving the distributed task scheduler to move the task package assigned to the corresponding node from the running task queue to the task package queue according to the received real-time import process information and the heartbeat monitoring result; S7: when the file data blocks of each parent task are imported, the unimported file data blocks are determined according to the received real-time import process information, the unimported file data blocks are parsed, and the InSAR deformation monitoring data obtained by the parsing is stored in the warehouse; S8: Return to the task package allocation step until all task packages in the task package queue are allocated, and generate an import task execution report according to the received real-time import information.

2. The distributed storage method for massive InSAR deformation monitoring data according to claim 1, characterized in that: Step S1: Distribute and store the original InSAR deformation monitoring data files to construct an InSAR deformation monitoring data file list, which specifically includes: S11: Use distributed storage technology to store the original InSAR deformation monitoring data files; S12: reading the InSAR deformation monitoring data file list in the distributed storage, and obtaining all file data block feature data of each InSAR deformation monitoring data file based on the technical features of the file data blocks in the distributed storage; The file data block characteristic data includes a block sequence number, a starting offset and a block length.

3. The distributed storage method for massive InSAR deformation monitoring data according to claim 1, characterized in that: Step S2: Taking each InSAR deformation monitoring data file as a parent task and the file data block of each InSAR deformation monitoring data file as a subtask of the parent task, a task package queue is formed, which specifically includes: S21: Taking each InSAR deformation monitoring data file as a parent task and the file data block of each InSAR deformation monitoring data file as a child task of the parent task, an InSAR deformation monitoring task queue is constructed; S22: Start the distributed task scheduler to load the InSAR deformation monitoring task queue, and pack all subtasks according to a preset number to form a task package queue.

4. The distributed storage method for massive InSAR deformation monitoring data according to claim 1, characterized in that: Step S3: Use the distributed task scheduler to load the imported node information in the cluster, extract the node status of each node and establish heartbeat monitoring between the distributed task scheduler and each node in the cluster, assign task packages to nodes in the cluster whose node status is idle, move the task packages received and confirmed by the nodes to the running task queue, and mark the relationship between the task packages and the nodes, specifically including: S31: using the distributed task scheduler to load the imported node information in the cluster, extracting the node status of each node, and establishing heartbeat monitoring between the distributed task scheduler and the nodes in the cluster; wherein the node status includes idle and occupied; S32: The distributed task scheduler allocates task packages to nodes in the cluster whose node status is idle. When the number of task packages is less than the number of idle nodes, all task packages are allocated. When the number of task packages is more than the number of idle nodes, part of the task packages corresponding to the number of idle nodes are allocated. S33: After each node receives the task package, it first synchronizes its own status to occupied, and replies to the distributed task scheduler with confirmation information of task package reception. The distributed task scheduler moves the confirmed task package to the running task queue, marks the relationship between the task package and the node, and puts the task package that has not received confirmation back into the task package queue to wait for the next allocation.

5. The distributed storage method for massive InSAR deformation monitoring data according to claim 1, characterized in that: Step S4: After receiving the task package, each node constructs a file data block import task, sequentially traverses and reads the file data blocks in the task package, and parses the obtained InSAR deformation monitoring data for storage processing, which specifically includes: S41: After receiving the confirmation, each node parses the task package and constructs a file data block import task through a thread pool according to the number of subtasks in the task package; S42: Each file data block import task in the node traverses and reads the complete data block in the file data block in sequence, parses the key information including spatial features, temporal features, and shape variable features, and constructs a spatiotemporal index for the feature data before warehousing.

6. The distributed storage method for massive InSAR deformation monitoring data according to claim 5, characterized in that: Step S42: Each file data block import task in the node sequentially traverses and reads the complete data block in the file data block, parses the key information including spatial features, temporal features, and shape variable features, and constructs a spatiotemporal index for the feature data before storing it in the database, specifically including: S421: For the incomplete data block, after the node completes importing the complete information, it synchronizes the incomplete data block to the main thread; S422: After the node main thread completes the import of all task packages, it synchronizes its own state to idle, and packages the incomplete data blocks and file data block subtasks, and sends the whole to the distributed task scheduler.

7. The distributed storage method for massive InSAR deformation monitoring data according to claim 6, characterized in that: Step S5: When the node detects that the file data block import fails, the subtask corresponding to the file data block is sent back to the distributed task scheduler, driving the distributed task scheduler to put the subtask into the parent task to wait for the next round of task package allocation, specifically including: S51: If the node fails to import a certain file block during the import process, the node sends the task block information back to the distributed task scheduler; S52: For the file database subtask that failed to be imported, the distributed task scheduler puts it back into the parent task for the next round of allocation. Tasks that failed to be allocated and executed after reaching the target number of times will no longer be allocated and will be put into the failed task queue.

8. The distributed storage method for massive InSAR deformation monitoring data according to claim 7, characterized in that: Step S6: When the node executes the import action of the file data block, the real-time import process information of the task package is sent to the distributed task scheduler, and the distributed task scheduler is driven to move the task package assigned to the corresponding node from the running task queue to the task package queue according to the received real-time import process information and the heartbeat monitoring result, specifically including: S61: When the node executes the import action of the file data block, the real-time import process information of the task package is sent to the distributed task scheduler, and the distributed task scheduler reorganizes and merges the import process information according to the file and the file data block; wherein the real-time import process information includes the import process status, log and progress information; S62: During the import process, if it is detected that the node does not feedback the real-time import process information within the target time or the heartbeat monitoring is abnormal, the task assigned to the corresponding node will be moved from the running task queue to the task package queue to wait for the next reallocation.

9. The distributed storage method for massive InSAR deformation monitoring data according to claim 8, characterized in that: Step S7: when the file data block of each parent task is imported, the unimported file data blocks are determined according to the received real-time import process information, the unimported file data blocks are parsed, and the InSAR deformation monitoring data obtained by the parsing is stored in the warehouse, which specifically includes: S71: After the import of the file data blocks of each parent task is completed, the distributed task scheduler sorts all the incomplete data blocks received based on the order of the data blocks, and estimates the first processing time for merging the data blocks corresponding to each file data block when considering the CPU processing speed and the second processing time for warehousing each file data block when considering the IO port transmission speed, according to the data size of the data blocks corresponding to each file data block and the predicted value of the CPU share and the IO port transmission speed of the task scheduler at each subsequent moment. Based on the first processing time and the second processing time, according to the principle that the CPU processing and the IO port transmission are parallel and the time consumption is the shortest, an optimal sequential control strategy for merging and warehousing of all incomplete data blocks is generated, and the optimal sequential control strategy is used to select the corresponding data blocks from the sorting results in turn and merge them into complete file data blocks; The principle of parallel CPU processing and IO port transmission with the shortest time consumption is configured as follows: the first group of data blocks is first merged, and each subsequent group of data blocks is merged while the previous group of data blocks is stored, until all data blocks have completed the merge and storage processes; and the sum of the second processing time of the previous group of data blocks after merging to obtain file data blocks and the first processing time of the next group of data blocks, whichever has the larger value, plus the first processing time of the first group of data blocks and the second processing time of the last group of data blocks, is the shortest; S72: traverse and read the complete data blocks in the file data block in sequence, parse the key information including spatial features, time features, and deformation variable features, and construct a spatiotemporal index for the feature data and then perform storage processing to complete the corresponding InSAR deformation monitoring data file data storage operation, update the parent task status to completed, and move the corresponding data packet task to the completed task queue.

10. A distributed storage system for massive InSAR deformation monitoring data, characterized in that: include: A construction module is used to distribute and store the original InSAR deformation monitoring data files to construct an InSAR deformation monitoring data file list; wherein the InSAR deformation monitoring data file list stores a plurality of file data blocks of each InSAR deformation monitoring data file; A task package queue forming module is used to form a task package queue by taking each InSAR deformation monitoring data file as a parent task and taking a file data block of each InSAR deformation monitoring data file as a child task of the parent task; A task package allocation module is used to use a distributed task scheduler to load the imported node information in the cluster, extract the node status of each node and establish heartbeat monitoring between the distributed task scheduler and each node in the cluster, allocate task packages to nodes in the cluster whose node status is idle, move the task packages received and confirmed by the nodes to the running task queue, and mark the relationship between the task packages and the nodes; The first storage module is used for each node to construct a file data block import task after receiving the task package, traverse and read the file data blocks in the task package in sequence, and parse the obtained InSAR deformation monitoring data for storage processing; A sending module, used for sending the subtask corresponding to the file data block back to the distributed task scheduler when the node detects that the file data block fails to be imported, so as to drive the distributed task scheduler to put the subtask into the parent task and wait for the next round of task package allocation; A moving module, used for sending the real-time import process information of the task package to the distributed task scheduler when the node executes the import action of the file data block, driving the distributed task scheduler to move the task package assigned to the corresponding node from the running task queue to the task package queue according to the received real-time import process information and the heartbeat monitoring result; The second storage module is used to determine the unimported file data blocks according to the received real-time import process information when the file data blocks of each parent task are imported, parse the unimported file data blocks, and store the InSAR deformation monitoring data obtained by the analysis; The repetition module is used to return to the task package allocation step until all the task packages in the task package queue are allocated, and generate an import task execution report according to the received real-time import information.

Citation Information

Patent Citations

  • Massive geoscientific data parallel processing method based on distributed file system

    CN103198097A

  • Data synchronization method

    CN107491565A

  • Task scheduling method and device

    CN111355751A

  • Distributed high-performance data ETL device and control method

    CN112199432A

  • Dynamic task scheduling system for space control simulation model calculation

    CN115454649A