A distributed storage method and system for multimodal mapping data
Through the HDFS and MapReduce framework combined with quad-tree and octree structures and exception monitoring systems, the efficient storage and processing of multimodal mapping data is solved, and efficient data management and system reliability are achieved.
Patent Information
- Application Number
- CN202310913819.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-24
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-07-24
AI Technical Summary
The prior art is difficult to effectively manage and store multimodal, unstructured mapping data, especially when the data volume is huge, it is difficult for traditional databases to achieve efficient storage and processing.
HDFS is used as a distributed file system, and two-dimensional and three-dimensional mapping files are divided through quad-tree and octree structures, and abnormal monitoring is carried out in combination with Prometheus and Grafana frameworks to build a dual-machine hot standby mechanism to ensure system reliability.
The management, storage and processing efficiency of multimodal mapping data is improved, and abnormal monitoring of nodes and high-reliability operation of the system is realized.
Smart Images

Figure CN116955307B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of distributed storage technology, and in particular to a distributed storage method and system for multimodal mapping data. Background Art
[0002] With the advancement of science and technology, more and more mapping data are being used in geographic information processing, smart city construction and other aspects.
[0003] However, mapping data is characterized by being unstructured, semi-structured, and multimodal. It is also typically massive and complex to process. Traditional databases struggle to effectively store, manage, and process multimodal mapping data, necessitating the adoption of distributed storage and computing technologies. Hadoop is an open-source framework that supports data-intensive distributed applications. Hadoop provides an excellent distributed file system (HDFS) that is widely used for distributed storage.
[0004] How to use Hadoop for organization, management and distributed storage of different types of mapping data is an urgent problem to be solved. Summary of the Invention
[0005] The present invention provides a distributed storage method and system for multimodal mapping data, which are used to solve the defects of using traditional databases for mapping data in the prior art.
[0006] In a first aspect, the present invention provides a method for distributed storage of multimodal mapping data, comprising:
[0007] Use HDFS to store multimodal mapping data;
[0008] Organizing and managing the multimodal mapping data based on HDFS, and using a multi-level file directory to store the multimodal mapping data;
[0009] Perform abnormal monitoring on multiple storage nodes in HDFS and determine the dual-machine hot standby operation mode of the multiple storage nodes.
[0010] According to a distributed storage method for multimodal mapping data provided by the present invention, HDFS is used to store multimodal mapping data, including:
[0011] The HDFS includes a data I / O module, a data backup and recovery module, a distributed file storage module, and an anomaly detection module;
[0012] The data I / O module is used to receive data input from the client and return access data or stored mapping data processing results to the client;
[0013] The data backup and recovery module is used to create multiple redundant copies of data and adopt a dual-machine hot standby solution for data security storage and data recovery;
[0014] The distributed file storage module is used to provide file storage space, provide file organization and management, and organize and manage multimodal mapping data;
[0015] The anomaly detection module is used to monitor the operating status of multiple nodes in the Hadoop cluster and perform anomaly processing.
[0016] According to a distributed storage method for multimodal mapping data provided by the present invention, the multimodal mapping data is organized and managed based on HDFS, and the multimodal mapping data is stored in a multi-level file directory, including:
[0017] Read two-dimensional mapping files or three-dimensional mapping files through the data I / O module;
[0018] Using a quadtree structure to split the two-dimensional mapping file to obtain a quadtree file, and using an octree structure to split the three-dimensional mapping file to obtain an octree file;
[0019] The quadtree file or the octree file is organized according to a pyramid structure and stored in HDFS.
[0020] According to a distributed storage method for multimodal mapping data provided by the present invention, a quadtree structure is used to segment the two-dimensional mapping file to obtain a quadtree file, comprising:
[0021] Obtaining the center point position of the two-dimensional mapping file;
[0022] Determine the maximum block size in HDFS;
[0023] If it is determined that the data file block to be divided is larger than the maximum block size, the image of the two-dimensional mapping file is divided into four equal parts based on the two-dimensional coordinate axis with the center point position as the origin, and the quadtree node number of the divided data block is obtained;
[0024] The quadtree file is output according to the quadtree node number.
[0025] According to a distributed storage method for multimodal mapping data provided by the present invention, an octree structure is used to segment the three-dimensional mapping file to obtain an octree file, comprising:
[0026] Calculate the outer bounding box of the three-dimensional mapping file and obtain the center point position of the outer bounding box;
[0027] Determine the maximum block size in HDFS;
[0028] If it is determined that the data file block to be divided is larger than the maximum block size, traversing all three-dimensional space coordinate positions in the three-dimensional mapping file, comparing the corresponding position relationship between each three-dimensional space coordinate position and the center point position, and obtaining the child node number of the corresponding position relationship;
[0029] The octree file is output according to the child node number.
[0030] According to a multimodal mapping data distributed storage method provided by the present invention, multiple storage nodes in HDFS are monitored for abnormalities, and a dual-machine hot standby operation mode of the multiple storage nodes is determined, including:
[0031] An anomaly monitoring system is constructed using the service monitoring framework Prometheus and the dashboard graphic editor Grafana, and the anomaly monitoring system performs anomaly monitoring on the multiple storage nodes;
[0032] A reliable coordination system ZooKeeper is used as a cluster coordinator to build a dual-machine hot standby mechanism, and a dual-machine hot standby operation mode of the multiple storage nodes is executed based on the dual-machine hot standby mechanism.
[0033] According to a multimodal mapping data distributed storage method provided by the present invention, Prometheus and Grafana are used to build an anomaly monitoring system, and the anomaly monitoring system performs anomaly monitoring on the multiple storage nodes, including:
[0034] Setting an active master node and a standby master node, wherein the active master node is a master node in a normal operating state, and the standby master node is a master node that takes over the work of the active master node;
[0035] Prometheus uses the core component Prometheus Server to regularly pull data from the indicator exposer of the cluster node, and uses Push Gateway to transfer data that cannot be pulled directly;
[0036] PromQL language is used to determine the alarm rules, and node abnormality alarm information is generated based on the alarm rules;
[0037] The monitoring data is preprocessed using the PromQL language to obtain preprocessed data, which is then transmitted to Grafana for visualization.
[0038] According to a distributed storage method for multimodal mapping data provided by the present invention, ZooKeeper is used as a cluster coordinator to build a dual-machine hot standby mechanism, and a dual-machine hot standby operation mode of the multiple storage nodes is executed based on the dual-machine hot standby mechanism, including:
[0039] Register named nodes in the ZooKeeper cluster and obtain the session identifier of each named node;
[0040] If it is determined that any named node fails, the session identifier is switched to an expired state and a failover is initiated;
[0041] Using the exclusive lock mechanism, it is determined that only one of the two naming nodes is active, and the standby naming node obtains the exclusive lock.
[0042] The JournalNode cluster mechanism is adopted to share data between naming nodes. The JournalNode cluster pulls data from the active naming node and saves the data in real time, and the standby naming node synchronizes data from the JournalNode cluster in real time.
[0043] In a second aspect, the present invention further provides a multimodal mapping data distributed storage system, comprising:
[0044] Build a module for storing multimodal mapping data using HDFS;
[0045] A management module, configured to organize and manage the multimodal mapping data based on HDFS, and store the multimodal mapping data in a multi-level file directory;
[0046] The monitoring module is used to monitor multiple storage nodes in the HDFS for abnormalities and determine the dual-machine hot standby operation mode of the multiple storage nodes.
[0047] In a third aspect, the present invention also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements any of the multimodal mapping data distributed storage methods described above.
[0048] The distributed storage method and system for multimodal mapping data provided by the present invention organize and manage two- and three-dimensional mapping data by using HDFS as the distributed storage file system for mapping data, and adopting the MapReduce framework to implement distributed parallel computing, thereby greatly improving the management, storage, and processing efficiency of multimodal, unstructured spatial data. The system also includes an anomaly monitoring system based on the Prometheus and Grafana frameworks to monitor the operating status of each node, and combines a dual-machine hot standby design to ensure system reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0050] Figure 1 It is a flowchart of the distributed storage method for multimodal mapping data provided by the present invention;
[0051] Figure 2 This is a schematic diagram of the HDFS distributed file storage system module provided by the present invention;
[0052] Figure 3 This is a flow chart of the anomaly detection module provided by the present invention;
[0053] Figure 4 This is a structural diagram of the dual-machine hot standby mechanism provided by the present invention;
[0054] Figure 5 It is a schematic diagram of the structure of the multimodal mapping data distributed storage system provided by the present invention;
[0055] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0056] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0057] Figure 1 FIG. 1 is a flow chart of a distributed storage method for multimodal mapping data provided by an embodiment of the present invention. Figure 1 Shown, including:
[0058] Step 100: Use HDFS to store multimodal mapping data;
[0059] Step 200: Organizing and managing the multimodal mapping data based on HDFS, and using a multi-level file directory to store the multimodal mapping data;
[0060] Step 300: Perform abnormality monitoring on multiple storage nodes in HDFS and determine the dual-machine hot standby operation mode of the multiple storage nodes.
[0061] Specifically, the distributed storage method for multimodal mapping data proposed in the embodiment of the present invention is based on the Hadoop overall distributed file storage system as the overall distributed file storage system for multimodal mapping data.
[0062] Based on the storage system, a two-dimensional and three-dimensional mapping data organization and management model suitable for distributed file storage is proposed, and an abnormality monitoring and reliability assurance method for distributed file storage is proposed.
[0063] The present invention uses HDFS as a distributed storage file system for mapping data to organize and manage two- and three-dimensional mapping data, and adopts the MapReduce framework to implement distributed parallel computing, which greatly improves the management, storage, and processing efficiency of multimodal and unstructured spatial data. It also includes an anomaly monitoring system based on the Prometheus and Grafana frameworks to monitor the operating status of each node, and combines a dual-machine hot standby design to ensure system reliability.
[0064] Based on the above embodiment, HDFS is used to store multimodal mapping data, including:
[0065] The HDFS includes a data I / O module, a data backup and recovery module, a distributed file storage module, and an anomaly detection module;
[0066] The data I / O module is used to receive data input from the client and return access data or stored mapping data processing results to the client;
[0067] The data backup and recovery module is used to create multiple redundant copies of data and adopt a dual-machine hot standby solution for data security storage and data recovery;
[0068] The distributed file storage module is used to provide file storage space, provide file organization and management, and organize and manage multimodal mapping data;
[0069] The anomaly detection module is used to monitor the operating status of multiple nodes in the Hadoop cluster and perform anomaly processing.
[0070] Specifically, if Figure 2 As shown, the HDFS used in this embodiment of the present invention is used as the overall distributed file storage system for multimodal mapping data, including the following modules:
[0071] The data I / O module is responsible for receiving data input from the client and returning the data the client needs to access or the mapping data processing results stored in HDFS. The data I / O module supports the input, output, and display of common spatial data formats such as SHP (Open Spatial Data Format), TIFF (Tagged Image File Format), and OSGB (Oblique Photography Binary Data Format).
[0072] The data backup and recovery module is responsible for creating multiple redundant copies of data, providing secure data storage and reliable recovery mechanisms. In this embodiment, an HDFS dual-host hot standby design is used. In this design, two coexisting name nodes are set up: one in active state and the other in standby state. If the active name node fails, the standby name node automatically starts up immediately.
[0073] The distributed file storage module provides file storage space, stores mapping data files in this file system, and manages and organizes files. Multimodal mapping data is divided by data dimension, with different management methods for 2D and 3D data, which are then stored as HBase metadata.
[0074] The anomaly detection module manages the entire distributed storage system and monitors the operating status and load of each node in the Hadoop cluster. It uses Prometheus and Grafana to monitor the operating status of nodes in the Hadoop distributed cluster. Prometheus, an open-source node monitoring framework, monitors each node's CPU utilization, load, read / write status, and disk usage. Grafana, the visualization component of the node monitoring dashboard, provides a more visually appealing and intuitive anomaly monitoring interface, allowing the anomaly detection module to easily identify and address anomalies.
[0075] Based on the above embodiment, the multimodal mapping data is organized and managed based on HDFS, and a multi-level file directory is used to store the multimodal mapping data, including:
[0076] Read two-dimensional mapping files or three-dimensional mapping files through the data I / O module;
[0077] Using a quadtree structure to split the two-dimensional mapping file to obtain a quadtree file, and using an octree structure to split the three-dimensional mapping file to obtain an octree file;
[0078] The quadtree file or the octree file is organized according to a pyramid structure and stored in HDFS.
[0079] The method of using a quadtree structure to split the two-dimensional mapping file to obtain a quadtree file includes:
[0080] Obtaining the center point position of the two-dimensional mapping file;
[0081] Determine the maximum block size in HDFS;
[0082] If it is determined that the data file block to be divided is larger than the maximum block size, the image of the two-dimensional mapping file is divided into four equal parts based on the two-dimensional coordinate axis with the center point position as the origin, and the quadtree node number of the divided data block is obtained;
[0083] The quadtree file is output according to the quadtree node number.
[0084] The octree structure is used to segment the three-dimensional mapping file to obtain an octree file, including:
[0085] Calculate the outer bounding box of the three-dimensional mapping file and obtain the center point position of the outer bounding box;
[0086] Determine the maximum block size in HDFS;
[0087] If it is determined that the data file block to be divided is larger than the maximum block size, traversing all three-dimensional space coordinate positions in the three-dimensional mapping file, comparing the corresponding position relationship between each three-dimensional space coordinate position and the center point position, and obtaining the child node number of the corresponding position relationship;
[0088] The octree file is output according to the child node number.
[0089] Specifically, the present invention adopts different management processes for two-dimensional mapping files and three-dimensional mapping files.
[0090] (1) Two-dimensional mapping files are organized and managed using a quadtree structure, including:
[0091] First, read the two-dimensional mapping file through the data I / O module;
[0092] Then, the maximum block size (blockSize) in HDFS is used as the split threshold, which is set to 64M. If the size of the split data file block is larger than 64M, quadtree splitting is required.
[0093] Then use the quadtree structure to preliminarily split the 2D mapping file and obtain the position of the center point of the 2D mapping file (X mid ,Y mid ), and based on this point as the origin, the image is divided into four equal parts by axis division. In the process of segmentation, {1,2,3,4} is used as the quadtree node number of the segmented data block. The method of quadtree space division for a two-dimensional plane is as follows:
[0094]
[0095] num node =2 1 *b y +2 0 *bx
[0096] For the split file blocks, the complexity of their internal data is determined, and recursive splitting is continued for those with high complexity until the data complexity of each file block is less than the split threshold.
[0097] The quadtree files are organized according to the pyramid structure and stored in the HDFS distributed file system.
[0098] (2) The 3D mapping files are organized and managed using an octree structure, including:
[0099] First, read the 3D mapping file through the data I / O module;
[0100] Then, the maximum block size (blockSize) in HDFS is used as the split threshold, which is set to 64M. If the size of the split data file block is larger than 64M, octree splitting is required.
[0101] Then use the octree structure to preliminarily segment the 3D mapping file, calculate the outer bounding box of the corresponding 3D mapping data, and obtain the center point coordinates of the outer bounding box (X mid ,Y mid ,Z mid ). Next, traverse all the three-dimensional space coordinate positions (X, Y, Z) in the three-dimensional mapping data, and determine the corresponding positional relationship between its coordinates and the center point, so as to determine the sub-node number in the three-dimensional space of the corresponding previous node. The judgment method is to judge by the size of the corresponding three-dimensional coordinates and the size of the center point coordinates. If the corresponding coordinate is greater than the center point coordinate, it takes 1, and if it is less than, it takes 0. Then, the result obtained by comparing the three-dimensional coordinates is cross-combined to obtain a three-digit data, which is actually the binary code of the sub-node number at the corresponding position. The binary code is decoded to obtain the corresponding sub-node number, and it is stored in the corresponding sub-node three-dimensional data block. The specific formula is as follows:
[0102]
[0103] num node =2 2 *b z +2 1 *b y +2 0 *b x
[0104] For the split file blocks, the complexity of their internal data is determined, and recursive splitting is continued for those with high complexity until the data complexity of each file block is less than the split threshold.
[0105] The octree files are organized in a pyramid structure and stored in the HDFS distributed file system.
[0106] The present invention adopts different organization and management models to organize and manage different types of surveying and mapping data, and adopts a multi-level file directory method to store them, thereby realizing unified data organization and management.
[0107] Based on the above embodiment, abnormality monitoring is performed on multiple storage nodes in HDFS, and a dual-machine hot standby operation mode of the multiple storage nodes is determined, including:
[0108] An anomaly monitoring system is constructed using the service monitoring framework Prometheus and the dashboard graphic editor Grafana, and the anomaly monitoring system performs anomaly monitoring on the multiple storage nodes;
[0109] A reliable coordination system ZooKeeper is used as a cluster coordinator to build a dual-machine hot standby mechanism, and a dual-machine hot standby operation mode of the multiple storage nodes is executed based on the dual-machine hot standby mechanism.
[0110] Wherein, Prometheus and Grafana are used to build an anomaly monitoring system, and the anomaly monitoring system performs anomaly monitoring on the multiple storage nodes, including:
[0111] Setting an active master node and a standby master node, wherein the active master node is a master node in a normal operating state, and the standby master node is a master node that takes over the work of the active master node;
[0112] Prometheus uses the core component Prometheus Server to regularly pull data from the indicator exposer of the cluster node, and uses Push Gateway to transfer data that cannot be pulled directly;
[0113] PromQL language is used to determine the alarm rules, and node abnormality alarm information is generated based on the alarm rules;
[0114] The monitoring data is preprocessed using the PromQL language to obtain preprocessed data, which is then transmitted to Grafana for visualization.
[0115] Among them, ZooKeeper is used as a cluster coordinator to build a dual-machine hot standby mechanism, and the dual-machine hot standby operation mode of the multiple storage nodes is executed based on the dual-machine hot standby mechanism, including:
[0116] Register named nodes in the ZooKeeper cluster and obtain the session identifier of each named node;
[0117] If it is determined that any named node fails, the session identifier is switched to an expired state and a failover is initiated;
[0118] Using the exclusive lock mechanism, it is determined that only one of the two naming nodes is active, and the standby naming node obtains the exclusive lock.
[0119] The JournalNode cluster mechanism is adopted to share data between naming nodes. The JournalNode cluster pulls data from the active naming node and saves the data in real time, and the standby naming node synchronizes data from the JournalNode cluster in real time.
[0120] Specifically, since the Hadoop architecture implements the construction of a cluster using multiple low-cost devices, during the HDFS management of cluster storage, machine nodes within the cluster may fail or have access anomalies, which may result in client access being blocked and user services being unable to be completed in a timely manner. For the Hadoop cluster node structure provided by the present invention, an open source monitoring system and a visualization interface are built using the Prometheus and Grafana frameworks to monitor and visualize the status of all cluster nodes, allowing managers to more intuitively observe the operating status of data nodes. For the reliability guarantee of the master node, the present invention adopts a dual-machine hot standby strategy, setting two master nodes, one of which is in an active state and serves as the master node under normal operating conditions, while the other is in a standby state. When the master node fails or is abnormal, it takes over the master node to work.
[0121] like Figure 3 As shown in the figure, the Prometheus framework uses the core component, Prometheus Server, to regularly pull metrics data exposed by the cluster node's exporters. This component includes three components: Retrieval, TSDB, and HTTP Server. For data sent by applications that cannot be directly pulled by the server, a Push Gateway is used to transfer Push metric data from short-term tasks and allow the server to pull it, thus enabling monitoring of metric data from all cluster nodes.
[0122] Furthermore, the present invention uses PromQL language to customize alarm rules, analyzes indicator data, and generates alarm indications when nodes are found to be abnormal according to the alarm rules. On the other hand, the monitoring data is pre-processed using PromQL language and visualized by Grafana.
[0123] For node anomalies detected by Prometheus, the present invention adopts automatic standby named node activation and switching based on HDFS dual-machine hot standby design. Figure 4 As shown, its specific implementation is:
[0124] ZooKeeper is used as the cluster coordinator. Each name node, including the primary and backup NameNodes, is registered in the ZooKeeper cluster and obtains a session identifier. If a name node fails, the session identifier expires and a failover is initiated. Furthermore, the ZooKeeper cluster's exclusive lock mechanism ensures that only one of the two name nodes is active. During a failover, the backup name node acquires the ZooKeeper cluster's exclusive lock.
[0125] Furthermore, to ensure data consistency and synchronization between the two name nodes, the present invention uses a JournalNode cluster mechanism to share data between the name nodes (sharing the primary NameNode data). The JournalNode cluster pulls and stores data from the primary name node in real time. At the same time, the backup name node synchronizes data from the JournalNode cluster in real time, thereby achieving data synchronization between the two name nodes. The DataNode cluster reports the data block status to the primary NameNode and the backup NameNode.
[0126] The following describes the multimodal mapping data distributed storage system provided by the present invention. The multimodal mapping data distributed storage system described below and the multimodal mapping data distributed storage method described above can refer to each other.
[0127] Figure 5 Schematic diagram of the structure of the multimodal mapping data distributed storage system provided by the embodiment of the present invention. Figure 5 As shown, it includes: a construction module 51, a management module 52 and a monitoring module 53, wherein:
[0128] The construction module 51 is used to use HDFS to store multimodal mapping data; the management module 52 is used to organize and manage the multimodal mapping data based on HDFS, and use a multi-level file directory to store the multimodal mapping data; the monitoring module 53 is used to monitor the multiple storage nodes in HDFS for abnormalities and determine the dual-machine hot standby operation mode of the multiple storage nodes.
[0129] Figure 6 An example of a physical structure diagram of an electronic device is shown below. Figure 6As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 may call logic instructions in the memory 630 to execute a distributed storage method for multimodal mapping data, which includes: using HDFS to store multimodal mapping data; organizing and managing the multimodal mapping data based on HDFS, and using multi-level file directories to store the multimodal mapping data; and performing abnormal monitoring on multiple storage nodes in the HDFS to determine a dual-machine hot standby operation mode for the multiple storage nodes.
[0130] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0131] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the distributed storage method for multimodal mapping data provided by the above methods, which includes: using HDFS to store multimodal mapping data; organizing and managing the multimodal mapping data based on HDFS, and using a multi-level file directory to store the multimodal mapping data; performing abnormal monitoring on multiple storage nodes in HDFS, and determining the dual-machine hot standby operation mode of the multiple storage nodes.
[0132] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is implemented to execute the distributed storage method for multimodal mapping data provided by the above-mentioned methods. The method includes: using HDFS to store multimodal mapping data; organizing and managing the multimodal mapping data based on HDFS, and using a multi-level file directory to store the multimodal mapping data; performing abnormal monitoring on multiple storage nodes in HDFS, and determining the dual-machine hot standby operation mode of the multiple storage nodes.
[0133] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0134] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0135] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A distributed storage method for multimodal mapping data, characterized in that: include: Use Hadoop distributed file system HDFS to store multimodal mapping data; Organizing and managing the multimodal mapping data based on HDFS, and using a multi-level file directory to store the multimodal mapping data; Perform abnormal monitoring on multiple storage nodes in HDFS and determine the dual-machine hot standby operation mode of the multiple storage nodes; The multimodal mapping data is organized and managed based on HDFS, and a multi-level file directory is used to store the multimodal mapping data, including: Read two-dimensional mapping files or three-dimensional mapping files through the data I / O module; Using a quadtree structure to split the two-dimensional mapping file to obtain a quadtree file, and using an octree structure to split the three-dimensional mapping file to obtain an octree file; Organizing the quadtree file or the octree file according to a pyramid structure and storing the files in HDFS; The two-dimensional mapping file is divided into a quadtree file using a quadtree structure, comprising: Obtaining the center point position of the two-dimensional mapping file; Determine the maximum block size in HDFS; If it is determined that the data file block to be divided is larger than the maximum block size, the image of the two-dimensional mapping file is divided into four equal parts based on the two-dimensional coordinate axis with the center point position as the origin, and the quadtree node number of the divided data block is obtained; Output the quadtree file according to the quadtree node number; The octree structure is used to split the three-dimensional mapping file to obtain an octree file, including: Calculate the outer bounding box of the three-dimensional mapping file and obtain the center point position of the outer bounding box; Determine the maximum block size in HDFS; If it is determined that the data file block to be divided is larger than the maximum block size, traversing all three-dimensional space coordinate positions in the three-dimensional mapping file, comparing the corresponding position relationship between each three-dimensional space coordinate position and the center point position, and obtaining the child node number of the corresponding position relationship; The octree file is output according to the child node number.
2. The distributed storage method for multimodal mapping data according to claim 1, characterized in that: HDFS is used to store multimodal mapping data, including: The HDFS includes a data I / O module, a data backup and recovery module, a distributed file storage module, and an anomaly detection module; The data I / O module is used to receive data input from the client and return access data or stored mapping data processing results to the client; The data backup and recovery module is used to create multiple redundant copies of data and adopt a dual-machine hot standby solution for data security storage and data recovery; The distributed file storage module is used to provide file storage space, provide file organization and management, and organize and manage multimodal mapping data; The anomaly detection module is used to monitor the operating status of multiple nodes in the Hadoop cluster and perform anomaly processing.
3. The distributed storage method for multimodal mapping data according to claim 1, characterized in that: Performing abnormal monitoring on multiple storage nodes in HDFS and determining the hot standby operation mode of the multiple storage nodes includes: An anomaly monitoring system is constructed using the service monitoring framework Prometheus and the dashboard graphic editor Grafana, and the anomaly monitoring system performs anomaly monitoring on the multiple storage nodes; A reliable coordination system ZooKeeper is used as a cluster coordinator to build a dual-machine hot standby mechanism, and a dual-machine hot standby operation mode of the multiple storage nodes is executed based on the dual-machine hot standby mechanism.
4. The distributed storage method for multimodal mapping data according to claim 3, characterized in that: Prometheus and Grafana are used to build an anomaly monitoring system, which monitors the multiple storage nodes for anomalies, including: Setting an active master node and a standby master node, wherein the active master node is a master node in a normal operating state, and the standby master node is a master node that takes over the work of the active master node; Prometheus uses the core component Prometheus Server to regularly pull data from the indicator exposer of the cluster node, and uses Push Gateway to transfer data that cannot be pulled directly; PromQL language is used to determine the alarm rules, and node abnormality alarm information is generated based on the alarm rules; The monitoring data is preprocessed using the PromQL language to obtain preprocessed data, which is then transmitted to Grafana for visualization.
5. The distributed storage method for multimodal mapping data according to claim 3, characterized in that: ZooKeeper is used as a cluster coordinator to build a dual-machine hot standby mechanism, and a dual-machine hot standby operation mode of the multiple storage nodes is executed based on the dual-machine hot standby mechanism, including: Register named nodes in the ZooKeeper cluster and obtain the session identifier of each named node; If it is determined that any named node fails, the session identifier is switched to an expired state and a failover is initiated; Using the exclusive lock mechanism, it is determined that only one of the two naming nodes is active, and the standby naming node obtains the exclusive lock. The JournalNode cluster mechanism is adopted to share data between naming nodes. The JournalNode cluster pulls data from the active naming node and saves the data in real time, and the standby naming node synchronizes data from the JournalNode cluster in real time.
6. A multimodal mapping data distributed storage system, based on the multimodal mapping data distributed storage method according to any one of claims 1 to 5, characterized in that: include: Build a module for storing multimodal mapping data using HDFS; A management module, configured to organize and manage the multimodal mapping data based on HDFS, and store the multimodal mapping data in a multi-level file directory; The monitoring module is used to monitor multiple storage nodes in the HDFS for abnormalities and determine the dual-machine hot standby operation mode of the multiple storage nodes.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the multimodal mapping data distributed storage method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Cloud platform data organization and retrieval method for 3D (three-dimensional) urban building data
CN103955511A
Data storage system based on Hadoop architecture
CN107800808A