Massive Ship Historical Trajectory Data Storage System and Query Method

Through the combination of distributed architecture and local index structure, the problems of backward data storage models and inefficient query efficiency are solved for massive ship historical trajectory, and efficient data storage and query optimization are achieved.

CN115905242BActive Publication Date: 2025-07-08NAVAL UNIV OF ENG PLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211659538.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-22
Publication Date
2025-07-08
Estimated Expiration
2042-12-22

AI Technical Summary

Technical Problem

In the prior art, the storage model of ship historical trajectory data is backward and the query efficiency is inefficient, which cannot meet the efficient storage and real-time query requirements of massive data.

Method used

A massive ship historical trajectory data storage system was designed, and a distributed architecture was adopted, combining B+ trees, R trees and hash tables to build a local index structure. The query process was optimized through parallel query methods to reduce communication overhead between nodes and realize the rapid search of time, space and ship identification.

Benefits of technology

It improves the storage efficiency and query speed of massive ship historical trajectory data, reduces query delay, and meets the real-time query needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115905242B_ABST
    Figure CN115905242B_ABST
Patent Text Reader

Abstract

A mass storage system for ship historical trajectory data designed by the present invention includes a trajectory storage module, a local index module, and a data maintenance module. According to the characteristics of the ship trajectory data structure, the orderly organization and storage of ship historical trajectory data are completed. Combining B+ trees, R trees, and hash tables according to common query types, a local index structure supporting multiple query types is constructed. At the same time, the local index and the corresponding data storage partition are maintained in the same node, reducing communication overhead. Based on the constructed storage and index structure, a query method is implemented based on a parallel query method to optimize the queries based on time, space, and ship identification. Since the model of storing the same ship data in the same node effectively reduces the communication overhead between nodes, and the row key and index structure ensure the rapid search for these three types of keywords, the query latency can be effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ship data processing, and specifically to a storage system and query method for a large amount of historical ship trajectory data. Background Art

[0002] In recent years, with the rapid development of the economy, the number of ships, ship traffic volume, and ship traffic density have increased rapidly, the maritime traffic environment has become increasingly complex, and the frequency of maritime traffic accidents has also increased. To reduce the occurrence of maritime traffic accidents, more and more ships are equipped with the Automatic Identification System (AIS). AIS is a new type of digital navigation aid system and equipment that integrates network technology, modern communication technology, computer technology, and electronic information display technology. It broadcasts ship dynamic information such as ship position, speed, course change rate, and course, as well as ship static information such as ship name, call sign, draft, and dangerous goods, to nearby ships and shore stations via very high frequency radio waves in cooperation with GPS, enabling nearby ships and shore stations to promptly grasp the dynamic and static information of all ships on the sea surface, thereby ensuring maritime traffic safety.

[0003] Today, hundreds of thousands of ships globally have installed AIS, and the monthly AIS data volume globally can reach several hundred GB. Among them, each AIS message recording position information can be regarded as a trajectory point of a ship. Therefore, these accumulated AIS data contain a large amount of historical ship trajectory data. Through the analysis and mining of historical ship trajectory data, good solutions can be provided for route optimization, abnormal behavior monitoring, false target identification, port throughput statistical analysis, etc.

[0004] Effective analysis and mining of a large amount of historical ship trajectory data cannot be achieved without the support of efficient storage and query technologies. However, traditional centralized storage and query solutions have been unable to meet the query requirements of a large amount of historical ship trajectory data due to deficiencies in computing, reading and writing, communication, and scalability, and have the following problems:

[0005] 1. The storage model is backward. Existing historical ship trajectory data is usually stored in a relational database or a relational database based on spatial expansion, deployed on a single computer. Since the capacity of a single computer is limited and it is difficult for traditional relational databases to be expanded, it is not suitable for large-scale data storage.

[0006] 2. The query efficiency is low. Due to the large scale and diverse types of existing historical ship trajectory data, it is difficult for traditional query methods to implement parallel query means, with low efficiency and unable to meet real-time query requirements.

[0007] As can be seen from the above content, the backward storage model and low query efficiency are the two major problems to be solved in the current methods for storing and querying ship historical trajectory data. Summary of the Invention

[0008] The purpose of the present invention is to provide a storage system and a query method for massive ship historical trajectory data, so as to solve the problems of backward storage model and low query efficiency faced in the storage and query of ship historical trajectory data, and to achieve efficient storage and query of massive ship historical trajectory data.

[0009] To achieve this purpose, the massive ship historical trajectory data storage system designed by the present invention includes a trajectory storage module, a local index module, and a data maintenance module. Among them, the trajectory storage module includes trajectory storage partitions distributed in several nodes of the cluster, and each trajectory storage partition stores the trajectory data of several ships;

[0010] The local index module includes several local index partitions, and all local index partitions are stored in memory. Each local index partition corresponds to a trajectory storage partition, and the local index partition and the trajectory storage partition with a corresponding relationship are stored in the same node;

[0011] Each local index partition includes a spatio-temporal index and a ship identification time index. During the query process, the spatio-temporal index locates the query range according to the spatio-temporal query keyword, and the ship identification time index locates the query range according to the ship identification keyword and the time keyword; the data maintenance module is used to synchronously process the splitting and migration of the trajectory storage partition and the corresponding local index partition;

[0012] When a certain trajectory storage partition is split due to the amount of stored data exceeding the storage threshold to form a new trajectory storage partition, the data maintenance module is used to synchronously split the corresponding local index partition into new local index partitions, so that the new local index partitions correspond to the new trajectory storage partitions; when the trajectory storage partition is migrated from one node to another node, the data maintenance module synchronously migrates the corresponding local index partition to the same computer.

[0013] A method for parallel querying of massive ship historical trajectory data using the above system includes the following steps:

[0014] Step 1: Send the query time keyword qtk, the query space keyword qsk, and the query ship identification set qms to each local index partition of the local index module. When the user initiates a query, at least one of the three query conditions, namely the query time keyword qtk, the query space keyword qsk, and the query ship identification set qms, is not empty. If the query time keyword qtk is empty, the query covers the entire time range; if the query space keyword qsk is empty, the query covers the entire space range; if the query ship identification set qms is empty, the query covers all ships.

[0015] Step 2: Start parallel query. The query process in each local index partition is as follows:

[0016] Step 201: For each local index partition, first query the spatio-temporal index through the query time keyword qtk and the query space keyword qsk to obtain a set of candidate row keys for the query result, crks0, crks0 = {rk 0,1 , rk 0,2 , …, rk m,n}, where rk m,n represents the row key of the trajectory segment generated by the ship with MMSI m in the time interval ti n .

[0017] Step 202: Query the ship identification time index through the query time keyword qtk and the query ship identification set qms to obtain a set of candidate row keys for the query result, crks1, crks1 = {rk 0,1 , rk 0,2 , …, rk i,j}, where rk i,j represents the row key of the trajectory segment generated by the ship with MMSI i in the time interval ti j .

[0018] Step 203: Calculate the intersection candidate row key set crks2 of the set of candidate row keys crks0 and the set of candidate row keys crks1, and obtain the set of candidate row keys mcrks0 through the set of candidate row keys crks2.

[0019] Step 204: Sequentially read each element of the set of candidate row keys mcrks0. For the current element mcrk g , query the corresponding trajectory storage partition of the local index partition through the candidate row key belonging to mcrk g to obtain the trajectory data of the rows corresponding to all candidate row keys in the set of candidate row keys mcrks0. Then filter the trajectory data through the query time keyword qtk and the query space keyword qsk to obtain the query result of the partition.

[0020] Step 3: Aggregate the query results of each partition obtained to get a query result set.

[0021] Advantages of the present invention:

[0022] Aiming at the problems of backward storage model and low query efficiency existing in the use of the existing ship historical trajectory data storage device and query method, according to the characteristics of the ship trajectory data structure, the present invention completes the orderly organization and storage of ship historical trajectory data, and combines the B+ tree, R tree and hash table according to common query types to construct a local index structure supporting multiple query types. At the same time, the local index and the corresponding data storage partition are maintained in the same node, reducing communication overhead. Based on the constructed storage and index structure, a query method is implemented based on a parallel query method, realizing the optimization of queries based on time, space and ship identification. Since the model of storing the same ship data in the same node effectively reduces the communication overhead between nodes, and the row key and index structure ensure the rapid search for these three types of keywords, the query latency can be effectively reduced. Description of the drawings

[0023] Figure 1 It is a schematic structural diagram of a massive ship historical trajectory data storage system of the present invention.

[0024] Among them, 1 - trajectory storage module, 2 - local index module, 3 - data maintenance module, 4 - trajectory storage partition; 5 - local index partition; 6 - spatio-temporal index; 7 - ship identification time index. Specific embodiments

[0025] The following further elaborates on the present invention in detail with reference to the drawings and specific embodiments:

[0026] As Figure 1 shown in the massive ship historical trajectory data storage system, it includes a trajectory storage module 1, a local index module 2 and a data maintenance module 3. Among them, the trajectory storage module 1 includes trajectory storage partitions 4 distributed in several nodes of the cluster (the cluster is composed of multiple computers connected together, and a node refers to a single computer in the cluster), and each trajectory storage partition 4 stores the trajectory data of several ships;

[0027] The local index module 2 includes several local index partitions 5. All local index partitions 5 are stored in the memory. Each local index partition 5 corresponds to a trajectory storage partition 4. "Corresponds to" means that one local index partition only indexes the data in one trajectory storage partition, and they are on the same node. It can be said that there is a one-to-one correspondence between them. Since they are stored on the same node, there is no need to cross nodes during communication, which can reduce the overhead of packet encapsulation, network transmission, and packet parsing required for cross-node communication. The local index partition 5 and the trajectory storage partition 4 with a corresponding relationship are stored in the same node;

[0028] Each local index partition 5 includes a spatio-temporal index 6 and a vessel identification time index 7. During the process of querying vessel historical trajectory data, the spatio-temporal index 6 locates the spatio-temporal range of the query according to the spatio-temporal query keywords, and the vessel identification time index 7 locates the query time range and vessel identification according to the vessel identification keywords and time keywords; The data maintenance module 3 is used to synchronously process the splitting and migration of the trajectory storage partition 4 and the corresponding local index partition 5 (if the trajectory data written into the partition is too much, resulting in the partition occupying more space than the threshold, which is not conducive to system maintenance at this time, then the over-large partition needs to be split into two, and one of the split partitions is moved to other nodes in the cluster to ensure the load balance of the entire cluster);

[0029] When a certain trajectory storage partition 4 is split due to the stored data volume exceeding the storage threshold (try to ensure that the storage capacities of the new partitions are the same during splitting, and it is required that the data of the same vessel after splitting is on one partition), forming new trajectory storage partitions 4, the data maintenance module 3 is used to synchronously split the corresponding local index partition 5 into new local index partitions 5, so that the new local index partitions 5 correspond to the new trajectory storage partitions 4; When the trajectory storage partition 4 is migrated from one node to another node, the data maintenance module 3 synchronously migrates the corresponding local index partition 5 to the corresponding node. Moving a partition from one node to another to ensure load balance. Migration occurs after splitting. The migration method is to migrate the partition to the node with the smallest load after calculating the load of each node in the computing cluster. When the index and data partition are on one node, the processing overhead caused by cross-node query can be reduced during query, and the query efficiency can be improved.

[0030] In the above technical solution, the cluster is a group of computers that work together either loosely or tightly (loosely means that the smallest data unit stored in a node is a file; tightly means that the smallest data unit in a node is a data block, that is, a file is split into multiple data blocks and then stored on different nodes). A single computer in the cluster is regarded as a node, and the nodes are connected through a local area network. The purpose of the cluster is to coordinate multiple computers to complete tasks and improve work efficiency through the method of distributed parallelism. Compared with a single computer, multiple computers have stronger capabilities in terms of computing and storage when working together.

[0031] In the above technical solution, the trajectory storage module 1 is implemented based on the non-relational database HBase. The data in HBase is stored in tables, and a table consists of rows and columns. Columns of the same type can form column families. A separate HBase table Tab t is used to store ship trajectory data. The HBase table Tab t is divided into several trajectory storage partitions distributed in the cluster. Each trajectory storage partition stores the trajectory data of several ships. The trajectory data of the same ship is continuously and orderly stored in a trajectory storage partition 4 in the form of trajectory segments. Continuously means that the data of the same ship is continuously distributed in the storage space without inserting the data of other ships in the middle. Orderly means that different trajectory data of the same ship is stored in chronological order;

[0032] The trajectory segment is the continuous and orderly ship trajectory data generated by a ship in a certain time interval. The time interval is obtained by equally dividing the time dimension, and the length of all time intervals is 2 hours. In a certain time interval ti i the trajectory segment TS i,j ={p1, p2, …, p k} means that the sequence of sampled trajectory points of ship sh j in the time interval ti i The trajectory points p1, p2, …, p k are sorted in chronological order;

[0033] The trajectory segment is uniquely identified by combining the ship identification and the time interval attribute. Since the HBase table Tab t uses the row key (Row Key) to determine the location of the data in the storage partition, when the trajectory data is stored in the form of trajectory segments, the row key rk is implemented in the form of combining the ship identification mmsi and the time interval ti attribute:

[0034] rk = mmsi + ti

[0035] Among them, the mmsi (Maritime Mobile Service Identify, MMSI) is the identification code for maritime mobile communication services and can be regarded as the identification of a ship. ti represents the time interval. Since the mmsi information in rk is used as a prefix, according to the characteristics of the dictionary sorting of HBase row keys, the trajectory data of the same ship will be stored continuously and orderly in Tab t ;

[0036] Cooperating with the design of the row key in the HBase table Tab t In terms of column families and columns, one column family is used to store all the trajectory points in the trajectory segment. Each column in the column family uses the sampling time of the trajectory point as the column key to ensure that the trajectory points in the trajectory segment are stored in chronological order. The spatial position of the trajectory point and other information in the AIS data except for MMSI, time, longitude, and latitude (such as course, speed, destination, etc.) are recorded in the corresponding data unit of the HBase table. The data unit is a general data storage structure in HBase. A data unit can uniquely determine a cell through the row key, column family, and column key. This design can reduce the communication overhead and sorting overhead across nodes during query. Since the data of the same ship is continuously stored in the same partition, only one node needs to be queried to query the data of one ship; when querying the data of multiple ships, since the data of one ship is not on multiple nodes, there is no need to coordinate multiple nodes to aggregate the data, which can reduce the communication overhead. In addition, since the data of the same ship is stored in an orderly manner, there is no need to re-sort the results after querying.

[0037] In the above technical solution, the spatio-temporal index 6 in the local index partition 5 adopts a hierarchical hybrid structure, which is divided into upper and lower layers. The upper layer is a classification index based on time periods. The time periods are obtained by dividing the time dimension into several equal-length time segments. Each time period contains several time intervals, and the length of each time period is 24 hours. The lower layer is a trajectory segment index, which consists of several R-trees. These R-trees use the trajectory segments as the index objects. Each R-tree in the lower layer corresponds to a time period in the upper layer. Therefore, each R-tree only needs to maintain the trajectory segments within the time period;

[0038] In order to be able to apply the indexing of ship trajectory data, in the spatio-temporal index 6, the R-tree not only indexes the spatial location attributes, but also includes the time attribute within the indexing scope. The index entries of the classification index in the spatio-temporal index 6 are recorded in the form of a binary tuple A(tc, rtree), where tc represents the time period of the index entry, and rtree represents a pointer to the corresponding lower-level R-tree. The intermediate nodes of the R-tree in the spatio-temporal index 6 are recorded in the form of a triple A(tp, mbr, rcns), where tp represents the time range of the node, mbr (Minimum Bounding Rectangle) represents the spatial minimum bounding rectangle of the node, and rcns is a set of pointers to child nodes. The leaf nodes of the R-tree in the spatio-temporal index 6 are recorded in the form of a triple B(tp, mbr, tses), where tp represents the time range of the node, mbr represents the spatial minimum bounding rectangle of the node, and tses is a set of pointers to trajectory segment index entries. In tses, the pointers of the trajectory segment index entries are arranged in ascending order of the row key values of the trajectory segment index entries pointed to by the pointers. The trajectory segment index entry is recorded in the form of a binary tuple B(rk, mbr1), where rk is the row key of the trajectory segment, and mbr1 represents the spatial minimum bounding rectangle of the trajectory segment. Since the time interval information of the trajectory segment already exists in the row key, it is not recorded additionally. Since the spatio-temporal index 6 uses time and spatial attributes as indexing objects and is stored in memory, it has a fast access speed and can support the query of time keywords and spatial keywords.

[0039] In the above technical solution, the ship identification time index 7 in the local index partition 5 adopts a hierarchical hybrid structure, which is divided into upper and lower layers. The upper layer is the ship identification index, which adopts a hash table structure and is used to index ship identification attributes; the lower layer is the time index, which consists of several B+ trees, and each B+ tree corresponds to an upper-layer ship identification index entry and is used to index the time intervals corresponding to all trajectory segments of the corresponding ship. Since the ship identification time index 7 uses ship identification and time as indexing objects and is stored in memory, it has a fast access speed and can support the query of ship and time keywords.

[0040] The index entry of the ship identification in the ship identification time index 7 is recorded in the form of a binary tuple C(mmsi, bptree), where mmsi represents the ship identification, and bptree represents a pointer to the corresponding lower-level B+ tree. The intermediate nodes of the B+ tree in the ship identification time index 7 are recorded in the form of a binary tuple D(tp, bcns), where tp represents the time range of the node, and bcns represents a set of pointers to child nodes; this feature corresponds to the data distribution of the trajectory storage partition (MMSI is in the front and time is in the back in the row key), so the upper-layer structure of the ship identification time index 7 uses ship identification attributes as indexing objects.

[0041] The leaf nodes of the B+ tree in the ship identification time index 7 are recorded in the form of a quadruple (pre, next, tp, mes), where pre is a pointer to the previous leaf node, next is a pointer to the next leaf node, tp is the time range of the node, and mes represents the set of row keys corresponding to the trajectory segments within the node time range. This feature corresponds to the data distribution of the trajectory storage partition (MMSI comes first and time comes second in the row key). The lower-level structure uses the time attribute as the indexing object. In addition, the leaf nodes of the B+ tree are linked to each other, enabling quick implementation of time range queries.

[0042] A method for parallel query of massive ship historical trajectory data using the above system, which includes the following steps:

[0043] Step 1: Send the query time keyword qtk, the query space keyword qsk, and the query ship identification set qms to each local index partition 5 of the local index module 2. When the user initiates a query, at least one of the three query conditions, namely the query time keyword qtk, the query space keyword qsk, and the query ship identification set qms, is not empty. If the query time keyword qtk is empty, the query covers the entire time range; if the query space keyword qsk is empty, the query covers the entire space range; if the query ship identification set qms is empty, the query covers all ships. The query conditions involved in this method are ship identification, time, and space. It is possible that the user's query involves all three conditions, in which case the query keywords are directly sent to each partition to start the next step of processing. However, it is also possible that the user's query conditions are only one or two, in which case the query range of the unspecified condition needs to be set to all so that the correct query structure can be obtained in the subsequent processing steps;

[0044] Step 2: Start parallel query. The query process in each local index partition 5 is as follows:

[0045] Step 201: For each local index partition, first query the spatio-temporal index 6 through the query time keyword qtk and the query space keyword qsk to obtain the set crks0 of candidate row keys for the query result, crks0 = {rk 0,1 , rk 0,2 , …, rk m,n}, where rk m,n represents the row key of the trajectory segment generated by the ship with MMSI m in the time interval ti n ;

[0046] Step 202: Query the ship identification time index 7 through the query time keyword qtk and the query ship identification set qms to obtain the set crks1 of candidate row keys for the query result, crks1 = {rk 0,1 , rk 0,2 , …, rki,j}, where rk i,j represents the row key of the trajectory segment generated by the ship with MMSI i in the time interval ti j ;

[0047] Step 203: Calculate the intersection crks2 = crks0 ∩ crks1 of the set of candidate row keys crks0 and the set of candidate row keys crks1 to obtain the set of candidate row keys crks2, and obtain the set of candidate row keys mcrks0 through the set of candidate row keys crks2;

[0048] Step 204: Read each element of the set of candidate row keys mcrks0 in the lexicographical order (from smallest to largest in alphabetical and numerical order) of the row key values. For the current candidate row key element mcrk g , through mcrk g query the trajectory storage partition 4 corresponding to the local index partition 5, obtain the trajectory data of the rows corresponding to all candidate row keys in the set of candidate row keys mcrks0, and then filter the trajectory data of one row in the trajectory storage partition corresponding to a row key through the query time keyword qtk and the query space keyword qsk (obtain the data of all rows in the trajectory storage partition corresponding to mcrks0), to obtain the query result of the trajectory storage partition (a data set composed of multiple trajectory points. For each trajectory point in the query result, its MMSI attribute appears in qms, its time attribute is within the time range of the time keyword qtk, and its spatial position is within the range of the spatial keyword qsk);

[0049] Step 3: Aggregate the query results of each partition obtained, obtain the query result set, return it to the user, and the query ends.

[0050] In the above technical solution, the specific process of step 201 is as follows: For each local index partition 5, first query the spatio-temporal index 6 through the query time keyword qtk and the query space keyword qsk. When searching the spatio-temporal index 6, first search the classification index on the upper layer of the spatio-temporal index 6 through the query time keyword qtk. If a certain index item binary tuple A(tc, rtree) of the classification index satisfies that is, the time range of qtk intersects with the time range of the time period of the index item, then search the R tree rtree corresponding to this index item from top to bottom (starting from the root node query of the tree to the leaf node). When each intermediate node triple A(tp, mbr, rcns) of rtree is searched, if this intermediate node satisfies and That is, if the time range of qtk intersects with the time range of the intermediate node and the space range of qsk intersects with the space range of the intermediate node, then search for the children nodes of this node through the set rcns of pointers to child nodes, and so on; when each leaf node triple B(tp, mbr, tses) in the rtree is searched, if this leaf node satisfies and That is, if the time range of qtk intersects with the time range of the intermediate node and the space range of qsk intersects with the space range of the intermediate node, then search for the index entries included in this node through the set tses of pointers to trajectory segment index entries; when each index entry pair B(rk, mbr1) is searched, obtain the time interval ti of the indexed trajectory segment from the row key rk, if the index entry satisfies and then extract the row key of this index entry, regarded as a candidate row key, and regard this index entry as a candidate index entry. After the spatio-temporal index search is completed, obtain the set crks0 of candidate row keys, denoted as crks0 = {rk 0,1 , rk 0,2 , …, rk m,n}, where rk m,n represents the row key of the trajectory segment generated by the ship with MMSI m in the time interval ti n . Step 201 is to query the spatio-temporal index 6 through the time keyword and the space keyword, aiming to obtain the row keys of all trajectory segments whose time range intersects with the time keyword and whose space range intersects with the space keyword. Since the spatio-temporal index 6 is stored in memory and uses time and space attributes as index objects, it can quickly obtain the trajectory segments with spatio-temporal intersection.

[0051] In the above technical solution, the specific process of the said step 202 is: query the ship identification time index 7 through the query time keyword qtk and the query ship identification set qms. For each query ship identification mmsi q in the query ship identification set qms q , use mmsi q to search the ship identification time index 7. If the search for mmsi q hits, assume the hit index entry is (mmsi q , bptree q ), bptree q represents the lower-layer B+ tree corresponding to mmsi q in the ship identification time index 7. Then search bptree q from top to bottom through qtk. When each intermediate node (tp, bcns) of bptree That is, if the time range of qtk intersects with the time range tp of the intermediate node, search for the children of this node through the set bcns of pointers to child nodes, and so on; every time a leaf node quadruple (pre, next, tp, mes) in bptree q is found, if this leaf node satisfies That is, if the time range of qtk intersects with the time range tp of the leaf node, extract the row keys included in mes and regard them as candidate row keys. For the first leaf node whose time range intersects with the time range of qtk found by searching in a top-down manner, after completing the search of this leaf node, continue to search for the leaf nodes connected to this leaf node through the pointer pre to the previous leaf node and the pointer next to the next leaf node. Let (pre p , next p , tp p , mes p ) be the leaf node connected to (pre, next, tp, mes) through the pre pointer. pre p , next p , tp p , mes p respectively represent the pointer to the previous leaf node, the pointer to the next leaf node, the time range of the node, and the set of row keys corresponding to the trajectory segments within the node time range of the leaf node connected to (pre, next, tp, mes) through the pre pointer. Determine whether it holds. If it holds, extract the row keys included in mes p and regard them as candidate row keys, and continue to search for the connected leaf nodes through pre p , and so on, until the time ranges of the currently searched leaf nodes do not intersect; let (pre n , next n , tp n , mes n ) be the leaf node connected to (pre, next, tp, mes) through the next pointer. pre n , next n , tp n , mes n respectively represent the pointer to the previous leaf node, the pointer to the next leaf node, the time range of the node, and the set of row keys corresponding to the trajectory segments within the node time range of the leaf node connected to (pre, next, tp, mes) through the next pointer. Determine whether it holds. If it holds, extract the row keys included in mes n and regard them as candidate row keys, and continue to search for the connected leaf nodes through next nContinue to search for connected leaf nodes, and so on, until the time ranges of the currently searched leaf nodes do not intersect; when the leaf nodes searched through pre p and next n do not intersect with qtk, then stop searching the bptree q . After the search for the current query vessel identification is completed, read the next query vessel identification in the query vessel identification set qms, and complete the search process according to the same steps as above. After all elements in qms are searched, a set of candidate row keys crks1 is obtained, denoted as crks1 = {rk 0,1 , rk 0,2 , …, rk i,j}, where rk i,j represents the row key of the trajectory segment generated by the vessel with MMSI of i in the time interval ti j . Step 202 is to query the vessel identification time index 7 through the query time keyword and the query vessel identification set, aiming to obtain the row keys of all trajectory segments whose time ranges intersect with the query time keyword and whose vessel identification attributes are in the query vessel identification set. Since the vessel identification time index 7 is stored in memory and uses time and vessel identification attributes as index objects, it can quickly obtain the trajectory segments that intersect in time and have the specified vessel identification.

[0052] In the above technical solution, the specific process of obtaining the candidate row key set mcrks0 from the candidate row key set crks2 in step 203 is as follows: sort the row keys in the candidate row key set crks2 in ascending order of the row key values, and classify the row keys according to the MMSI attribute while sorting, to obtain the sorted and classified row key set mcrks0 = {mcrk0, mcrk1, …, mcrk k}, where mcrk k represents the candidate row key set of the vessel with MMSI of k, and in mcrk k , the row keys are arranged in ascending order. This design is to avoid subsequent classification and sorting processing of the query results. This is because the data in the trajectory storage partition is stored continuously and orderly. If the candidate row keys are processed as classified and ordered, then each result subset obtained by the query must belong to a vessel and is arranged in ascending order of time.

[0053] In the above technical solution, the specific process of step 204 is as follows: sequentially read each element of the selected row key set mcrks0. For the current element mcrk g , query the trajectory storage partition 4 corresponding to the local index partition 5 through the candidate row keys belonging to mcrk g , and obtain the trajectory data of the rows corresponding to all candidate row keys in the candidate row key set mcrks0, denoted as cTsg = {p g,1 , p g,2 , …, p g,f}, where p g,f represents the f-th trajectory point of the ship with MMSI g sorted in ascending order of time. Then, by querying the time keyword qtk and the spatial query keyword qsk, the trajectory points belonging to cTs g are filtered. For the trajectory point p g,d ∈ cTs g , if the sampling time of the trajectory point p g,d is not within the range of the query time keyword qtk or the spatial position of the trajectory point p g,d is not within the range of the query spatial keyword qsk, then the trajectory point p g,d is deleted from cTs g . After filtering, the trajectory query result Ts g of the ship g is obtained. Since the candidate row keys in mcrk g are ordered and the trajectory points in the trajectory storage partition 4 are stored in time series, the trajectory point query results in the trajectory query result Ts g will be automatically stored in time series. The trajectory query result Ts g is added to the partition query result Tss p . Tss p refers to the summary of all query results in a trajectory storage partition, which is composed of the trajectory query results Ts g of multiple ships. Continue to access the next element of mcrks0 and complete the processing according to the same steps as above until all elements in mcrks0 are processed.

[0054] The content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.

Claims

1. A storage system for a large amount of historical ship trajectory data, characterized in that: It includes a trajectory storage module (1), a local index module (2) and a data maintenance module (3). Among them, the trajectory storage module (1) contains trajectory storage partitions (4) distributed in several nodes of the cluster, and each trajectory storage partition (4) stores the trajectory data of several ships; The local index module (2) includes several local index partitions (5), and all local index partitions (5) are stored in memory. Each local index partition (5) corresponds to a trajectory storage partition (4), and the local index partition (5) and the trajectory storage partition (4) with a corresponding relationship are stored in the same node; Each local index partition (5) includes a spatio-temporal index (6) and a ship identification time index (7). During the process of querying the historical trajectory data of ships, the spatio-temporal index (6) locates the spatio-temporal range of the query according to the spatio-temporal query keywords, and the ship identification time index (7) locates the query time range and ship identification according to the ship identification keywords and time keywords; the data maintenance module (3) is used to synchronously process the splitting and migration of the trajectory storage partition (4) and the corresponding local index partition (5); When a certain trajectory storage partition (4) is split due to the amount of stored data exceeding the storage threshold to form a new trajectory storage partition (4), the data maintenance module (3) is used to synchronously split the corresponding local index partition (5) into new local index partitions (5) so that the new local index partitions (5) correspond to the new trajectory storage partitions (4); when the trajectory storage partition (4) is migrated from one node to another node, the data maintenance module (3) synchronously migrates the corresponding local index partition (5) to the corresponding node; The spatio-temporal index (6) in the local index partition (5) adopts a hierarchical hybrid structure, which is divided into upper and lower layers. The upper layer is a classification index based on time periods, and the time periods are obtained by dividing the time dimension into several equal-length time segments. Each time period contains several time intervals. The lower layer is a trajectory segment index, which consists of several R-trees. These R-trees use trajectory segments as index objects, and each R-tree in the lower layer corresponds to a time period in the upper layer; In order to be able to apply the indexing of ship trajectory data, in the spatio-temporal index (6), the R-tree not only indexes the spatial location attributes, but also includes the time attribute within the indexing scope. The index entries of the classification index in the spatio-temporal index (6) are recorded in the form of a binary tuple A(tc, rtree), where tc represents the time period of the index entry, and rtree represents the pointer to the corresponding lower-level R-tree. The intermediate nodes of the R-tree in the spatio-temporal index (6) are recorded in the form of a triple A(tp, mbr, rcns), where tp represents the time range of the node, mbr represents the spatial minimum bounding rectangle of the node, and rcns is the set of pointers to child nodes. The leaf nodes of the R-tree in the spatio-temporal index (6) are recorded in the form of a triple B(tp, mbr, tses), where tp represents the time range of the node, mbr represents the spatial minimum bounding rectangle of the node, and tses is the set of pointers to the trajectory segment index entries. In tses, the pointers of the trajectory segment index entries are arranged in ascending order of the row key values of the trajectory segments pointed to by the pointers. The trajectory segment index entry is recorded in the form of a binary tuple B(rk, mbr1), where rk is the row key of the trajectory segment, and mbr1 represents the spatial minimum bounding rectangle of the trajectory segment.

2. The massive ship historical trajectory data storage system according to claim 1, wherein: The cluster is a group of computers that work together either loosely or tightly connected. A single computer in the cluster is regarded as a node, and the nodes are connected through a local area network.

3. The massive ship historical trajectory data storage system according to claim 1, characterized in that: The trajectory storage module (1) is implemented based on the non-relational database HBase. The data in HBase is stored in tables, which are composed of rows and columns. Columns of the same type can form column families. A separate HBase table Tab t is used to store ship trajectory data. The HBase table Tab t is divided into several trajectory storage partitions distributed in the cluster. Each trajectory storage partition stores the trajectory data of several ships. The trajectory data of the same ship is continuously and orderly stored in a trajectory storage partition (4) in the form of trajectory segments. Continuity means that the data of the same ship is continuously distributed in the storage space without inserting the data of other ships in the middle. Orderliness means that different trajectory data of the same ship are stored in chronological order; The trajectory segment is the continuous and ordered ship trajectory data generated by a ship in a certain time interval, and the time interval is obtained by equally dividing the time dimension. In a certain time interval ti i the trajectory segment TS i,j ={p1, p2, …, p k} refers to the sequence formed by the sampled trajectory points of the ship sh j in the time interval ti i The trajectory points p1, p2, …, p k are sorted in chronological order; The trajectory segment is uniquely identified by combining the vessel identification and the time interval attribute. Since the HBase table Tab t uses the row key to determine the location of the data stored in the storage partition. Therefore, when the trajectory data is stored in the form of trajectory segments, the row key rk is implemented in the form of combining the vessel identification mmsi and the time interval ti attribute: rk = mmsi + ti Among them, mmsi is the identification code for maritime mobile communication services, and ti represents the time interval. Since the mmsi information in rk is used as the prefix, according to the characteristics of the dictionary sorting of HBase row keys, the trajectory data of the same ship will be stored continuously and orderly in Tab t ; Cooperate with the HBase table Tab t In the design of the row key, in terms of column families and columns, one column family is used to store all the trajectory points in the trajectory segment. Each column in the column family uses the sampling time of the trajectory point as the column key to ensure that the trajectory points within the trajectory segment are stored in chronological order. The spatial position of the trajectory points and other information in the AIS data except for MMSI, time, longitude, and latitude are recorded in the corresponding data units of the HBase table.

4. The mass ship historical trajectory data storage system according to claim 1, characterized in that: The ship identification time index (7) in the local index partition (5) adopts a hierarchical hybrid structure, which is divided into upper and lower layers. The upper layer is the ship identification index, which adopts a hash table structure and is used to index the ship identification attributes; The lower layer is the time index, which consists of several B+ trees. Each B+ tree corresponds to an upper-layer ship identification index entry and is used to index the time intervals corresponding to all the trajectory segments of the corresponding ship; The index entries of the ship identification in the ship identification time index (7) are recorded in the form of a binary tuple C(mmsi, bptree), where mmsi represents the ship identification, and bptree represents the pointer to the corresponding lower-level B+ tree. The intermediate nodes of the B+ tree in the ship identification time index (7) are recorded in the form of a binary tuple D(tp, bcns), where tp represents the time range of the node, and bcns represents the set of pointers to child nodes; The leaf nodes of the B+ tree in the ship identification time index (7) are recorded in the form of a quadruple (pre, next, tp, mes), where pre is the pointer to the previous leaf node, next is the pointer to the next leaf node, tp is the time range of the node, and mes represents the set of row keys corresponding to the trajectory segments within the node time range.

5. A method for parallel querying of massive historical ship trajectory data using the system according to claim 1, which includes the following steps: Step 1: Send the query time keyword qtk, the query space keyword qsk, and the query ship identification set qms to each local index partition (5) of the local index module (2). When the user initiates a query, at least one of the three query conditions of the query time keyword qtk, the query space keyword qsk, and the query ship identification set qms is not empty. If the query time keyword qtk is empty, the query covers the entire time range; if the query space keyword qsk is empty, the query covers the entire space range; if the query ship identification set qms is empty, the query covers all ships. Step 2: Start parallel query. The query process in each local index partition (5) is as follows: Step 201: For each local index partition, first query the spatio-temporal index (6) by using the query time keyword qtk and the query space keyword qsk, and obtain a set crks0 of candidate row keys of the query result, crks0 = {rk 0,1 , rk 0,2 , …, rk m,n}, where rk m,n represents the row key of the trajectory segment generated by the ship with MMSI m within the time interval ti n . Step 202: Query the vessel identification time index (7) by querying the time keyword qtk and the set of vessel identifications qms, and obtain the set of candidate row keys of the query result crks1, crks1 = {rk 0,1 , rk 0,2 , …, rk i,j}, where rk i,j represents the row key of the trajectory segment generated by the vessel with MMSI i in the time interval ti j ; Step 203: Calculate the intersection of the set of candidate row keys crks0 and the set of candidate row keys crks1 to obtain the set of candidate row keys crks2, and obtain the set of candidate row keys mcrks0 through the set of candidate row keys crks2. Step 204: Read each element of the candidate row key set mcrks0 in the lexicographical order of the row key values. For the current candidate row key element mcrk g , query the trajectory storage partition (4) corresponding to the local index partition (5) through mcrk g to obtain the trajectory data of the rows corresponding to all candidate row keys in the candidate row key set mcrks0. Then, filter the trajectory data of a row in a row key corresponding trajectory storage partition through the query time keyword qtk and the query space keyword qsk to obtain the query result of the trajectory storage partition; Step 3: Aggregate the query results of each partition obtained to obtain a query result set.

6. The method for parallelized query of massive historical ship trajectory data according to claim 5, wherein: The specific process of the said step 201 is as follows: For each local index partition (5), first query the spatio-temporal index (6) through the query time keyword qtk and the query space keyword qsk. When searching the spatio-temporal index (6), first search the classification index at the upper layer of the spatio-temporal index (6) through the query time keyword qtk. If a certain index item binary tuple A(tc, rtree) in the classification index satisfies then search the R tree rtree corresponding to this index item from top to bottom. When each intermediate node triple A(tp, mbr, rcns) of the rtree is searched, if this intermediate node satisfies and then search the child nodes of this node through the set rcns of pointers to child nodes; when each leaf node triple B(tp, mbr, tses) in the rtree is searched, if this leaf node satisfies and then search the index items contained in this node through the set tses of pointers to trajectory segment index items; when each index item binary tuple B(rk, mbr1) is searched, obtain the time interval ti of the indexed trajectory segment from the row key rk. If the index item satisfies and then extract the row key of this index item, regarded as a candidate row key, and regard this index item as a candidate index item. After the spatio-temporal index search is completed, obtain the set crks0 of candidate row keys, denoted as crks0 = {rk 0,1 , rk 0,2 , …, rk m,n}, where rk m,n represents the row key of the trajectory segment generated by the ship with MMSI m in the time interval ti n .

7. The method for performing parallelized queries on a large amount of historical ship trajectory data according to claim 6, wherein: The specific process of step 202 is as follows: query the vessel identification time index (7) by querying the time keyword qtk and the query vessel identification set qms. For each query vessel identification mmsi in the query vessel identification set qms q , use mmsi q to search the vessel identification time index (7). If the mmsi q search hits, set the hit index entry as (mmsi q , bptree q ). Bptree q represents the lower-level B+ tree in the vessel identification time index (7) corresponding to mmsi q . Then, search bptree q in a top-down manner through qtk. Each time an intermediate node (tp, bcns) of bptree q is searched, if the intermediate node (tp, bcns) satisfies , then search the child nodes of this node through the set bcns of pointers to child nodes; each time a leaf node quadruple (pre, next, tp, mes) in bptree q is searched, if the leaf node satisfies , then extract the row key contained in mes and regard it as a candidate row key. For the first leaf node whose time range intersects with the qtk time range searched in a top-down manner, after completing the search of this leaf node, continue to search the leaf nodes connected to this leaf node through the pointer pre to the previous leaf node and the pointer next to the next leaf node. Let (pre p , next p , tp p , mes p ) be the leaf node connected to (pre, next, tp, mes) through the pre pointer. Pre p , next p , tp p , mes p respectively represent the pointer to the previous leaf node, the pointer to the next leaf node, the time range of the node, and the set of row keys corresponding to the trajectory segments within the node time range of the leaf node connected to (pre, next, tp, mes) through the pre pointer. Determine whether holds. If it holds, then extract the row key contained in mes p , regard it as a candidate row key, and continue to search the connected leaf nodes through pre p , and so on until the time range of the currently searched leaf node does not intersect; let (pre n , next n ,tp n ,mes n ) is a leaf node connected to (pre, next, tp, mes) through the next pointer, pre n ,next n ,tp n ,mes n respectively represent the pointer to the previous leaf node, the pointer to the next leaf node, the time range of the node, the set of row keys corresponding to the track segments within the node time range, and determine whether it holds. If it holds, then extract the row keys included in mes n , regarded as candidate row keys, and continue to search for connected leaf nodes through next n until the time ranges of the currently searched leaf nodes do not intersect; when the leaf nodes searched through pre p and next n do not intersect with qtk, then stop searching the bptree q . After the current query ship identification search is completed, read the next query ship identification in the query ship identification set qms. When all elements in qms are searched, obtain the set of candidate row keys crks1, denoted as crks1 = {rk 0,1 ,rk 0,2 ,…,rk i,j}, where rk i,j represents the row key of the track segment generated by the ship with MMSI i in the time interval ti j .

8. The method for parallelizing the query of massive historical ship trajectory data according to claim 7, wherein: The specific process of obtaining the candidate row key set mcrks0 from the candidate row key set crks2 in step 203 is as follows: Sort the row keys in the candidate row key set crks2 in ascending order according to the row key values, and classify the row keys according to the MMSI attribute while sorting, to obtain the sorted and classified row key set mcrks0 = {mcrk0, mcrk1, …, mcrk k}, where mcrk k represents the candidate row key set of the ship with MMSI of k, and in mcrk k , the row keys are arranged in ascending order.

9. The method for parallelizing the query of a large amount of historical ship trajectory data according to claim 8, wherein: The specific process of step 204 is as follows: sequentially read each element of the row selection key set mcrks0, and for the current element mcrk g , query the trajectory storage partition (4) corresponding to the local index partition (5) through the candidate row keys belonging to mcrk g , and obtain the trajectory data of the rows corresponding to all candidate row keys in the candidate row key set mcrks0, denoted as cTs g = {p g,1 , p g,2 , …, p g,f}, where p g,f represents the f-th trajectory point of the ship with MMSI g sorted in ascending order of time. Then, filter the trajectory points belonging to cTs g through the query time keyword qtk and the spatial query keyword qsk. For the trajectory point p g,d ∈ cTs g , if the sampling time of the trajectory point p g,d is not within the range of the query time keyword qtk or the spatial position of the trajectory point p g,d is not within the range of the query spatial keyword qsk, then delete the trajectory point p g,d from cTs g . After the filtering is completed, obtain the trajectory point query result Ts g of the ship g. Since the candidate row keys in mcrk g are ordered and the trajectory points in the trajectory storage partition (4) are stored in time series, the trajectory point query result in the trajectory query result Ts g will be automatically stored in time series. Add the trajectory query result Ts g to the partition query result Tss p , and continue to access the next element of mcrks0 until all elements in mcrks0 are processed.

Citation Information

Patent Citations

  • Method for storing and searching mass sensor data

    CN102651020A

  • Spatio-temporal data processing systems and methods

    US20140156806A1