A community data management method and system based on HBase
By using a hybrid compression algorithm based on column feature clustering and a clustered index storage model in the HBase table, massive community data are compressed and indexed, which solves the problem of data discrete and granularity selection, and efficient data compression and query are achieved.
Patent Information
- Application Number
- CN202111416610.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-11-25
AI Technical Summary
The existing data compression algorithms have difficulties in data discreteness and granularity selection in massive community data management, and have failed to fully utilize the characteristics of classification compression to optimize query indexes.
The hybrid compression algorithm based on column feature clustering is used to compress the data in the HBase table, and the clustering center point is stored as an index in the HBase table. The index is written to disk or cache according to the access frequency through the cluster index storage model.
It effectively solves the problems of data discrete and granularity selection, improves data compression efficiency and query efficiency, and meets the real-time reading and writing needs of massive community data.
Smart Images

Figure CN114253966B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data management, and in particular to a community data management method and system based on HBase. Background Art
[0002] The statements in this section merely provide background art related to the present invention and do not necessarily constitute prior art.
[0003] Data management technology mainly refers to the technology of modeling real objects and the relationships between objects in the real world, and using computer-related technologies to read and write the models. Its significance lies in digitizing the characteristic information of objects and providing data and business support for large-scale service applications. The granularity of data and management strategies determine the precision and efficiency of related service applications. Early data management technologies were usually based on relational database technology, storing the characteristic information of real objects in fixed models for operation. However, with the development of computer technology and the growth of data volume, relational databases are no longer suitable for the management of massive and complex data, so non-relational database-related management technologies have emerged.
[0004] Community data mainly includes semi-structured data such as community residents, real estate, community vehicles, nearby points of interest, and unstructured data such as community surveillance videos and images. Therefore, relational databases for structured data are not suitable for community service scenarios, while HBase database is a typical representative of non-relational databases. HBase manages data in the form of columns and can adapt to complex and changing data structures. Compared with relational databases, HBase can ensure higher efficiency and real-time performance in specific application scenarios such as random queries, multi-table associations, and historical data analysis. In addition, since HBase data is stored in columns, the stored data in the same column often has the same form and a limited range of feature values, which is more suitable for data compression than row-based data. The purpose of data compression is to organize the data in the same column so that the data distribution in the column is more dense, thereby saving space and increasing the data transmission rate.
[0005] However, existing data compression algorithms often require prior knowledge for subjective classification, and overly detailed classification methods often lead to higher computational complexity. At the same time, the algorithm does not fully utilize the characteristics of classification compression to optimize query indexes. The data compression efficiency of HBase depends on the data preprocessing method and compression granularity, and the index establishment method affects the efficiency of data query. Summary of the invention
[0006] In order to address the deficiencies of the prior art, the present invention provides a community data management method and system based on HBase, which solves the problems of data discreteness and granularity selection in the data compression process, realizes the management of massive and diversely structured community service platform related data, and enables community service related applications to call community data in real time.
[0007] In order to achieve the above object, the present invention adopts the following technical solution:
[0008] A first aspect of the present invention provides a community data management method based on HBase.
[0009] A community data management method based on HBase includes the following processes:
[0010] Obtain community data and store relevant data in the HBase table according to the structure of community data;
[0011] A hybrid compression algorithm based on column feature clustering is used to compress the data in the HBase table, and the cluster center points generated during the compression process are stored as indexes in the HBase table.
[0012] The cluster index storage model is used to write the cluster center point index to disk or cache according to the access frequency.
[0013] Furthermore, the structure of the community data includes semi-structured data and unstructured data;
[0014] For semi-structured data, store the columns or column clusters of the semi-structured data into HBase tables;
[0015] For unstructured data, the unstructured data is converted into binary files and stored in the HDFS file system. At the same time, a keyword contact table related to the binary file characteristics is established and stored in the HBase table.
[0016] Furthermore, the keyword association table includes two parts: a file name column and a characteristic keyword column cluster;
[0017] The file name column stores the extension of the binary file;
[0018] The feature keyword column cluster stores features that the binary file contains or may contain.
[0019] Furthermore, the specific steps of the hybrid compression algorithm based on column feature clustering are:
[0020] Divide the data in the HBase table into multiple sub-tables by column;
[0021] Encode the data in the subtable and calculate feature similarity;
[0022] A hierarchical clustering algorithm is used to cluster the data in the subtable.
[0023] Furthermore, the cluster index storage model also uses a consistent hash algorithm to optimize the cache address location process, specifically:
[0024] The consistent hash algorithm organizes the entire hash value space into a ring area to continuously store the cluster center point index and the column data around it, and the address of the cache server is also arranged in the ring area;
[0025] When the index to be retrieved is located in the annular area through the Hash algorithm, the first server encountered in the clockwise direction is the server where the index to be retrieved should be located, so that the index to be retrieved and the column data in the surrounding continuous space are accessed in the server.
[0026] Furthermore, it also includes allocating the query task to multiple threads to obtain multiple subtasks; each subtask accesses the corresponding index through the cluster index storage model according to the data batch allocated to it, and then reads the column data stored continuously around the index; finally, the query results of multiple subtasks are merged into the total query result.
[0027] A second aspect of the present invention provides a community data management system based on HBase.
[0028] A community data management system based on HBase, comprising:
[0029] The community data storage module is configured to: obtain community data, and store relevant data into an HBase table according to the structure of the community data;
[0030] A data compression module is configured to: compress the data in the HBase table using a hybrid compression algorithm based on column feature clustering, and store the cluster center points generated during the compression process as indexes in the HBase table;
[0031] The cluster index storage module is configured to: adopt a cluster index storage model to write the cluster center point index into a disk or cache according to the access frequency.
[0032] Furthermore, it also includes a query module, which is configured to: assign query tasks to multiple threads to obtain multiple subtasks; each subtask accesses the corresponding index through the cluster index storage model according to the data batch assigned to it, and then reads the column data stored continuously around the index; finally, the query results of multiple subtasks are merged into the total query result.
[0033] Furthermore, the structure of the community data includes semi-structured data and unstructured data;
[0034] For semi-structured data, store the columns or column clusters of the semi-structured data into HBase tables;
[0035] For unstructured data, it is converted into binary files and stored in the HDFS file system. At the same time, a keyword contact table related to the binary file characteristics is established and stored in the HBase table.
[0036] Furthermore, the keyword association table includes two parts: a file name column and a characteristic keyword column cluster;
[0037] The file name column stores the extension of the binary file;
[0038] The feature keyword column cluster stores features that the binary file contains or may contain.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] 1. The HBase-based community data management method and system described in the present invention improve the data compression process and index-based query process of HBase in order to make up for the shortcomings of HBase, and propose a hybrid compression algorithm based on column feature clustering and a cluster index storage model, which can highly meet the requirements of high dynamics and real-time performance, thereby effectively managing massive and diversely structured community data, and providing efficient and real-time data docking and business support for community service applications.
[0041] 2. The HBase-based community data management method and system described in the present invention adopts a hybrid compression algorithm based on column feature clustering to compress column data and distribute column data with similar feature values in a continuous storage space. It can organize discrete column data in a clustered manner with appropriate granularity to save storage space.
[0042] 3. The HBase-based community data management method and system described in the present invention stores the cluster center points generated in the clustering process to the disk or cache according to the access frequency for the compressed columns. Through the cluster center point index, the system user can obtain column data similar to the center point feature value. Compared with the general query index, the query method based on the cluster center point index has a shorter search length and can meet the real-time requirements of the system. It can realize efficient and refined search of data without prior knowledge, thereby improving the query accuracy and efficiency.
[0043] 4. The HBase-based community data management method and system described in the present invention adopts a cluster index storage model to manage cluster center indexes in order to optimize the access efficiency and storage efficiency of the index itself, making the storage and call of cluster center indexes more efficient.
[0044] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0046] Figure 1 This is a flowchart of a community data management method based on HBase according to Embodiment 1 of the present invention;
[0047] Figure 2 A conceptual diagram of a hierarchical clustering algorithm according to Embodiment 1 of the present invention;
[0048] Figure 3 This is a conceptual diagram of the consistent hash algorithm of Example 1 of the present invention. DETAILED DESCRIPTION
[0049] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0050] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0051] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0052] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0053] Embodiment 1:
[0054] like Figure 1As shown, Example 1 of the present invention provides a community data management method based on HBase, which makes full use of the characteristics of various semi-structured and unstructured data contained in the community data, optimizes the compression storage process and query process of the data, and proposes a hybrid compression algorithm based on column feature clustering and a clustering index storage model. Among them, the hybrid compression algorithm based on column feature clustering first splits the columns in the data table into different tables, and uses a hierarchical clustering algorithm to aggregate columns with similar features in a continuous storage space according to the feature distribution in each column, and at the same time integrates the cluster center points generated by the clustering process into a multi-level tree structure. Thereafter, the clustering index storage model is used to store the formed cluster center points according to their access frequency to the disk or cache, and the consistent Hash algorithm is used to optimize the cache address positioning process to improve the real-time performance of the query. A community data management method based on HBase provided in this embodiment not only inherits the advantages of HBase, but also optimizes the compression storage and index establishment process, making it more suitable for real-time reading and writing of massive and diversely structured community data. Specifically comprising the following steps:
[0055] Step 1: Use HBase to build a community database.
[0056] Obtain community data and store relevant data in HBase tables according to the structure of community data. Community data mainly includes semi-structured data such as community residents, real estate, community vehicles, and nearby points of interest, as well as unstructured data such as community surveillance videos and images. Specifically:
[0057] For semi-structured data, use HBase table structure to store its columns or column clusters; that is, for semi-structured data, establish a corresponding HBase table for known column attributes, and set up an extensible column cluster on this basis. The function of this column cluster is to manage changeable data columns; store the basic feature columns and extended feature column clusters of semi-structured data in HBase tables.
[0058] For unstructured data such as videos and images, the whole is stored in the form of binary files, and its keyword association is established in the HBase table; that is, for unstructured data, it is usually converted into binary files and stored in the HDFS file system, and a keyword association table related to the binary file features is established at the same time; the keyword association table is also stored using the HBase table structure, where the table structure of the keyword association table includes two parts: the file name column and the feature keyword column cluster; the file name column stores the extension of the binary file, and the feature keyword column cluster stores the features that the binary file contains and may contain, such as file size and type, and the column cluster can be deleted or expanded according to the actual situation. The file name column and the extended feature column cluster of structured data are stored in the HBase table.
[0059] After the above steps, the community data from the community property system are stored in an appropriate manner according to their structure.
[0060] Step 2: Use the hybrid compression algorithm based on column feature clustering to compress the data in the HBase table, and at the same time store the cluster center points generated during the compression process as indexes in the HBase table. The overall process of the hybrid compression algorithm based on column feature clustering is as follows:
[0061] (1) Divide the data in the HBase table by columns and create a table. For each table divided by columns, that is, divide the data table in the community database into multiple sub-tables by columns to prepare for clustering compression; the composite row key of the sub-table is represented in the form of (columnID_rowID_row-key), where columnID represents the position of the original column, rowID represents the new position of the row data in the sub-table, and row-key represents the position of the original row; in addition, after the composite row key column of the sub-table is the column value to be divided; take the community resident data table as an example, as shown in Table 1 and Table 2, assuming that one of the divided sub-tables is the name table, the structure of the sub-table is (composite row key column, name column).
[0062] Table 1 Example table
[0063] row-key col:name col:gender col:age … 1fg45kh Li Ming male 34 … 3sad723 Wang Lili female 22 … … … … … …
[0064] Table 2 Subtable
[0065]
[0066]
[0067] (2) After the data table is divided into sub-tables, a suitable clustering algorithm is used for each sub-table to cluster the column data with similar eigenvalues, and the clustered column data is stored in a continuous space for easy retrieval. Specifically, it includes:
[0068] (201) Encode the data in the sub-table;
[0069] Since the clustering algorithm performs clustering based on the similarity between feature values, it is necessary to encode the data in the subtable according to the data type of each column to calculate the feature similarity. In the application scenario involved in the present invention, for non-restricted string data such as addresses and links, prefix encoding in simple dictionary encoding is usually used to convert them, while date, time or other data types with small spacing are usually converted using incremental encoding.
[0070] (202) After encoding the data, the data in the sub-table is clustered, and an appropriate criterion function is used to calculate the similarity of the eigenvalues during the clustering process.
[0071] For columns with discrete eigenvalues, a criterion function should be set according to actual conditions to measure the similarity between eigenvalues (such as "red" and "pink"); the criterion function takes the encoding corresponding to the two eigenvalues and the tolerable difference as input, and outputs the similarity between the two eigenvalues.
[0072] As an implementation method, a hierarchical clustering algorithm is used to cluster the column data. Hierarchical clustering gradually merges or splits the data points in the subtable by calculating the similarity between the data points to generate a nested clustering tree. Compared with other forms of clustering algorithms, hierarchical clustering is not sensitive to the measurement method of similarity, so it is suitable for clustering many different types of data points, and its time complexity is relatively small. In addition, Figure 2 As shown in the figure, the hierarchical clustering algorithm can integrate the cluster center points generated by the clustering process into a tree structure, so the cluster index storage model can directly store the cluster center point index in the form of a tree, which speeds up the index calling process. Since the initial column data usually exists in a discrete form and prior classification knowledge is often lacking, the top-down clustering form is not suitable for accurate clustering of column data.
[0073] Therefore, the hybrid compression algorithm based on column feature clustering proposed in the present invention uses a hierarchical clustering method to cluster the column data in the subtable from bottom to top. In each iteration process, several points or clustering areas with similarity reaching a preset value are clustered, and this process is repeated after a new clustering area is formed until the clustering area covers the entire column data. In each iteration process, the center point of the generated clustering area will be stored as an index in the HBase table for calling. Through the cluster center point index, the points in the clustering area can be quickly retrieved.
[0074] Step 3: Design a cluster index storage model to optimize the storage and call process of the cluster center index. The cluster center index generated in step 2 will be written to disk or memory according to the access frequency for reading. The purpose of the cluster index storage model is to improve the efficiency of index call during the query process.
[0075] In order to optimize the storage and call process of the cluster center point index generated in step 2 to improve the query efficiency, the present invention particularly proposes a cluster index storage model to improve the speed of index access. In step 2, the cluster center point index itself is associated in a tree structure, similar to B-tree, and its own retrieval rate is higher than that of the general index storage model, but it does not fully and effectively utilize the hardware resources of the community database. Although the clustered column data is stored in the continuous space around the cluster center point according to the similarity, the storage location of the cluster center point itself is random. Therefore, in order to further improve the retrieval efficiency, the cluster center point index is transferred to the disk or cache according to the access frequency. Among them, the disk uses a persistent distributed storage method, and the cache uses a consistent Hash algorithm for address mapping. The disk mainly performs long-term persistent storage of the cluster center point index, while the cache stores the index with a higher access frequency, that is, the center point index with a higher access frequency is stored in the cache, and the low center point index is stored in the disk. In addition, if the corresponding index is not retrieved in the cache during the access process, the retrieval result in the disk is transferred to the cache for access.
[0076] Cache servers can speed up the index access process, but the storage requirements of community databases are not static. When the number of cache servers changes, the address of the index will change, and the large amount of cache data stored in it needs to be re-mapped through the Hash algorithm again, and the computing pressure of the server will increase dramatically. Therefore, the present invention applies the consistent Hash algorithm to the mapping process of the index address in the cache server. Figure 3As shown in the figure, the consistent hash algorithm organizes the entire hash value space into a ring area to continuously store the cluster center point index and the column data around it, and the address of the cache server is also arranged in the ring area. When the index to be retrieved is located in the ring area through the hash algorithm, the first server it encounters in the clockwise direction is the server where the current index should be located, so the current index and the column data in the surrounding continuous space are accessed in this server. If the cache server changes, only the index stored in the changed server needs to be relocated, and the data in other servers will not be affected. In summary, the consistent hash algorithm shows good fault tolerance and scalability in the cache service area, and is suitable for the storage of cluster center point indexes.
[0077] Step 4: Use parallel mechanism to optimize the query process. The query mechanism based on single thread cannot meet the real-time query requirements of massive data, so it is necessary to divide the query task into multiple threads to perform a comprehensive search of the HBase database.
[0078] In step 4, it is necessary to use a parallel mechanism to optimize the query process to meet the real-time requirements of the query. The present invention distributes the query tasks from community service applications to multiple threads to obtain multiple subtasks, and each subtask accesses the corresponding index through the clustered index storage model according to the data batch assigned to it, and then reads the column data stored continuously around the index. Finally, the query results of multiple subtasks are merged into the total query results and fed back to the community service application via the corresponding external interface. In summary, the community database based on HBase and the community service application complete the business docking with high real-time performance.
[0079] Through the hybrid compression algorithm based on column feature clustering and the cluster index storage model proposed by the present invention, community data is stored in a continuous space according to the similarity of feature values, and the access process is accelerated by an efficient attribute index management mechanism. While saving storage space, the efficiency of data transmission is improved, and real-time dynamic business support can be provided for community service applications. The HBase-based community data management method of the present invention can accept query requests from community service applications, allocate query tasks in the form of multi-threaded parallelism, and finally summarize the retrieval results of each query subtask and feed them back to the community service application.
[0080] Embodiment 2:
[0081] Embodiment 2 of the present invention provides a community data management system based on HBase, including a community data storage module, a data compression module, a cluster index storage module and a query module.
[0082] The community data storage module is configured to: obtain community data, and store relevant data into an HBase table according to the structure of the community data;
[0083] A data compression module is configured to: compress the data in the HBase table using a hybrid compression algorithm based on column feature clustering, and store the cluster center points generated during the compression process as indexes in the HBase table;
[0084] The cluster index storage module is configured to: adopt a cluster index storage model to write the cluster center point index into a disk or cache according to the access frequency.
[0085] The HBase-based community database is connected to community service applications, and appropriate interfaces are used to implement business support for applications in different development environments. Finally, the HBase-based community data management system is completed.
[0086] The query module is configured to: assign query tasks from community service applications to multiple threads to obtain multiple subtasks, and each subtask accesses the corresponding index through the clustered index storage model according to the data batch assigned to it, and then reads the column data stored continuously around the index; finally, the query results of multiple subtasks are merged into the total query result, which is fed back to the community service application via the corresponding external interface.
[0087] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A community data management method based on HBase, characterized in that: The process includes: Obtain community data and store relevant data in the HBase table according to the structure of community data; The structural forms of the community data include semi-structured data and unstructured data; For semi-structured data, store the columns or column clusters of the semi-structured data into HBase tables; For unstructured data, convert the unstructured data into binary files and store them in the HDFS file system. At the same time, establish a keyword contact table related to the binary file features and store the keyword contact table in the HBase table. A hybrid compression algorithm based on column feature clustering is used to compress the data in the HBase table, and the cluster center points generated during the compression process are stored as indexes in the HBase table. The specific steps of the hybrid compression algorithm based on column feature clustering are: Divide the data in the HBase table into multiple sub-tables by column; Encode the data in the subtable and calculate feature similarity; A hierarchical clustering algorithm is used to cluster the data in the subtable from bottom to top; The cluster index storage model is used to write the cluster center point index to disk or cache according to the access frequency; The cluster index storage model also uses a consistent hash algorithm to optimize the cache address location process. Specifically: The consistent hash algorithm organizes the entire hash value space into a ring area to continuously store the cluster center point index and the column data around it, and the address of the cache server is also arranged in the ring area; When the index to be retrieved is located in the annular area through the Hash algorithm, the first server encountered in the clockwise direction is the server where the index to be retrieved should be located, so that the index to be retrieved and the column data in the surrounding continuous space are accessed in the server.
2. The HBase-based community data management method according to claim 1, characterized in that: The keyword contact table includes two parts: a file name column and a characteristic keyword column cluster; The file name column stores the extension of the binary file; The feature keyword column cluster stores features that the binary file contains or may contain.
3. The HBase-based community data management method according to claim 1, characterized in that: The method also includes allocating the query task to multiple threads to obtain multiple subtasks; each subtask accesses the corresponding index through the cluster index storage model according to the data batch allocated to it, and then reads the column data stored continuously around the index; and finally the query results of multiple subtasks are merged into the total query result.
4. A community data management system based on HBase using the method of claim 1, characterized in that: include: The community data storage module is configured to: obtain community data, and store relevant data into an HBase table according to the structure of the community data; A data compression module is configured to: compress the data in the HBase table using a hybrid compression algorithm based on column feature clustering, and store the cluster center points generated during the compression process as indexes in the HBase table; The cluster index storage module is configured to: adopt a cluster index storage model to write the cluster center point index into a disk or cache according to the access frequency.
5. A community data management system based on HBase as claimed in claim 4, characterized in that: It also includes a query module, which is configured to: assign query tasks to multiple threads to obtain multiple subtasks; each subtask accesses the corresponding index through the cluster index storage model according to the data batch assigned to it, and then reads the column data stored continuously around the index; finally, the query results of multiple subtasks are merged into the total query result.
6. A community data management system based on HBase as claimed in claim 4, characterized in that: The structural forms of the community data include semi-structured data and unstructured data; For semi-structured data, store the columns or column clusters of the semi-structured data into HBase tables; For unstructured data, it is converted into binary files and stored in the HDFS file system. At the same time, a keyword contact table related to the binary file characteristics is established and stored in the HBase table.
7. A community data management system based on HBase as claimed in claim 6, characterized in that: The keyword contact table includes two parts: a file name column and a characteristic keyword column cluster; The file name column stores the extension of the binary file; The feature keyword column cluster stores features that the binary file contains or may contain.
Citation Information
Patent Citations
Cloud management database establishment method and system for health behavior intervention
CN110504031A