Precise Query Method for Theater Data Based on HBase Distributed Storage System

By using the Thrift interface and PARTITION algorithm in the HBase distributed storage system for direct connection query, and combining the iterative computing advantages of the Spark platform, a large number of complex data storage and query are solved, and efficient and accurate data query and processing are achieved.

CN114328610BActive Publication Date: 2025-06-10ZHEJIANG UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111675934.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-06-10
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The prior art lacks the storage consideration of large amounts of data and the optimization processing of query methods, resulting in complex data query.

Method used

The HBase distributed storage system is used for data storage, and HBase is accessed using C++ through the Thrift interface, and direct connection query is performed in combination with the PARTITION algorithm to reduce data processing time. At the same time, Spark distributed computing platform is used to perform calculation-intensive and large-data operations to achieve efficient query.

Benefits of technology

It realizes high-reliability and high-performance data storage and query, and builds large-scale storage clusters through horizontal scale, improves user random reading and access efficiency, reduces query cost, and achieves accurate and efficient data query and processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114328610B_ABST
    Figure CN114328610B_ABST
Patent Text Reader

Abstract

The present invention discloses an optimized method for precise query of theater data based on HBase, comprising the following steps: 1) accessing HBase using C++ through Thrift; 2) using Thrift to generate a file writing interface and placing the HBase.thritf file in the source code of HBase under the home directory of the root user; 3) accessing HBase in the program; opening the service of HBase thrift2service on the server side, and writing an interface program by calling the API with reference to the function definitions in the generated file. The client applies to the server for accessing the database and calls the thrift function to open the connection to the database; 4) performing a direct connection operation using the PARTITION algorithm to reduce data processing time; 5) using the Spark distributed computing platform for compute-intensive and large-volume data operations, and realizing efficient query of data through the iterative computing advantage of the Spark platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of storage and query optimization of theater data, and particularly relates to a precise query optimization method for theater data based on HBase. Background Art

[0002] There is information indicating that for the management of specific theater data, there are a large number of employees and actors in the theater, and there are various types and a large number of performing arts equipment. Therefore, these data need to be stored in a database in a standardized manner, and the data in the database can be quickly queried to facilitate managers to perform operations such as adding, deleting, or modifying these data in a timely manner, so as to ensure the timeliness and stability of daily management and performances. Therefore, the distributed storage system Hbase is used to store the data based on columns, and the direct connection query strategy is selected to provide precise data management services.

[0003] Chinese Patent Document CN111914155A discloses a "query conversion system and its method". The query steps include: establishing a query conversion system based on keyword matching; the user sends a fuzzy query request to the service query server through the client, and the fuzzy query request includes the query statement input by the user; the query request conversion module receives the fuzzy query request; the service query rule matching unit retrieves the service query rules in the service query rule database and matches each keyword; the query statement generation unit traverses the query logical connectors and connects the parameter items and query condition values corresponding to all the obtained keywords to generate an exact query statement; the query request execution module retrieves the exact query statement and performs query processing on the exact query statement. The above technical solution lacks the consideration of the need to store a large amount of data and the optimization processing of the query method, making the data more complex to query. Summary of the Invention

[0004] The present invention mainly solves the problem that the original technical solution lacks the consideration of the need to store a large amount of data and the optimization processing of the query method, making the data more complex to query. The present invention provides a precise query method for theater data based on the Hbase distributed storage system. The HBase platform can randomly access data, and both writing and reading data can be completed according to requirements. It can fully store unstructured data, allowing a flexible and dynamic data model. During the storage process, it does not care about the data type. At the same time, when performing connection operations in a local area network or a dedicated line network, the entire relationship is transmitted from the local site to another site, and the optimization for this is direct connection query optimization, reducing the cost during query and achieving precise and efficient data query and processing.

[0005] The above technical problems of the present invention are mainly solved by the following technical solutions: The present invention includes the following steps:

[0006] S1 accesses HBase using C++ through Thrift;

[0007] S2 uses Thrift to generate a file writing interface;

[0008] S3 accesses HBase in the program;

[0009] S4 adopts the PARTITION algorithm to perform a direct join operation on two or more relations, reducing data processing time:

[0010] S5 uses the Spark distributed computing platform for compute-intensive and large-data volume operations. Through the iterative computing advantage of the Spark platform, efficient data query can be achieved.

[0011] Preferably, the step S1 specifically includes:

[0012] S1.1 Write the HBase.thrift file to define the HBase data structure and service interface;

[0013] S1.2 Use the Thrift code generator to generate several codes that conform to the agreed communication format;

[0014] S1.3 Use the thrift application framework, including the library functions provided by thrift itself, other third-party support libraries, and the generated codes to access the database. The HBase installation package comes with the HBase.thrift file for generating API codes. Using the Thrift code generator requires using the Thrift software.

[0015] Preferably, in the step S1.2, the user only needs to declare his own service in Thrift, and the corresponding code block will be automatically generated through the compilation function of Thrift. Finally, the user only needs to implement the coding according to the mode of "the client calls the service and the server calls the service". The Thrift underlying architecture includes the following parts:

[0016] 1) protocol protocol layer: defines the data transmission format. The protocol is the rules, standards or agreements established for data exchange in the network. This protocol can be used for communication between entities in different systems;

[0017] 2) transport transport layer: defines the data transmission method, which can be TCP / IP transmission, memory sharing or file sharing;

[0018] 3) processor server processing layer: encapsulates the operations of reading data from the input data stream and writing data to the data data stream. The read and write data streams are represented by protocol objects.

[0019] Preferably, step S2 specifically includes:

[0020] S2.1 Place the HBase.thrift file in the HBase source code in the home directory of the root user. This file defines how to access HBase through the Thrift interface. Using the HBase.thrift file, the C++ files and lib libraries required to access HBase can be generated.

[0021] S2.2 Place the generated series of files in the ~ / gen-cpp directory;

[0022] S2.3 Generate the gen-cpp directory in the home directory of root. View this directory, and the C++ files required to access HBase will be seen.

[0023] Preferably, Thrift generates a total of 7 files in step S2.3, 3.h files and 4.cpp files, namely: HBase_constants.cpp, HBase.cpp, HBase_server.skeleton.cpp, HBase_types.h, HBase_constants.h, HBase.h, HBase_types.cpp; Among the above 7 files, except for HBase_server.skeleton.cpp, the rest of the files are relevant when writing HBase application programs in C++. Copy them to the project directory.

[0024] Preferably, step S3 specifically includes opening the service of HBase thrift2 service on the server side, and referring to the function definitions in the generated files to call the API to write the interface program. The client applies to the server to access the database and calls the thrift function to open the connection to the database.

[0025] Preferably, in step S4, in order to reduce the data processing time, a direct connection method is adopted, and the PARTITION algorithm is adopted in this method. For a join operation involved in a distributed query, the PARTITION algorithm can divide two or more relations into non-overlapping fragments on a certain join attribute, and distribute these fragments to a group of processing sites, and copy the other relations involved in the query to these sites as well. In this way, the query can be processed in parallel on these sites.

[0026] Preferably, in step S5, Spark is a computing framework based on distributed memory. Due to its good memory computing characteristics, it is suitable for operations with multiple iterations. When facing compute-intensive and large-data-volume operations, Spark can effectively utilize the repeated iteration of memory to achieve very good performance. Spark has strong advantages in processing big data with its diversified computing performance. The iterative computing advantage of the Spark platform can be used to achieve efficient data query.

[0027] N represents the average time to complete a query within the Spark platform, as shown in formula (1). Num_Result is the number of tuples covered by the query output, T Spark represents the total time to complete the query. The Spark cluster has three nodes, T Ave_Spark is used to save the performance of one node, as shown in formula (2):

[0028]

[0029] T Ave_Spark = T Spark ×Num_Node (2)

[0030] The beneficial effects of the present invention are as follows: The Hbase distributed storage system is adopted to achieve highly reliable, high-performance, column-oriented, and scalable storage of unstructured and semi-structured loose data. Through horizontal expansion, cheap servers are used to build a large-scale storage cluster to achieve efficient random reading and access efficiency for users; the PARTITION algorithm is adopted for direct connection to reduce data processing time, and finally, the iterative computing advantage of the Spark platform can be used to achieve efficient data query. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is the flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The technical solution of the present invention will be further specifically described below through embodiments and in conjunction with the drawings.

[0033] Embodiment:

[0034] An optimized method for accurate query of theater data based on HBase, as Figure 1 shown, includes the following steps:

[0035] S1 Access HBase using C++ through Thrift;

[0036] S1.1 Write the HBase.thrift file to define the HBase data structure and service interface;

[0037] S1.2 Use the Thrift code generator to generate several pieces of code that conform to the agreed communication format;

[0038] S1.3 Use the thrift application framework, including the library functions provided by thrift itself, other third-party support libraries, and the generated code to access the database. The HBase installation package comes with the HBase.thrift file for generating API code, and the Thrift software is required to use the Thrift code generator.

[0039] Users only need to declare their services in Thrift, and the corresponding code blocks will be automatically generated through the compilation function of Thrift. Finally, users only need to implement the coding according to the mode of "the client calls the service, and the server calls the service". The underlying architecture of Thrift includes the following parts:

[0040] 1) protocol protocol layer: defines the data transmission format. The protocol is the rules, standards, or agreements established for data exchange in the network, and this protocol can be used for communication between entities in different systems;

[0041] 2) transport transport layer: defines the data transmission method, which can be TCP / IP transmission, memory sharing, or file sharing;

[0042] 3) processor server processing layer: encapsulates the operations of reading data from the input data stream and writing data to the data data stream. The read and write data streams are represented by protocol objects.

[0043] S2 Use Thrift to generate a file write interface;

[0044] S2.1 Place the HBase.thritf file in the HBase source code in the home directory of the root user. This file defines how to access HBase through the Thrift interface. With the HBase.thrift file, the C++ files and lib libraries required to access HBase can be generated.

[0045] S2.2 Place the generated series of files in the ~ / gen-cpp directory;

[0046] S2.3 Generate the gen-cpp directory under the root's home directory. View this directory and you will see the C++ files required to access HBase. Thrift generates a total of 7 files, 3.h files and 4.cpp files, namely: HBase_constants.cpp, HBase.cpp, HBase_server.skeleton.cpp, HBase_types.h, HBase_constants.h, HBase.h, HBase_types.cpp. Among the above 7 files, except for HBase_server.skeleton.cpp, the rest of the files are relevant when writing HBase application programs in C++. Copy them to the project directory.

[0047] S3 Access HBase in the program;

[0048] Open the service of HBase thrift2 service on the server side. By referring to the function definitions in the generated files, call the API to write the interface program. The client applies to the server to access the database and calls the thrift function to open the connection to the database.

[0049] S4 Adopt the direct connection method to reduce the data processing time. The PARTITION algorithm is adopted in this method;

[0050] For a join operation involved in a distributed query, the PARTITION algorithm can divide two or more relations into disjoint segments on a certain join attribute and distribute these segments to a group of processing sites, and copy the other relations involved in the query to these sites as well. In this way, the query can be processed in parallel on these sites.

[0051] The PARTITION algorithm realizes the partitioning of two or more relations and distributes them to several sites for parallel processing, thus reducing the communication cost and local processing cost of the query. The main process of the PARTITION algorithm is to first determine a value privot, which is the basis for dividing the array. Those greater than or equal to privot are on the right side of privot, and those less than or equal to privot are on the left side of privot. Finally, just return the final position where privot is located. According to this algorithm principle, the algorithm scans from left to right to partition the array, defines a pointer position pos, and during the scanning process, pos points to the position of the element that does not meet the condition and needs to be exchanged. Once it is found that an exchange is needed, continuously exchange the elements. Among them, the function variable privot_posisiton represents the initial position of the privot element in the array.

[0052] S5 uses the Spark distributed computing platform for computationally intensive operations with large amounts of data; Spark is a distributed memory-based computing framework, and due to its good memory computing characteristics, it is suitable for operations with multiple iterations. When faced with computationally intensive operations with large amounts of data, Spark can effectively utilize the repeated iterations of memory to achieve very good performance. Spark has strong advantages in facing big data processing with its diverse computing capabilities. Through the iterative computing advantages of the Spark platform, efficient query of data can be achieved.

[0053] N represents the average time taken to complete a query within the Spark platform, as shown in formula (3). Num_Result is the number of tuples covered by the query output, and T Spark represents the total time taken to complete the query. The Spark cluster has three nodes, and T Ave_Spark is used to save the performance of one node, as shown in formula (4):

[0054]

[0055] T Ave_Spark = T Spark × Num_Node (4)

[0056] The specific embodiments described in this article are merely illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

[0057] Although terms such as HBase distributed storage system and direct connection are used more frequently in this article, the possibility of using other terms is not excluded. The use of these terms is only to more conveniently describe and explain the essence of the present invention; interpreting them as any additional limitation is contrary to the spirit of the present invention.

Claims

1. An optimized method for accurate query of theater data based on HBase, characterized in that, it includes the following steps; 1) Use C++ to access HBase through Thrift: 1.1) Write the HBase.thrift file to define the HBase data structure and service interface; 1.2) Use the Thrift code generator to generate several codes that conform to the agreed communication format; 1.3) Use the thrift application framework, including the library functions provided by thrift itself, third-party support libraries, and the generated codes to access the database; the HBase installation package comes with the HBase.thrift file for generating API codes, and the Thrift software is required to use the Thrift code generator; 2) Use the Thrift-generated file write interface: 2.1) Place the HBase.thritf file in the HBase source code in the home directory of the root user. This file defines how to access HBase through the Thrift interface. Use the HBase.thrift file to generate the C++ files and lib libraries required to access HBase; 2.2) Place the generated files in the ~ / gen-cpp directory; 2.3) Generate the gen-cpp directory in the home directory of root. View this directory and you will see the C++ files required to access HBase; 3) Access HBase in the program: 3.1) Open the service of the HBase thrift2 service on the server side. By referring to the function definitions in the generated files, call the API to write the interface program. The client applies to the server to access the database and calls the thrift function to open the connection to the database; 4) Adopt the PARTITION algorithm to perform a direct join operation on two or more relationships to reduce data processing time: The PARTITION algorithm divides two or more relationships into non-overlapping fragments on a certain join attribute, distributes the divided fragments to a group of processing sites, and copies the other relationships involved in the query to the processing sites as well, so that the query can be processed in parallel on the processing sites; 5) Use the Spark distributed computing platform to perform computationally intensive and large-data volume operations. Through the iterative computing advantage of the Spark platform, efficient query of data can be achieved; N represents the average time taken for query completion within the Spark platform, as shown in formula (3), Num_Result is the number of tuples covered by the query output, T Spark represents the total query completion time. The Spark cluster has three nodes, T Ave_Spark is used to save the performance of a single node, as shown in formula (4): T Ave_Spark = T Spark × Num_Node(4).

2. The optimized method for accurate query of theater data based on HBase according to claim 1, characterized in that, the underlying architecture of the Thrift includes the following parts: 1) The protocol layer: defines the data transmission format; the protocol is a rule, standard, or convention established for data exchange in the network, and this protocol can be used for communication between entities in different systems; 2) The transport layer: defines the data transmission method; it is TCP / IP transmission, memory sharing, or file sharing; 3) Processor server processing layer: Encapsulates the operations of reading data from the input data stream and writing data to the data data stream. The read and write data streams are represented by protocol objects.

Citation Information

Patent Citations

  • Keyword matching-based query conversion system and keyword matching-based query conversion method

    CN111914155A

  • HBase-based big data storage and retrieval method and system

    CN104915450A

  • Hadoop-based mass log data processing method

    CN106709003A