A connection query method and system for Elasticsearch
The method and system for Elasticsearch enable connection queries through a query parsing layer and hash join algorithm, addressing the lack of SQL-style query support in Elasticsearch to enhance data association for security analysis.
Patent Information
- Application Number
- CN202211524893.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-11-30
AI Technical Summary
Elasticsearch does not support SQL-style connection queries, resulting in poor correlation capabilities in multi-source data analysis scenarios, making it difficult to meet the needs of security analysis and other fields.
A connection query method and system for Elasticsearch is designed. The RESTful client sends syntax-compliant connection query statements, uses a hash connection algorithm to process the data, and writes the results to Elasticsearch, supporting keywords and parameters such as join, on, rename, pick and max, and implementing efficient connection query for multi-source data.
It realizes the perfection and flexible expression of Elasticsearch's connection query requirements, improves the correlation capabilities of multi-source data, supports efficient data analysis operations, and can be lightweightly deployed in conjunction with existing Elasticsearch clusters.
Smart Images

Figure CN115878617B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data retrieval, and in particular, to a connection query method and system for Elasticsearch. Background Art
[0002] Nowadays, it has gradually become normal for organizations and institutions to deploy multiple types and multiple network security monitoring devices to protect their networks. Although a wide variety of monitoring devices can control all aspects of the network all the time, they also generate a huge amount of multi-source data, which brings great difficulties and challenges to network management and security analysis.
[0003] From the perspective of security analysis, the correlation of multi-source data is a necessary prerequisite step. The key clues of attacks and potential attack chains can only be discovered by security analysts after correlating multiple data sources (such as intrusion detection system alerts, firewall logs, vulnerability scan reports, etc.) and restoring a sufficiently comprehensive network state. Therefore, security analysts need a tool that not only has excellent data retrieval capabilities but also supports efficient connection queries to retrieve and analyze multi-source security data. As an efficient full-text search engine, Elasticsearch is used by many well-known organizations and institutions for retrieving massive amounts of data. However, due to its flat data storage design, Elasticsearch does not support SQL (Structured Query Language) style connection queries on data, resulting in its excellent data retrieval capabilities being difficult to be applied in fields with multi-source data analysis requirements such as security analysis. Summary of the Invention
[0004] The purpose of the present invention is to solve the problem that the prior art cannot perform connection queries through Elasticsearch and has poor correlation capabilities in multi-source data analysis scenarios; to provide a connection query method and system for Elasticsearch to solve the data correlation problem in some data analysis scenarios. The technical solutions are as follows:
[0005] A connection query method for Elasticsearch, comprising the following steps:
[0006] S1: A user sends a connection query statement that conforms to the syntax to a query statement parsing layer through a RESTful client;
[0007] S2: The query statement parsing layer parses the received connection query statement, extracts the target index and additional parameter information, and sends this information to an Elasticsearch data processing framework;
[0008] S3: The Elasticsearch data processing framework configures parameters according to the received information and retrieves data of the target index from Elasticsearch;
[0009] S4: Process the original data using the hash join algorithm and write the processing result into Elasticsearch;
[0010] S5: The user directly queries the join result from Elasticsearch through the RESTful client and performs data analysis operations supported by Elasticsearch on the join result.
[0011] Furthermore, the core keywords of the join query syntax are join, on, and index; join and index are used to specify the two target indexes to be joined; on is used to specify the common join field name of the target indexes.
[0012] Even further, the additional optional parameters of the join query syntax are rename, pick, and max; rename is used to specify another join field; pick is used to select the fields to be joined; max is used to specify the number of data matching times.
[0013] Even further, the specific process of step S4 is as follows:
[0014] S4.1: Based on the data of target index A, establish two hash tables. The first hash table is used to classify documents according to the hash code of the join field value, and the second hash table records the number of times each join field value has been matched;
[0015] S4.2: Match according to the hash code of the join field value in target index B in the two hash tables. If the match is successful, then according to the number of times of matching recorded in the second hash table and the max parameter value input by the user, determine whether to join the document arrays in the first hash table;
[0016] S4.3: During the joining process, filter and add prefix operations to the join fields according to the pick parameter input by the user;
[0017] S4.4: After the joining is completed, write the join result into a newly created index in Elasticsearch. The name of the newly created index is composed of the concatenation of the two target index names.
[0018] Even further, use the hashCode method provided by Java to calculate the hash code of the join field value to avoid too high a hash collision rate.
[0019] Further, in step S2, the query statement parsing layer receives the user's connection query statement in the form of an HTTP request, and sends the processed connection query statement to the Elasticsearch data processing framework in the JSON data format.
[0020] A connection query system for Elasticsearch includes a RESTful client, a query statement parsing layer, an Elasticsearch data processing framework, and Elasticsearch.
[0021] The RESTful client is used to send a connection query statement that conforms to the syntax to the query statement parsing layer, and is used to directly query the connection result from Elasticsearch and perform data analysis operations supported by Elasticsearch on the connection result.
[0022] The query statement parsing layer includes a syntax analysis tree and a traverser; the query statement parsing layer extracts the query statement in the HTTP request sent by the RESTful client, and performs parameter extraction and exception handling on it through the syntax analysis tree and the traverser; the parsed operation object and additional parameter data are encapsulated in JSON format and sent to the Elasticsearch data processing framework.
[0023] The Elasticsearch data processing framework includes a parameter configuration unit, a data acquisition unit, a connection algorithm unit, and a data writing unit; the parameter configuration unit performs parameter configuration according to the user input parameters encapsulated in JSON format obtained, the data acquisition unit assembles the corresponding Elasticsearch query request, and sends the query request to the Elasticsearch cluster specified by the user according to the configuration file to obtain the corresponding data; the connection algorithm unit processes the obtained original data using the hash join algorithm, and the data writing unit writes the processing result into Elasticsearch.
[0024] Compared with the prior art, the beneficial effects of the present invention are:
[0025] 1. By encapsulating an additional query statement parsing layer for Elasticsearch, the present invention solves the problem that the native syntax of Elasticsearch cannot express the connection query requirements, and realizes the perfect and flexible expression and interpretation of the connection query requirements.
[0026] 2. Based on the idea of the hash join algorithm, the present invention designs an Elasticsearch data processing framework, and realizes the efficient connection query that Elasticsearch originally does not support.
[0027] 3. The entire mechanism of the present invention is designed with lightweight deployment as the goal and can be used in conjunction with any deployed Elasticsearch cluster by modifying the configuration file, simply and efficiently realizing the function expansion of Elasticsearch. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is the overall architecture diagram of the Elasticsearch connection query method of the present invention.
[0029] Figure 2 is a schematic diagram showing the implementation idea of the hash join algorithm of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] The technical solution of the present invention will be further specifically described below through embodiments in conjunction with the drawings.
[0031] Embodiment:
[0032] A connection query method for Elasticsearch in this embodiment includes the following steps:
[0033] S1: The user sends a connection query statement that conforms to the syntax to the query statement parsing layer through a RESTful client.
[0034] The process of sending the connection query statement is as Figure 1 shown and includes the following process:
[0035] The user uses any software that can send RESTful-style HTTP requests to send an HTTP request to the specified port and interface of the host where the present invention is deployed, and the request body contains a connection query statement that conforms to the connection query syntax.
[0036] Figure 1 Legal connection query statement examples are given in
[0037] index=INDEX_NAME [pick] join index=INDEX_NAME [pick] on COLUMN[rename] [max]
[0038] The core keywords of this syntax are join, on, and index. Considering the intuitiveness, generality, and usability of the syntax, short and concise words are selected as the core keywords. join and index are used to specify the two target indexes to be joined, that is, INDEX_NAME; on is used to specify the common connection field name of the target indexes, that is, COLUMN. Based on the natural language syntax, the core keywords are given grammatical meanings that conform to their single-word semantics.
[0039] The additional optional parameters of this syntax are rename, pick, and max. To improve the usability and flexibility of the syntax, corresponding additional optional parameters are provided for specific usage scenarios. Rename is used to specify another connection field; pick is used to select the fields to be connected; max is used to specify the number of data matching times. Since the syntax function is directly related to the data to be processed, the additional optional parameters can also improve the query efficiency to a certain extent.
[0040] S2: The query statement parsing layer parses the received join query statement, extracts information such as the target index and additional parameters, and sends this information to the Elasticsearch data processing framework.
[0041] The parsing process of the query statement is as Figure 1 shown and includes the following processes:
[0042] The query statement parsing layer extracts the query statement from the HTTP request sent by the RESTful client and performs parameter extraction and exception handling on it through the syntax analysis tree and traverser. The parsed operation object and additional parameter data are encapsulated in JSON format and sent to the Elasticsearch data processing framework. This ensures the efficient transmission of data between modules.
[0043] Among them, the syntax analysis tree and traverser are automatically generated by ANTLR 4 according to the designed syntax rules. To ensure the efficient parsing of the query statement, a reasonable and easily extensible syntax branch structure is designed, which can quickly identify the statement format input by the user during parsing and achieve accurate and efficient parameter extraction.
[0044] S3: The Elasticsearch data processing framework configures the parameters according to the received information and retrieves the data of the target index from Elasticsearch.
[0045] The process of retrieving the target data is as Figure 1 shown and includes the following processes:
[0046] The Elasticsearch data processing framework configures and assembles the corresponding Elasticsearch query request according to the user input parameters encapsulated in JSON format obtained, and sends the query request to the Elasticsearch cluster specified by the user according to the configuration file to obtain the corresponding data.
[0047] S4: Use the hash join algorithm to process the original data and write the processing result into Elasticsearch.
[0048] The data connection and writing process is asFigure 1 and Figure 2 as shown in the figure, including the following steps:
[0049] Step ① and step ② are carried out simultaneously. Hash table 1 and hash table 2 are respectively established according to the data of target index A, which are used to classify documents according to the hash code of the connection field value and record the number of times each connection field value has been matched;
[0050] After steps ① and ② are executed, steps ③ and ④ are executed simultaneously. Match according to the hash code of the connection field value in target index B in hash table 1 and hash table 2. If the match is successful, judge whether to connect the document array in hash table 1 according to the number of times of matching recorded in hash table 2 and the max parameter value input by the user;
[0051] During the connection process, the connection field also needs to be filtered in combination with the pick parameter input by the user, as well as other operations such as adding prefixes;
[0052] After steps ③ and ④ are executed, the connected data, that is, the processing result of the algorithm processing module, will be generated in the memory, waiting for the data writing module to write it into Elasticsearch.
[0053] To avoid too high a hash collision rate, the hashCode method provided by Java is used to calculate the hash code of the connection field value.
[0054] To improve the data writing efficiency, the data writing module will write the data into Elasticsearch in a batch processing manner while the data is generated; on this basis, to avoid too much memory space occupied by a large number of intermediate data, the data that has been written into Elasticsearch will be deleted to release the occupied memory space.
[0055] S5: The user directly queries the connection result from Elasticsearch through the RESTful client and performs data analysis operations supported by Elasticsearch on the connection result.
[0056] The data query process is as Figure 1 shown in the figure, including the following steps:
[0057] The user directly sends a data query request to the Elasticsearch cluster through any software that can send RESTful-style HTTP requests. The format of the query index name is "original target index A-join-original target index B"; on this basis, the user can use any function supported by Elasticsearch to perform subsequent processing on the connected data, including but not limited to sorting, statistics, filtering, etc.
[0058] The solution of this embodiment solves the problem that the native syntax of Elasticsearch cannot express the requirements of join queries by encapsulating an additional query statement parsing layer for Elasticsearch, and realizes the expression and interpretation of the requirements of join queries. On this basis, based on the idea of the hash join algorithm, an Elasticsearch data processing framework is designed to realize the join query function that Elasticsearch originally does not support. The whole mechanism aims at lightweight deployment and can be used in conjunction with any deployed Elasticsearch cluster by modifying the configuration file, simply and efficiently realizing the function extension of Elasticsearch.
[0059] It should be understood that the embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
Claims
1. A connection query method for Elasticsearch, characterized in that, It includes the following steps: S1: The user sends a connection query statement that conforms to the syntax to the query statement parsing layer through a RESTful client; S2: The query statement parsing layer parses the received connection query statement, extracts the target index and additional parameter information, and sends this information to the Elasticsearch data processing framework; S3: The Elasticsearch data processing framework configures the parameters according to the received information and obtains the data of the target index from Elasticsearch; S4: Use the hash join algorithm to process the original data and write the processing result into Elasticsearch; The specific process of step S4 is as follows: S4.1: According to the data of target index A, two hash tables are established. The first hash table is used to classify documents according to the hash code of the connection field value, and the second hash table records the number of times each connection field value has been matched; S4.2: Match according to the hash code of the connection field value in target index B in the two hash tables. If the match is successful, judge whether to join the document array in the first hash table according to the number of times of matching recorded in the second hash table and the max parameter value input by the user; S4.3: During the joining process, filter the connection field and add a prefix operation according to the pick parameter input by the user; S4.4: After the joining is completed, write the joining result into a newly created index in Elasticsearch. The name of the newly created index is composed of the concatenation of the two target index names; And use the hashCode method provided by Java to calculate the hash code of the connection field value to avoid too high a hash collision rate; S5: The user directly queries the joining result from Elasticsearch through the RESTful client and performs data analysis operations supported by Elasticsearch on the joining result.
2. The connection query method for Elasticsearch according to claim 1, wherein The core keywords of the connection query syntax are join, on, and index; join and index are used to specify the two target indexes to be joined; on is used to specify the common connection field name of the target indexes.
3. The connection query method for Elasticsearch according to claim 1, wherein The additional optional parameters of the connection query syntax are rename, pick, and max; rename is used to specify another connection field; pick is used to select the fields to be joined; max is used to specify the number of data matches.
4. The connection query method for Elasticsearch according to claim 1, wherein In step S2, the query statement parsing layer receives the user's connection query statement in the form of an HTTP request and sends the processed connection query statement to the Elasticsearch data processing framework in the JSON data format.
5. A connection query system for Elasticsearch, characterized in that, It includes a RESTful client, a query statement parsing layer, an Elasticsearch data processing framework, and Elasticsearch; The RESTful client is used to send a connection query statement that conforms to the syntax to the query statement parsing layer, and is used to directly query the joining result from Elasticsearch and perform data analysis operations supported by Elasticsearch on the joining result; The query statement parsing layer includes a syntax analysis tree and a traverser; the query statement parsing layer extracts the query statement from the HTTP request sent by the RESTful client, and performs parameter extraction and exception handling on it through the syntax analysis tree and the traverser; the parsed operation object and additional parameter data are encapsulated in JSON format and sent to the Elasticsearch data processing framework; The Elasticsearch data processing framework includes a parameter configuration unit, a data acquisition unit, a connection algorithm unit, and a data writing unit; the parameter configuration unit configures parameters according to the user input parameters encapsulated in JSON format obtained, the data acquisition unit assembles the corresponding Elasticsearch query request, and sends the query request to the Elasticsearch cluster specified by the user according to the configuration file to obtain the corresponding data; The connection algorithm unit processes the obtained raw data using the hash join algorithm, and the data writing unit writes the processing result into Elasticsearch; Processing the obtained raw data using the hash join algorithm specifically means: According to the data of the target index A, two hash tables are established. The first hash table is used to classify documents according to the hash code of the join field value, and the second hash table records the number of times each join field value has been matched; Match according to the hash code of the join field value in the target index B in the two hash tables. If the match is successful, judge whether to join the document array in the first hash table according to the number of times of matching recorded in the second hash table and the max parameter value input by the user; During the joining process, filter the join field and add a prefix operation according to the pick parameter input by the user; After the joining is completed, write the joining result into a newly created index in Elasticsearch, and the name of the newly created index is formed by concatenating the names of the two target indexes; And use the hashCode method provided by Java to calculate the hash code of the join field value to avoid too high a hash collision rate.
Citation Information
Patent Citations
An HBase secondary index creation method and system based on a coprocessor
CN109165222A
A big data query optimization method based on Preto and Elasticsearch, in particular to a big data query optimization method based on Preto and Elasticsearch
CN109739882A