Data writing method, device, electronic device and storage medium
Searching and replacing the index alias in the Elasticsearch cluster through the Flink-Elasticsearch connector solves the problem of repeated writing of data processed by Flink in the Elasticsearch cluster, and reducing data uniqueness and storage costs.
Patent Information
- Application Number
- CN202310923847.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-07-25
AI Technical Summary
When the data processed by Flink is written to the Elasticsearch cluster, since each row of data contains an index alias, row data of the same primary key field is repeatedly written under the same index alias in the Elasticsearch cluster, resulting in data redundancy and increased storage costs.
Through the Flink-Elasticsearch connector, you can first find the specific physical index in the Elasticsearch cluster based on the primary key field and index alias of the row data to be written, replace the index alias to this physical index and then perform the write operation to ensure that the hot and cold index scrolling is sensed and the data is avoided repeatedly written.
Ensure that only one row of data with the same primary key field in the Elasticsearch cluster exists, which reduces the storage cost of indexed data, improves data writing efficiency, and increases the complexity of downstream third-party query data.
Smart Images

Figure CN117131101B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and in particular to a data writing method, apparatus, electronic device and storage medium. Background Art
[0002] In the field of big data, it is usually necessary to write the data calculated in real time by the open-source stream real-time processing framework Flink into the Elasticsearch cluster for other third-party systems to query the data online through the Elasticsearch cluster. Moreover, in order to ensure the orderliness and reliability of data storage in the Elasticsearch cluster, not only are physical indexes generated by rolling over time under each same index alias in the Elasticsearch cluster, and the index documents under each physical index are used to store the row data processed by Flink, but also the Elasticsearch cluster supports the management of the data life cycle. The latest rolled physical index under the same index alias is the hot index, and each row data will be written into the hot index under the corresponding index alias. For example, when two physical indexes, index01 and index02, are generated by rolling under the index alias (alias) in the Elasticsearch cluster, the row data will be written into the hot index index02; when three physical indexes, index01, index02, and index03, are generated by rolling under the same index alias in the Elasticsearch cluster, the row data will be written into the hot index index03.
[0003] However, since each row data processed by Flink includes the index alias written into the Elasticsearch cluster, each row data will be directly written into the index document corresponding to the hot index under the corresponding alias, resulting in duplicate row data with the same primary key field being queried by downstream third parties. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems existing in the related art. To this end, the present invention proposes a data writing method. By means of the Flink-Elasticsearch connector, first, according to the primary key field and index alias carried in the row data to be written, the specific physical index in the Elasticsearch cluster is found, and then the index alias carried in the row data to be written is replaced with this specific physical index before performing the writing operation, so as to ensure that the Flink-Elasticsearch connector can sense the hot and cold index rolling of each physical index in the Elasticsearch cluster, thereby ensuring that there is only one row of data with the same primary key field in the Elasticsearch cluster, avoiding the drawback of repeated writing of row data with the same primary key field under the same index alias in the Elasticsearch cluster, effectively reducing the storage cost of index data in the Elasticsearch cluster, improving the data writing efficiency, and at the same time increasing the complexity of downstream third-party querying data.
[0005] The present invention also proposes a data writing device.
[0006] The present invention also proposes an electronic device.
[0007] The present invention also proposes a non-transitory computer-readable storage medium.
[0008] According to the data writing method of the first aspect embodiment of the present invention, which is applied to the Flink-Elasticsearch connector, the method includes:
[0009] In response to a data writing instruction, determine the target primary key field and target alias included in the row data to be written;
[0010] Based on the matching result of the target primary key field and the target alias with the Elasticsearch cluster, determine the target physical index from the physical indexes under the target alias in the Elasticsearch cluster;
[0011] Replace the target alias with the target physical index, determine the target row data to be written, and send the target row data to be written to the Elasticsearch cluster;
[0012] Wherein, the Elasticsearch cluster is used to write the target row data to be written into the index document corresponding to the target physical index.
[0013] According to the data writing method of the embodiments of the present invention, this method first finds the specific physical index in the Elasticsearch cluster according to the primary key field and index alias carried by the row data to be written through the Flink-Elasticsearch connector, and then replaces the index alias carried by the row data to be written with this specific physical index and then performs the writing operation, ensuring that the Flink-Elasticsearch connector can perceive the hot and cold index rolling of each physical index in the Elasticsearch cluster, so as to ensure that there is only one row of data with the same primary key field in the Elasticsearch cluster, avoiding the drawback of repeated writing of row data with the same primary key field under the same index alias in the Elasticsearch cluster, effectively reducing the storage cost of index data in the Elasticsearch cluster, improving the data writing efficiency, and at the same time increasing the complexity of downstream third-party querying data.
[0014] According to an embodiment of the present invention, determining the target physical index from the physical indexes under the target alias in the Elasticsearch cluster based on the matching result of the target primary key field and the target alias with the Elasticsearch cluster includes:
[0015] Sending the target primary key field and the target alias to the Elasticsearch cluster, where the Elasticsearch cluster is used to match the target primary key field with the index documents corresponding to the physical indexes under the target alias;
[0016] Receiving the matching result fed back by the Elasticsearch cluster, and determining the target physical index from the physical indexes under the target alias in the Elasticsearch cluster based on the matching result.
[0017] According to an embodiment of the present invention, determining the target physical index from the physical indexes under the target alias in the Elasticsearch cluster based on the matching result includes:
[0018] Based on the first matching result that the target primary key field already exists in the Elasticsearch cluster under the target alias, determining the physical index corresponding to the index document with the target primary key field existing under the target alias in the Elasticsearch cluster as the target physical index;
[0019] Wherein, the first matching result is specifically used to indicate that the target primary key field exists in the index documents corresponding to the physical indexes under the target alias in the Elasticsearch cluster.
[0020] According to an embodiment of the present invention, the method further includes:
[0021] Based on the first matching result, instruct the Elasticsearch cluster to update the row data containing the target primary key field in the index document corresponding to the target physical index under the target alias to the target row data to be written.
[0022] According to an embodiment of the present invention, the determining the target physical index from the physical indexes under the target alias in the Elasticsearch cluster based on the matching result further includes:
[0023] Based on the second matching result that the target primary key field does not exist in the index documents corresponding to all the physical indexes under the target alias in the Elasticsearch cluster, determine the hot index under the target alias in the Elasticsearch cluster;
[0024] Determine the physical index corresponding to the hot index as the target physical index;
[0025] Wherein, the second matching result is specifically used to characterize that the target primary key field does not exist in the index documents corresponding to all the physical indexes under the target alias in the Elasticsearch cluster, and the life cycle indexes corresponding to all the physical indexes, and each life cycle index is one of a hot index, a warm index, and a cold index.
[0026] According to an embodiment of the present invention, the method further includes:
[0027] Based on the second matching result, instruct the Elasticsearch cluster to add the target row data to be written to the index document corresponding to the target physical index under the target alias.
[0028] According to an embodiment of the present invention, the data write instruction is a batch data write instruction. Before determining the target primary key field and the target alias included in the row data to be written in response to the data write instruction, the method further includes:
[0029] Receive the row data to be written transmitted after real-time calculation by the Fllink framework;
[0030] Determine that the cumulative byte count of the received multiple rows of data to be written reaches the byte count threshold, and automatically generate a batch data write instruction.
[0031] According to an embodiment of the present invention, after receiving the to-be-written row data transmitted after real-time calculation by the Fllink framework, the method further includes:
[0032] Determine that the cumulative duration of the received multiple to-be-written row data reaches a preset duration, and automatically generate the batch data writing instruction.
[0033] According to the data writing device of the second aspect embodiment of the present invention, which is applied to the Flink-Elasticsearch connector, the device includes:
[0034] An information determination module, configured to determine a target primary key field and a target alias included in the to-be-written row data in response to a data writing instruction;
[0035] An index determination module, configured to determine a target physical index in the Elasticsearch cluster that matches the target primary key field and the target alias;
[0036] A data writing module, configured to replace the target alias with the target physical index, determine target to-be-written row data including the target primary key field and the target physical index, and send the target to-be-written row data to the Elasticsearch cluster;
[0037] Wherein, the Elasticsearch cluster is configured to write the target to-be-written row data into an index document corresponding to the target physical index.
[0038] According to the data writing device of the embodiment of the present invention, the device first finds a specific physical index in the Elasticsearch cluster according to the primary key field and index alias carried by the to-be-written row data through the Flink-Elasticsearch connector, and then replaces the index alias carried by the to-be-written row data with the specific physical index and then performs the writing operation, ensuring that the Flink-Elasticsearch connector can perceive the hot and cold index rolling of each physical index in the Elasticsearch cluster, so as to ensure that there is only one row of data with the same primary key field in the Elasticsearch cluster, avoiding the drawback of repeated writing of row data with the same primary key field under the same index alias in the Elasticsearch cluster, effectively reducing the storage cost of index data in the Elasticsearch cluster, improving the data writing efficiency, and at the same time improving the complexity of downstream third-party query data.
[0039] One or more of the above technical solutions in the embodiments of the present invention have at least one of the following technical effects: By means of the Flink-Elasticsearch connector, first find the specific physical index in the Elasticsearch cluster according to the primary key field and index alias carried by the row data to be written, and then replace the index alias carried by the row data to be written with the specific physical index before performing the write operation, ensuring that the Flink-Elasticsearch connector can perceive the hot and cold index rolling of each physical index in the Elasticsearch cluster, so as to ensure that there is only one row of data with the same primary key field in the Elasticsearch cluster, avoiding the drawback of repeated writing of row data with the same primary key field under the same index alias in the Elasticsearch cluster, effectively reducing the storage cost of index data in the Elasticsearch cluster, improving the data writing efficiency, and at the same time increasing the complexity of downstream third-party query data.
[0040] Further, the Flink-Elasticsearch connector determines the target physical index corresponding to the target alias in the Elasticsearch cluster by sending the target alias and target primary key field contained in the row data to be written to the Elasticsearch cluster for matching and receiving the matching result feedback from the Elasticsearch cluster. In this way, not only the flexible interactivity between the Flink-Elasticsearch connector and the Elasticsearch cluster is improved, but also the reliability and accuracy of determining the target physical index can be improved.
[0041] Furthermore, when the Flink-Elasticsearch connector determines that there is a target primary key field under the target alias in the Elasticsearch cluster, the physical index corresponding to the index document with the target primary key field under the target alias in the Elasticsearch cluster is determined as the target physical index. In this way, the convenience and accuracy of determining the target physical index can be improved.
[0042] Still further, when the Flink-Elasticsearch connector determines that there is no target primary key field under the target alias in the Elasticsearch cluster, the physical index corresponding to the hot index under the target alias in the Elasticsearch cluster is determined as the target physical index. In this way, by means of the Flink-Elasticsearch connector's ability to dynamically perceive the hot and cold index rolling mode of each physical index in the Elasticsearch cluster, the flexibility and accuracy of determining the target physical index can be further improved.
[0043] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the related art, the following will briefly introduce the drawings required for use in the description of the embodiments or the related art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0045] Figure 1 is a schematic flowchart of the data writing method provided by an embodiment of the present invention;
[0046] Figure 2 is a data processing logic diagram of the data writing method provided by an embodiment of the present invention;
[0047] Figure 3 is a schematic structural diagram of the data writing device provided by an embodiment of the present invention;
[0048] Figure 4 is a schematic structural diagram of the electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention fall within the scope of protection of the present invention.
[0050] In the embodiments of the present invention, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B may be singular or plural. In the written description of the present invention, the character " / " generally represents an "or" relationship between the associated objects before and after. In addition, it should be noted that the serial numbers assigned to the objects described in the present invention itself, such as "first", "second", etc., are only used to distinguish the described objects and do not have any sequential or technical meaning.
[0051] In the field of big data, it is usually necessary to write the data calculated in real time by the open-source stream real-time processing framework Flink into an Elasticsearch cluster for other third-party systems to query the data online through the Elasticsearch cluster. Moreover, to ensure the orderliness and reliability of data storage in the Elasticsearch cluster, not only are physical indexes generated by rolling over time under each same index alias in the Elasticsearch cluster, and the index documents under each physical index are used to store the row data processed by Flink, but also the Elasticsearch cluster supports managing the data life cycle. The latest physically generated index under the same index alias is the hot index, and each row data will be written into the hot index under the corresponding index alias. For example, when two physical indexes, index01 and index02, are generated by rolling under the index alias (alias) in the Elasticsearch cluster, the row data will be written into the hot index index02; when three physical indexes, index01, index02, and index03, are generated by rolling under the same index alias in the Elasticsearch cluster, the row data will be written into the hot index index03.
[0052] In the related art, the process of the Flink framework writing data into the Elasticsearch cluster includes (1) to (3).
[0053] (1) The Flink framework provides the RichSinkFunction interface. Implementing the RichSinkFunction interface can output the row data processed in real time by the Flink framework to other external systems such as the Elasticsearch cluster. The ElasticsearchSink class in the Flink-connector (connector)-Elasticsearch module, which is responsible for writing the row data processed in real time by the Flink framework into the Elasticsearch cluster, implements this RichSinkFunction interface.
[0054] (2) When the Flink framework starts a real-time computing task, it will initially instantiate the ElasticsearchSink class. The startup process includes: checking and saving the connection configuration information of the Elasticsearch cluster, determining whether to enable the BulkProcessor configuration. After the BulkProcessor configuration is enabled, it can accelerate the performance of the client writing to the Elasticsearch cluster; then, using the connection configuration information of the Elasticsearch cluster, instantiate the client that connects the Flink framework to the Elasticsearch cluster. Before configuring the client to write to the Elasticsearch cluster, it is necessary to call the listener BulkProcessorListener object to process the feedback results of the Elasticsearch cluster when the Flink framework writes row data each time, such as failure retry or throwing an exception, etc.
[0055] (3) When the Flink framework's real-time computing task is running, it calls the invoke function of the ElasticsearchSink class to process the row data that needs to be written to the Elasticsearch cluster. The processing process includes: First, check whether there is an exception in the client that the Flink framework is currently connected to the Elasticsearch cluster. When there is no exception, assemble the row data to be written into a RequestIndexer object; then, add the row data to be written to the Elasticsearch cluster to the RequestIndexer object, and determine whether to write the row data to the Elasticsearch cluster by flushing the cache; finally, the Elasticsearch cluster receives the BulkRequest request, parses the BulkRequest request to obtain all DocWriteRequest objects, and sends the DocWriteRequest objects to the target physical index for data update.
[0056] As can be seen from the above steps (1) to (3), when writing the row data processed by the Flink framework into the Elasticsearch cluster, it can only be written into the Elasticsearch cluster according to the physical index or index alias specified for each row of data, and it is impossible to adapt to the scenario of multiple physical indexes under one index alias when the Elasticsearch cluster manages the index life cycle for data; that is, since each row of data processed by Flink includes the index alias written into the Elasticsearch cluster, each row of data will be directly written into the hot index under the corresponding index alias. Then, when the row data containing the same primary key field is repeatedly written into the same index alias in the Elasticsearch cluster, there will inevitably be a defect that the data is repeatedly written into the same physical index or different physical indexes under the index alias, resulting in duplicate row data containing the same primary key field being queried by downstream third parties.
[0057] To solve the above technical problems, the present invention provides a data writing method, device, electronic device and storage medium. The following will be combined with Figures 1 to 4 Describe the data writing method, device, electronic device and storage medium of the present invention. The execution subject of the data writing method is the Flink-Elasticsearch connector. This Flink-Elasticsearch connector can be set between the Flink framework and the Elasticsearch cluster, or can be set in the same server together with the Flink framework in the form of a module. The Flink framework is an open-source stream processing framework developed by the Apache Software Foundation, and its core is a distributed stream data flow engine written in Java and Scala; Flink executes any stream data program in a data parallel and pipeline manner, and the pipeline runtime system of the Flink framework can execute batch processing and stream processing programs; in addition, the runtime of the Flink framework itself also supports the execution of iterative algorithms. The Elasticsearch cluster is a distributed, highly scalable, and highly real-time search and data analysis engine, which can conveniently and quickly enable a large amount of data to have the ability to search, analyze and explore, and can also make the data more valuable in the production environment by making full use of horizontal scalability. The Flink-Elasticsearch connector at least has functions such as data life cycle processing monitoring, data receiving, interface calling, function adding, data processing and data sending. And the following method embodiments will be described by taking the execution subject as the Flink-Elasticsearch connector as an example.
[0058] To facilitate the understanding of the data writing method provided by the embodiments of the present invention, the data writing method provided by the embodiments of the present invention will be described in detail through the following several exemplary embodiments. It can be understood that these several exemplary embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0059] Referring to Figure 1 , which is a schematic flowchart of the data writing method provided by the embodiments of the present invention. As Figure 1 shown, the data writing method includes the following steps 110 to step 130.
[0060] Step 110: In response to a data writing instruction, determine the target primary key field and the target alias included in the row data to be written.
[0061] Among them, the row data to be written can be row data generated after the Flink framework processes other data such as business data or log data. The row data to be written can include but is not limited to the target alias, operation type, target primary key field, and specific content. For example, the target alias in the row data to be written is index_alias, the operation type is insert, the target primary key field (_id) is 1, and the specific content is abc; the target alias in the row data to be written is index_alias, the operation type is delete, the target primary key field (_id) is 2, and the specific content is efg; the index alias in the row data to be written is index_alias, the operation type is update, the target primary key field (_id) is 3, and the specific content is hij. The target alias is the index alias corresponding to the row data to be written that can be written into the Elasticsearch cluster. For example, when the target alias is index_alias, the corresponding row data to be written is written under the index alias index alias in the Elasticsearch cluster. The target primary key field can be the primary key field corresponding to the row data to be written. For example, when the primary key field in the row data to be written is 1, _id = 1; when the primary key field in the row data to be written is 2, _id = 2; when the primary key field in the row data to be written is 3, _id = 3. In addition, the number of rows of data to be written can be 1 or multiple; no specific limitation is made here.
[0062] Specifically, the Flink-Elasticsearch connector can automatically generate a data writing instruction when receiving the row data to be written sent by the Flink framework, or directly receive the data writing instruction sent by the Flink framework; the data writing instruction can carry the target alias, operation type, target primary key field, and specific content that make up the row data to be written. Further, in response to this data writing instruction, the Flink-Elasticsearch connector can specifically parse this data writing instruction to determine the target primary key field and target alias in the row data to be written from the data writing instruction.
[0063] It should be noted that the Flink-Elasticsearch connector can specifically be the Flink-Elasticsearch7 connector. The Flink-Elasticsearch7 connector is a connector that connects the Flink framework and the Elasticsearch7 cluster, and is a connector formed after the Flink framework encapsulates the code logic for writing data to the Elasticsearch7 cluster into a module. The Elasticsearch7 cluster is a high-availability solution for the Elasticsearch database. That is, a system composed of multiple Elasticsearch instances can provide high-availability and high-performance index services. The Elasticsearch7 cluster can perform index lifecycle management on different physical indexes (index) under the same index alias. Index lifecycle management was first introduced in Elasticsearch 6.6 (public beta) and officially launched in Elasticsearch 6.7. It is mainly used to manage the index documents in each physical index under each index alias, and uses a hot-cold separation architecture to split the index documents contained in each of the multiple physical indexes into hot, warm, and cold indexes, so that the index documents continuously roll dynamically from the hot stage -> warm stage -> cold stage -> delete stage. For example, when two physical indexes, index01 and index02, are generated by rolling under the index alias (index_alias) in the Elasticsearch cluster, index02 is the hot index and index01 is the warm index; when three physical indexes, index01, index02, and index03, are generated by rolling under the index alias (index_alias) in the Elasticsearch cluster, index03 is the hot index, index03 is the warm index, and index01 is the cold index.
[0064] Step 120: Based on the matching result of the target primary key field and the target alias with the Elasticsearch cluster, determine the target physical index from the physical indexes under the target alias in the Elasticsearch cluster.
[0065] Specifically, the Flink-Elasticsearch connector can pre-establish an index lifecycle rolling reporting protocol with the Elasticsearch cluster. This index lifecycle rolling reporting protocol is used to indicate that when the Elasticsearch cluster detects a new physical index under a certain index alias and the lifecycle fields of each physical index change, it will send the corresponding index alias and all the primary key fields of the index documents in each physical index under this index alias to the Flink-Elasticsearch connector. For example, when the Elasticsearch cluster detects that the physical indexes under the index alias (index_alias) roll and new index03 is added from index01 and index02, index02 changes from a hot index to a warm index, index01 changes from a warm index to a cold index, and the newly added index03 is a hot index. At this time, the Elasticsearch cluster can report this index alias (index_alias) and all the primary key fields in the index documents of index01, index02, and index03 under this index_alias to the Flink-Elasticsearch connector. Based on this, the Flink-Elasticsearch connector can match the index alias reported by the Elasticsearch cluster and all the primary key fields of the index documents in each physical index under this index alias with the target primary key field and the target alias respectively, so as to determine the target physical index from the physical indexes under the target alias in the Elasticsearch cluster, that is, the target physical index in the Elasticsearch cluster that matches both the target primary key field and the target alias.
[0066] Step 130: Replace the target alias with the target physical index, determine the target row data to be written, and send the target row data to be written to the Elasticsearch cluster.
[0067] Among them, the Elasticsearch cluster is used to write the target row data to be written into the index document corresponding to the target physical index. In addition to the target primary key field and the target physical index, the target row data to be written may also include the operation type and the specific content. For example, for the target alias in the row data to be written being index_alias, the operation type being insert, _id being 1, and the specific content being abc, if index_alias is replaced with the target physical index index_01, the corresponding determined target row data to be written includes the target physical index index_01, the operation type insert, the target primary key field with _id being 1, and the specific content abc; for the target alias in the row data to be written being index_alias, the operation type being delete, _id being 2, and the specific content being efg, if index_alias is replaced with the target physical index index_02, the corresponding determined target row data to be written includes the target physical index index_02, the operation type delete, the target primary key field with _id being 2, and the specific content efg; for the index alias in the row data to be written being index_alias, the operation type being update, _id being 3, and the specific content being hij, if index_alias is replaced with the target physical index index_03, the corresponding determined target row data to be written includes the target physical index index_03, the operation type update, the target primary key field with _id being 3, and the specific content hij.
[0068] Specifically, when the Flink-Elasticsearch connector determines the target alias, it can replace the target alias in the row data to be written with the target physical index, thereby updating the row data to be written including the target primary key field and the target alias to the target row data to be written including the target primary key field and the target physical index; then send this target row data to be written to the Elasticsearch cluster to instruct the Elasticsearch cluster to write this target row data to be written into the index document corresponding to the target physical index.
[0069] It should be noted that the ElasticsearchSink class for performing data writing operations is pre-set in the Flink-Elasticsearch connector. An internal class of a lifecycle processing listener is newly added to the ElasticsearchSink class, and the internal class of the lifecycle processing listener inherits the BulkProcessor listener interface; the implementation of the BulkProcessor listener interface is as follows: before the Flink-Elasticsearch connector writes the row data to be written into the Elasticsearch cluster, it determines the target physical index in the Elasticsearch cluster that matches the target primary key field and the target alias of the row data to be written, and then after replacing the target alias with the target physical index, it determines the target row data to be written that includes the target primary key field and the target physical index. The implementation of the BulkProcessor listener interface represents that the index lifecycle processing of the row data to be written is monitored.
[0070] Exemplarily, although the ElasticsearchSink class inherits from the ElasticsearchSinkBase class, all the variables defined in the ElasticsearchSinkBase class are private variables (private). In the Java specification, private cannot be accessed by subclasses, so two core functions are added to ElasticsearchSink to bypass the Java specification restrictions. The two core functions added here are the getSuperFieldValue function and the setSuperFieldValue function. The getSuperFieldValue function is used to obtain the private defined in the ElasticsearchSinkBase class through the getSuperclass function of the Java object class; the setSuperFieldValue function is used to reassign the private defined in the ElasticsearchSinkBase class through the setAccessible function of the Java object class. Based on this, the specific implementation of the newly added inner class of the lifecycle processing listener in the ElasticsearchSink class is as follows: by calling the getSuperFieldValue function and the setSuperFieldValue function, bypass the private defined in the ElasticsearchSinkBase class in the Java specification, and assign values to the variables required inside the lifecycle processing listener, so that the lifecycle processing listener after assignment can perform lifecycle processing monitoring on the rows of data to be written sent by the Flink framework. The lifecycle processing monitoring process here is the implementation process of the BulkProcessor listener interface. In this way, by modifying the source code of the Flink-Elasticsearch connector, when writing the initial rows of data to be written to the Elasticsearch cluster, it first queries the target physical index where the row data is stored in the Elasticsearch cluster and then performs index alias replacement, and then sends the generated target rows of data to be written after replacement to the Elasticsearch cluster. Thus, it is ensured that the index data with the same primary key field can be correctly updated or stored.
[0071] The data writing method provided by the embodiment of the present invention uses the Flink-Elasticsearch connector to first find the specific physical index in the Elasticsearch cluster according to the primary key field and index alias carried by the row data to be written, and then replaces the index alias carried by the row data to be written with the specific physical index and then performs the writing operation, ensuring that the Flink-Elasticsearch connector can perceive the hot and cold index rolling of each physical index in the Elasticsearch cluster, so as to ensure that there is only one row of data with the same primary key field in the Elasticsearch cluster, avoiding the drawback of repeated writing of row data with the same primary key field under the same index alias in the Elasticsearch cluster, effectively reducing the storage cost of index data in the Elasticsearch cluster, improving the data writing efficiency, and at the same time increasing the complexity of downstream third-party query data.
[0072] It can be understood that when the Flink-Elasticsearch connector determines the target primary key field and target alias contained in the row data to be written, it can determine the target physical index by instructing the Elasticsearch cluster to query whether there is a physical index containing both the target primary key field and the target alias at the same time. Based on this, the specific implementation process of step 120 may include:
[0073] First, send the target primary key field and target alias to the Elasticsearch cluster, and the Elasticsearch cluster is used to match the index document corresponding to the physical index under the target alias with the target primary key field; then further receive the matching result feedback by the Elasticsearch cluster, and based on the matching result, determine the target physical index from the physical indexes under the target alias in the Elasticsearch cluster.
[0074] Among them, each index document includes row data containing different primary key fields.
[0075] Specifically, the Flink-Elasticsearch connector can send the target primary key field and the target alias in the row data to be written to the Elasticsearch cluster, so as to instruct the Elasticsearch cluster to match the index document corresponding to the physical index under the target alias with the target primary key field to obtain a matching result, and then feedback this matching result to the Flink-Elasticsearch connector. At this time, the Flink-Elasticsearch connector can determine the target physical index from the physical index under the target alias in the Elasticsearch cluster based on the matching result. For example, when the matching result is that the target alias does not exist in the Elasticsearch cluster, there is naturally no target physical index corresponding to the index document containing the target primary key field under this target alias, that is, there is no target physical index in the Elasticsearch cluster that matches both the target alias and the target primary key field. At this time, the Flink-Elasticsearch connector can instruct the Elasticsearch cluster to create a new target alias and roll-generate the first physical index under the target alias. At this time, the first physical index under the new target alias created in the Elasticsearch cluster can be determined as the target physical index.
[0076] In the data writing method provided by the embodiments of the present invention, the Flink-Elasticsearch connector determines the target physical index corresponding to the target alias in the Elasticsearch cluster by sending the target alias and the target primary key field contained in the row data to be written to the Elasticsearch cluster for matching, and receiving the matching result feedback from the Elasticsearch cluster. In this way, not only the flexible interactivity between the Flink-Elasticsearch connector and the Elasticsearch cluster is improved, but also the reliability and accuracy of determining the target physical index can be improved.
[0077] It can be understood that in the case where the target alias exists in the Elasticsearch cluster and the target primary key field also exists in the index document corresponding to the physical index under the target alias, based on the matching result, determining the target physical index from the physical index under the target alias in the Elasticsearch cluster, the specific implementation process may include:
[0078] Based on the first matching result that the target primary key field already exists in the Elasticsearch cluster under the target alias, the physical index corresponding to the index document with the target primary key field existing under the target alias in the Elasticsearch cluster is determined as the target physical index.
[0079] Among them, the first matching result is specifically used to characterize that the target primary key field exists in the index document corresponding to the physical index under the target alias in the Elasticsearch cluster.
[0080] Specifically, when the Elasticsearch cluster traverses the index documents corresponding to each physical index under the target alias based on the target primary key field and determines that the target primary key field already exists in the first matching result under the target alias in the Elasticsearch cluster, the first matching result can be sent to the Flink-Elasticsearch connector. At this time, the Flink-Elasticsearch connector can determine the target physical index corresponding to the target primary key field under the target alias in the Elasticsearch cluster from the first matching result.
[0081] Exemplarily, for the target alias being index_alias and the target primary key field _id = 1, when the first matching result is that _id = 1 exists in index_alias of the Elasticsearch cluster, the physical index index_01 of the target primary key field where _id = 1 exists in index_alias can be determined as the target physical index.
[0082] In the data writing method provided by the embodiments of the present invention, when the Flink-Elasticsearch connector determines that there is a target primary key field under the target alias in the Elasticsearch cluster, the physical index corresponding to the index document with the target primary key field under the target alias in the Elasticsearch cluster is determined as the target physical index. In this way, the convenience and accuracy of determining the target physical index can be improved.
[0083] It can be understood that when it is determined that there is a target physical index containing the target primary key field under the target alias in the Elasticsearch cluster, the Elasticsearch cluster can update the existing row data in the index document corresponding to the target physical index. Based on this, the data writing method provided by the embodiments of the present invention may further include:
[0084] Based on the first matching result, instruct the Elasticsearch cluster to update the row data containing the target primary key field in the index document corresponding to the target physical index under the target alias to the target row data to be written.
[0085] Specifically, when it is determined that there is a target physical index containing the target primary key field under the target alias in the Elasticsearch cluster, the Flink-Elasticsearch connector can instruct the Elasticsearch cluster to update the row data containing the target primary key field in the index document corresponding to the target physical index under the target alias to the target row data to be written, that is, instruct the Elasticsearch cluster to update the row data containing the target primary key field in the index document corresponding to the target physical index under the target alias based on the target row data to be written. For example, for the target alias being index_alias and the target primary key field _id = 1, when the first matching result is that _id = 1 already exists in index_01 under index_alias in the Elasticsearch cluster, the row data containing _id = 1 in the index document corresponding to index_01 under index_alias in the Elasticsearch cluster can be updated with the target row data to be written. This achieves the purpose of overwriting at least one of the fields such as name, phone number, address, and age in the existing row data in the Elasticsearch cluster.
[0086] In the data writing method provided by the embodiments of the present invention, when it is determined that there is a target physical index containing the target primary key field under the target alias in the Elasticsearch cluster, the Flink-Elasticsearch connector ensures that the index data with the same primary key field can be correctly updated by instructing the Elasticsearch cluster to update the corresponding row data based on the target row data to be written, thereby improving the accuracy and timeliness of updating the index data in the Elasticsearch cluster.
[0087] It can be understood that in the case where there is a target alias in the Elasticsearch cluster and the target primary key field does not exist in all index documents under the target alias, based on the matching result, to determine the target physical index from the physical indexes under the target alias in the Elasticsearch cluster, the specific implementation process may further include:
[0088] First, based on the second matching result that the target primary key field does not exist in the Elasticsearch cluster under the target alias, determine the hot index under the target alias in the Elasticsearch cluster; further, determine the physical index corresponding to the hot index as the target physical index.
[0089] Among them, the second matching result is specifically used to characterize that the target primary key field does not exist in the index documents corresponding to all physical indexes under the target alias in the Elasticsearch cluster, and the life cycle indexes corresponding to all these physical indexes, where each life cycle index is one of a hot index, a warm index, and a cold index respectively.
[0090] Specifically, when the Elasticsearch cluster traverses the index documents corresponding to each physical index under the target alias based on the target primary key field and determines that the target primary key field does not exist in the second matching result of the target alias in the Elasticsearch cluster, the hot index under the target alias can be determined from the second matching result. The hot index can be the physically latest rolled-up index under the target alias; and the physical index corresponding to this hot index is determined as the target physical index corresponding to the target primary key field under the target alias in the Elasticsearch cluster.
[0091] Exemplarily, for the target alias being index_alias and the target primary key field _id = 2, when the second matching result is that _id = 2 does not exist under index_alias in the Elasticsearch cluster, the physical index index_03 corresponding to the hot index under index_alias in the Elasticsearch cluster can be determined as the target physical index.
[0092] In the data writing method provided by the embodiments of the present invention, when the Flink-Elasticsearch connector determines that the target primary key field does not exist under the target alias in the Elasticsearch cluster, the physical index corresponding to the hot index under the target alias in the Elasticsearch cluster is determined as the target physical index. In this way, through the Flink-Elasticsearch connector, the hot and cold index rolling methods of each physical index in the Elasticsearch cluster can be dynamically sensed, which can further improve the flexibility and accuracy of determining the target physical index.
[0093] It can be understood that when it is determined that there is no target physical index containing the target primary key field under the target alias in the Elasticsearch cluster, the Elasticsearch cluster can perform data addition. Based on this, the data writing method provided by the embodiments of the present invention may further include:
[0094] Based on the second matching result, instruct the Elasticsearch cluster to add the target row data to be written to the index document corresponding to the target physical index under the target alias.
[0095] Specifically, when it is determined that there is no target physical index containing the target primary key field under the target alias in the Elasticsearch cluster, the Flink-Elasticsearch connector can instruct the Elasticsearch cluster to add the target row data to be written to the index document corresponding to the target physical index under the target alias. For example, for the target alias index_alias and the target primary key field _id = 3, when the second matching result is that _id = 3 does not exist in index_02 under index_alias in the Elasticsearch cluster, the target row data to be written can be added to the index document corresponding to index_02 under index_alias in the Elasticsearch cluster.
[0096] In the data writing method provided by the embodiments of the present invention, when it is determined that there is no target physical index of the target primary key field under the target alias in the Elasticsearch cluster, the Flink-Elasticsearch connector ensures the accuracy and uniqueness of data writing by instructing the Elasticsearch cluster to add new data based on the target row data to be written, and improves the necessity and reliability of adding index data in the Elasticsearch cluster.
[0097] It can be understood that, in order to improve the data writing efficiency, data can be written to the Elasticsearch cluster in batches. At this time, a batch data writing instruction can be generated by receiving that the cumulative byte size of the row data to be written meets the preset requirements. Based on this, before step 110, the data writing method provided by the embodiments of the present invention can further include:
[0098] First, receive the row data to be written transmitted after real-time calculation by the Flink framework; then, determine that the cumulative number of bytes of the received multiple rows of data to be written reaches the byte number threshold, and automatically generate a batch data writing instruction.
[0099] Among them, the batch data writing instruction carries each row data to be written that reaches the byte number threshold, and each row data to be written contains the target primary key field and the target alias.
[0100] Specifically, in the process of receiving the row data to be written sent by Flink by the Flink-Elasticsearch connector, when each row data to be written is received, the number of bytes of the corresponding row data to be written can be counted, and the counted number of bytes of each row data to be written is accumulated until the accumulated number of bytes reaches the byte number threshold, and a batch data writing instruction is automatically generated. The batch data writing instruction here is used to represent that all the row data to be written that reaches the byte number threshold is written to the Elasticsearch cluster in batches.
[0101] In the data writing method provided by the embodiment of the present invention, when the Flink-Elasticsearch connector determines that the cumulative byte count of multiple rows of data to be written received reaches the byte count threshold, it automatically generates a batch data writing instruction. In this way, through the batch writing data writing method, the data writing efficiency can be effectively improved, and the data writing complexity is reduced.
[0102] It can be understood that when generating the batch writing instruction, in addition to determining whether the cumulative byte count of multiple initial data to be written reaches the preset requirement, the batch data writing instruction can also be generated by determining whether the cumulative duration of receiving multiple rows of data to be written meets the preset requirement. Based on this, after receiving the rows of data to be written transmitted after real-time calculation by the Fllink framework, the data writing method provided by the embodiment of the present invention may further include:
[0103] Determine that the cumulative duration of the multiple rows of data to be written received reaches the preset duration, and automatically generate a batch data writing instruction.
[0104] Among them, in the batch data writing instruction carrying the multiple rows of data to be written that reach the preset duration, each row of data to be written also contains a target primary key field and a target alias respectively.
[0105] Specifically, during the process of the Flink-Elasticsearch connector receiving the rows of data to be written sent by Flink, it can determine whether the cumulative duration of receiving multiple rows of data to be written reaches the preset duration. When it is determined that the cumulative duration of receiving multiple rows of data to be written reaches the preset duration, a batch data writing instruction is automatically generated. For example, a batch data writing instruction can be automatically generated when the cumulative duration of the multiple rows of data to be written received reaches 1 second, 1.5 seconds, or 2 seconds. The batch data writing instruction here is used to represent batch writing all the rows of data to be written received within the preset duration into the Elasticsearch cluster.
[0106] In the data writing method provided by the embodiment of the present invention, when the Flink-Elasticsearch connector determines that the cumulative duration of the multiple rows of data to be written received reaches the preset duration, it automatically generates a batch data writing instruction. This improves the flexibility and convenience of generating the batch data writing instruction, ensuring that subsequent data writing is more efficient and concise.
[0107] Refer to Figure 2 , for the data processing logic diagram of the data writing method provided by the embodiment of the present invention. In Figure 2In the process, when the Flink framework receives a submitted Flink task, it will perform real-time calculations on the data to be processed carried by the Flink task and send the calculation results to the Flink-Elasticsearch connector. The calculation results here are the original index data shown in Figure 2 . The original index data specifically refers to the 3 rows of data to be written corresponding to the generated batch data instruction. By sending the target primary key field (_id) and the target alias contained in each row of data to be written to the Elasticsearch cluster for query, the target physical index in the Elasticsearch cluster that matches each target primary key field and target alias respectively is determined. For example, the target physical index in the Elasticsearch cluster that matches the target alias index_alias and 1 is index_01, the target physical index in the Elasticsearch cluster that matches the target alias index_alias and 2 is index_02, and the target physical index in the Elasticsearch cluster that matches the target alias index_alias and 3 is index_03. After replacing the target aliases in the 3 rows of data to be written with the corresponding target physical indexes in this way, 3 target rows of data to be written are obtained. The 3 target rows of data to be written here are also the replaced index data in Figure 2 . Then, the 3 target rows of data to be written are sent to the Elasticsearch cluster so that the Elasticsearch cluster writes each target row of data to be written into the index document corresponding to the target physical index. The specific implementation process and technical effects involved can be referred to the foregoing embodiments and will not be elaborated here.
[0108] Refer to Figure 3 , which is a schematic structural diagram of the data writing device provided by the embodiment of the present invention. As shown in Figure 3 , the data writing device 300 includes an information determination module 310, an index determination module 320, and a data writing module 330.
[0109] The information determination module 310 is configured to determine the target primary key field and the target alias contained in the row of data to be written in response to the data writing instruction.
[0110] The index determination module 320 is configured to determine the target physical index from the physical indexes under the target alias in the Elasticsearch cluster based on the matching results of the target primary key field and the target alias with the Elasticsearch cluster.
[0111] The data writing module 330 is used to replace the target alias with the target physical index, determine the target data to be written in a row, and send the target data to be written in a row to the Elasticsearch cluster.
[0112] Among them, the Elasticsearch cluster is used to write the target data to be written in a row into the index document corresponding to the target physical index.
[0113] It can be understood that the index determination module 320 specifically sends the target primary key field and the target alias to the Elasticsearch cluster. The Elasticsearch cluster is used to match the index document corresponding to the physical index under the target alias with the target primary key field; receive the matching result fed back by the Elasticsearch cluster, and determine the target physical index from the physical indexes under the target alias in the Elasticsearch cluster based on the matching result.
[0114] It can be understood that the index determination module 320 is specifically further used to determine the physical index corresponding to the index document in which the target primary key field exists under the target alias in the Elasticsearch cluster based on the first matching result that the target primary key field already exists under the target alias in the Elasticsearch cluster; where the first matching result is specifically used to represent that the target primary key field exists in the index document corresponding to the physical index under the target alias in the Elasticsearch cluster.
[0115] It can be understood that the data writing module 330 is specifically used to instruct the Elasticsearch cluster to update the row data containing the target primary key field in the index document corresponding to the target physical index under the index alias to the target data to be written in a row based on the first matching result.
[0116] It can be understood that the index determination module 320 is specifically further used to determine the hot index under the target alias in the Elasticsearch cluster based on the second matching result that the target primary key field does not exist under the target alias in the Elasticsearch cluster; determine the physical index corresponding to the hot index as the target physical index corresponding to the target primary key field under the target alias in the Elasticsearch cluster; where the second matching result is specifically used to represent that the target primary key field does not exist in the index documents corresponding to all the physical indexes under the target alias in the Elasticsearch cluster, and for each lifecycle index corresponding to all the physical indexes, each lifecycle index is one of a hot index, a warm index, and a cold index.
[0117] It can be understood that the data writing module 330 is specifically further configured to, based on the second matching result, instruct the Elasticsearch cluster to add the target data of the row to be written to the index document corresponding to the target physical index under the target alias.
[0118] It can be understood that the data writing device provided by the embodiment of the present invention may further include an instruction generation module, configured to receive the data of the row to be written transmitted after real-time calculation by the Fllink framework; determine that the cumulative byte count of the received multiple data of the rows to be written reaches the byte count threshold, and automatically generate a batch data writing instruction.
[0119] It can be understood that the instruction generation module is specifically further configured to determine that the cumulative duration of the received multiple data of the rows to be written reaches the preset duration, and automatically generate a batch data writing instruction.
[0120] The data writing device 300 provided by the embodiment of the present invention can execute the technical solution of the data writing method in any of the above embodiments. Its implementation principle and beneficial effects are similar to those of the data writing method. For details, please refer to the implementation principle and beneficial effects of the data writing method, which will not be elaborated here.
[0121] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0122] Figure 4 An example of the physical structure diagram of an electronic device is shown as Figure 4 shown. The electronic device 400 may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call the logic instructions in the memory 430 to execute the following method:
[0123] In response to a data writing instruction, determine the target primary key field and the target alias included in the row data to be written; based on the matching results of the target primary key field and the target alias with the Elasticsearch cluster, determine the target physical index from the physical indexes under the target alias in the Elasticsearch cluster; replace the target alias with the target physical index, determine the target row data to be written, and send the target row data to be written to the Elasticsearch cluster; wherein, the Elasticsearch cluster is used to write the target row data to be written into the index document corresponding to the target physical index.
[0124] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the related technology, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0125] On the other hand, an embodiment of the present invention discloses a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided in the above-mentioned method embodiments, for example, including:
[0126] In response to a data writing instruction, determine the target primary key field and the target alias included in the row data to be written; based on the matching results of the target primary key field and the target alias with the Elasticsearch cluster, determine the target physical index from the physical indexes under the target alias in the Elasticsearch cluster; replace the target alias with the target physical index, determine the target row data to be written, and send the target row data to be written to the Elasticsearch cluster; wherein, the Elasticsearch cluster is used to write the target row data to be written into the index document corresponding to the target physical index.
[0127] In another aspect, an embodiment of the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the transmission methods provided in the above embodiments, for example, including:
[0128] In response to a data writing instruction, determine the target primary key field and the target alias included in the row data to be written; based on the matching results of the target primary key field and the target alias with the Elasticsearch cluster, determine the target physical index from the physical indexes under the target alias in the Elasticsearch cluster; replace the target alias with the target physical index to determine the target row data to be written, and send the target row data to be written to the Elasticsearch cluster; wherein, the Elasticsearch cluster is configured to write the target row data to be written into the index document corresponding to the target physical index.
[0129] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0130] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0131] Finally, it should be noted that the above embodiments are only used to illustrate the present invention, rather than to limit the present invention. Although the present invention has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that various combinations, modifications, or equivalent replacements of the technical solutions of the present invention do not depart from the spirit and scope of the technical solutions of the present invention, and should all be covered by the scope of the claims of the present invention.
Claims
1. A data writing method, characterized in that, applied to the Flink-Elasticsearch connector, the method includes: In response to a data writing instruction, determining a target primary key field and a target alias included in the to-be-written row data; wherein, the to-be-written row data is row data generated after processing data under the Flink framework; Based on the matching results of the target primary key field and the target alias with the Elasticsearch cluster, determining a target physical index from the physical indexes under the target alias in the Elasticsearch cluster; Replacing the target alias in the to-be-written row data with the target physical index, determining target to-be-written row data, and sending the target to-be-written row data to the Elasticsearch cluster; wherein, the target to-be-written row data includes the target primary key field and the target physical index; wherein, the Elasticsearch cluster is used to write the target to-be-written row data into the index document corresponding to the target physical index; The data writing instruction is a batch data writing instruction. Before responding to the data writing instruction and determining the target primary key field and the target alias included in the to-be-written row data, the method further includes: Receiving the to-be-written row data transmitted after real-time calculation by the Flink framework; Determining that the cumulative byte count of the received multiple to-be-written row data reaches a byte count threshold, and automatically generating a batch data writing instruction.
2. The data writing method according to claim 1, characterized in that, The determining the target physical index from the physical indexes under the target alias in the Elasticsearch cluster based on the matching results of the target primary key field and the target alias with the Elasticsearch cluster includes: Sending the target primary key field and the target alias to the Elasticsearch cluster, and the Elasticsearch cluster is used to match the target primary key field with the index document corresponding to the physical index under the target alias; Receiving the matching result fed back by the Elasticsearch cluster, and based on the matching result, determining the target physical index from the physical indexes under the target alias in the Elasticsearch cluster.
3. The data writing method according to claim 2, characterized in that, The determining the target physical index from the physical indexes under the target alias in the Elasticsearch cluster based on the matching result includes: Based on the first matching result that the target primary key field already exists in the Elasticsearch cluster under the target alias, determining the physical index corresponding to the index document in the Elasticsearch cluster under the target alias where the target primary key field exists as the target physical index; Among them, the first matching result is specifically used to characterize that the target primary key field exists in the index document corresponding to the physical index under the target alias in the Elasticsearch cluster.
4. The data writing method according to claim 3, wherein, the method further includes: Based on the first matching result, instruct the Elasticsearch cluster to update the row data containing the target primary key field in the index document corresponding to the target physical index under the target alias to the target row data to be written.
5. The data writing method according to claim 2, wherein, The determining the target physical index from the physical indexes under the target alias in the Elasticsearch cluster based on the matching result further includes: Based on the second matching result that the target primary key field does not exist in the Elasticsearch cluster under the target alias, determine the hot index in the Elasticsearch cluster under the target alias; Determine the physical index corresponding to the hot index as the target physical index; Among them, the second matching result is specifically used to characterize that the target primary key field does not exist in the index documents corresponding to all the physical indexes under the target alias in the Elasticsearch cluster, and the life cycle indexes corresponding to all the physical indexes, and each life cycle index is respectively one of a hot index, a warm index, and a cold index.
6. The data writing method according to claim 5, wherein, the method further includes: Based on the second matching result, instruct the Elasticsearch cluster to add the target row data to be written to the index document corresponding to the target physical index under the target alias.
7. The data writing method according to claim 1, wherein, After receiving the row data to be written transmitted after real-time calculation by the Flink framework, the method further includes: Determine that the cumulative duration of the received multiple rows of data to be written reaches a preset duration, and automatically generate the batch data writing instruction.
8. A data writing device, wherein, Applied to the Flink-Elasticsearch connector, the device includes: An information determination module, configured to determine a target primary key field and a target alias included in the row data to be written in response to a data writing instruction; wherein, the row data to be written is row data generated after processing the data under the Flink framework; An index determination module, configured to determine a target physical index from the physical indexes under the target alias in the Elasticsearch cluster based on the matching result between the target primary key field and the target alias and the Elasticsearch cluster; A data writing module, configured to replace the target alias in the to-be-written row data with the target physical index, determine the target to-be-written row data, and send the target to-be-written row data to the Elasticsearch cluster; wherein, the target to-be-written row data includes the target primary key field and the target physical index; Wherein, the Elasticsearch cluster is configured to write the target to-be-written row data into an index document corresponding to the target physical index; The data writing instruction is a batch data writing instruction, and the data writing device further includes an instruction generation module, configured to receive the to-be-written row data transmitted after real-time calculation by the Flink framework; determine that the cumulative byte count of the received multiple to-be-written row data reaches a byte count threshold, and automatically generate a batch data writing instruction.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, it implements the data writing method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, it implements the data writing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data writing method, electronic equipment and storage medium
CN116860742A