Data writing method, data migration method, device and electronic equipment
By leveraging distributed file systems in the Elasticsearch cluster, reducing word segmentation parsing and separate data storage of replica shards, the problems of high CPU consumption, high storage costs and slow scaling are solved, and more efficient data writing and migration are achieved.
Patent Information
- Application Number
- CN202111248654.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-26
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-10-26
AI Technical Summary
In the scenarios of massive data storage and high concurrent writes, the Elasticsearch cluster has problems such as high CPU resource consumption, high storage costs and slow scaling.
By using replica shards to provide query services using the storage directory of the main shard in the distributed file system in the elastic search computing node, the word segmentation analysis process and separate data storage of replica shards are reduced.
It reduces CPU resource consumption and storage costs, improves the capacity expansion performance and data equalization speed of ES computing nodes.
Smart Images

Figure CN113986878B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage, and in particular to a data writing method, a data migration method, a device, an electronic device and a readable storage medium. Background Art
[0002] Currently, the open-source Elasticsearch (ES, Elastic Search) cluster has the following problems in the scenarios of massive data storage and high-concurrency writing: 1. High-level CPU resource consumption: To ensure the reliability of the cluster, the data stored in the ES cluster is generally in the form of double replicas or triple replicas. During the writing process of ES, the same data needs to be parsed on the primary shard and the replica shards simultaneously, resulting in multiple times of CPU resource consumption. 2. High storage cost: In the scenario of storing massive logs, data of TB level or even PB level is usually encountered. In such a high-data-volume situation, to avoid data loss and ensure the reliability of the cluster, double replicas or even triple replicas are used for guarantee, thus resulting in multiple times of high storage resource consumption. 3. Slow scaling: When the computing resources of the cluster are insufficient, if it is necessary to scale the ES cluster, the data will be migrated and balanced between different ES nodes. The data is transmitted through the network, and it takes hours for TB-level data to be rebalanced.
[0003] Therefore, how to reduce the CPU resource consumption and storage cost and improve the expansion performance of ES computing nodes on the basis of ensuring the reliability of the ES cluster is an urgent problem to be solved nowadays. Summary of the Invention
[0004] The purpose of the present invention is to provide a data writing method, a data migration method, a device, an electronic device and a readable storage medium to reduce the CPU resource consumption and storage cost and improve the expansion performance of ES computing nodes on the basis of ensuring the reliability of the ES cluster.
[0005] To solve the above technical problems, the present invention provides a data writing method, including:
[0006] The first primary shard in the Elasticsearch computing node receives the primary shard writing data sent by the target Elasticsearch computing node; wherein, the first primary shard is any primary shard in the Elasticsearch computing node;
[0007] Send the index structure data to the first replica shard; wherein, the index structure data is obtained by parsing the primary shard writing data, and the first replica shard is the replica shard corresponding to the first primary shard;
[0008] Store the stored data in the storage directory corresponding to the first master shard in the distributed file system; wherein, the stored data includes the master shard write data and the index structure data.
[0009] In this solution, the replica shards use the storage directory of the master shard in the distributed file system to provide query services, reducing the word segmentation and parsing process and separate data storage of the replica shards, and reducing the CPU resource consumption and storage cost.
[0010] Optionally, the data writing method further includes:
[0011] The second replica shard in the elastic search computing node updates the saved replica index structure data according to the index structure data sent by the second master shard; wherein, the second replica shard is any replica shard in the elastic search computing node, and the second master shard is the master shard corresponding to the second replica shard;
[0012] Query the master shard write data in the storage directory corresponding to the second master shard in the distributed file system by using the replica index structure data.
[0013] In this solution, the replica shards update the replica index structure data stored by themselves according to the index structure data sent by the corresponding master shards, so as to be able to provide query services by using the latest replica index structure data, ensuring the accuracy of data query.
[0014] Optionally, when the target elastic search computing node is the elastic search computing node, before the first master shard in the elastic search computing node receives the master shard write data sent by the target elastic search computing node, it further includes:
[0015] The elastic search computing node receives the master shard write data sent by the client device;
[0016] Parse the master shard write data to determine the first master shard corresponding to the master shard write data;
[0017] Send the master shard write data to the first master shard.
[0018] In this solution, each elastic search computing node can parse the data to be written sent by the client device, automatically identify the data writing location, and ensure the efficiency of data writing.
[0019] Optionally, the data writing method further includes:
[0020] The second replica shard in the elastic search computing node is upgraded to a master shard according to the upgrade instruction sent by the elastic search master node; wherein, the second replica shard is any replica shard in the elastic search computing node.
[0021] In this solution, when the replica shards in the Elasticsearch computing nodes fail, they can be quickly upgraded to primary shards under the control of the Elasticsearch master node, ensuring the reliability of the Elasticsearch cluster.
[0022] Optionally, the distributed file system is specifically a distributed file system that uses erasure coding storage services.
[0023] In this solution, using a distributed file system that uses erasure coding storage services can reduce the storage cost of data.
[0024] Optionally, when the Elasticsearch computing node is the master node of the Elasticsearch cluster, the data writing method further includes:
[0025] After adding a new Elasticsearch computing node to the Elasticsearch cluster, determine the migration primary shards in the primary shards of the old Elasticsearch computing nodes;
[0026] Control the new Elasticsearch computing node to create replica shards corresponding to the migration primary shards; among them, the replica shards corresponding to the migration primary shards store the index structure data stored in their respective corresponding migration primary shards;
[0027] Migrate the stored data in the storage directory corresponding to the migration primary shards in the distributed file system to their respective target storage directories; among them, the target storage directory is the storage directory of the replica shards in each of the new Elasticsearch computing nodes corresponding to each of the migration primary shards in the distributed file system;
[0028] Upgrade the replica shards corresponding to the migration primary shards in the new Elasticsearch computing node to primary shards and close the migration primary shards.
[0029] In this solution, by migrating the stored data in the storage directory corresponding to the migration primary shards in the distributed file system to their respective target storage directories, the distributed file system is used to implement the migration of the stored data corresponding to the primary shards, improving the data balancing speed.
[0030] Optionally, the step of migrating the stored data in the storage directory corresponding to the migration primary shards in the distributed file system to their respective target storage directories includes:
[0031] Control the migration primary shards to stop processing their respective newly added write requests;
[0032] After waiting for the processing of the current write requests of the migration primary shards to be completed, migrate the stored data in the storage directory corresponding to the migration primary shards in the distributed file system to their respective target storage directories;
[0033] Correspondingly, after upgrading the replica shard corresponding to the migrated primary shard in the new Elasticsearch compute node to the primary shard and closing the migrated primary shard, the method further includes:
[0034] Controlling the primary shards in the new Elasticsearch compute node to resume processing their respective newly added write requests.
[0035] In this solution, during the migration of the primary shard, by controlling the migrated primary shard to stop processing its respective newly added write requests, the impact of the data migration of the migrated primary shard on the processing of write requests is reduced.
[0036] The present invention further provides a data writing device applied to an Elasticsearch compute node, including:
[0037] A data receiving module, configured to receive primary shard write data sent by a target Elasticsearch compute node by using a first primary shard; wherein, the first primary shard is any primary shard in the ES compute node;
[0038] A data sending module, configured to send the index structure data to the first replica shard; wherein, the index structure data is obtained by parsing the primary shard write data, and the first replica shard is the replica shard corresponding to the first primary shard;
[0039] A data storage module, configured to store storage data in a storage directory corresponding to the first primary shard in a distributed file system; wherein, the storage data includes the primary shard write data and the index structure data.
[0040] The present invention further provides a data migration method, including:
[0041] After adding a new Elasticsearch compute node to an Elasticsearch cluster, an Elasticsearch master node determines a migrated primary shard among the primary shards of an old Elasticsearch compute node;
[0042] Controlling the new Elasticsearch compute node to create a replica shard corresponding to the migrated primary shard; wherein, the replica shard stores the index structure data stored in its respective corresponding migrated primary shard;
[0043] Migrating the storage data in the storage directory corresponding to the migrated primary shard in the distributed file system to their respective target storage directories; wherein, the target storage directory is the storage directory of the replica shard in each of the new Elasticsearch compute nodes corresponding to each of the migrated primary shards in the distributed file system;
[0044] Upgrading the replica shard corresponding to the migrated primary shard in the new Elasticsearch compute node to the primary shard, and closing the migrated primary shard.
[0045] In this solution, by migrating the stored data in the storage directory corresponding to the migrating primary shard in the distributed file system to their respective target storage directories, the distributed file system is used to implement the migration of the stored data corresponding to the primary shard, improving the data balancing speed.
[0046] Optionally, the migrating of the stored data in the storage directory corresponding to the migrating primary shard in the distributed file system to their respective target storage directories includes:
[0047] Controlling the migrating primary shards to stop processing their respective newly added write requests;
[0048] After waiting for the processing of the current write requests of the migrating primary shards to be completed, migrating the stored data in the storage directory corresponding to the migrating primary shard in the distributed file system to their respective target storage directories;
[0049] Correspondingly, after upgrading the replica shards corresponding to the migrating primary shards in the new elastic search computing nodes to primary shards and shutting down the migrating primary shards, it further includes:
[0050] Controlling the primary shards in the new elastic search computing nodes to resume processing their respective newly added write requests.
[0051] In this solution, by controlling the migrating primary shards to stop processing their respective newly added write requests during the migration of the primary shards, the impact of the data migration of the migrating primary shards on the processing of write requests is reduced.
[0052] Optionally, the migrating of the stored data in the storage directory corresponding to the migrating primary shard in the distributed file system to their respective target storage directories includes:
[0053] Using the cut interface of the distributed file system to migrate the stored data in the storage directory corresponding to the migrating primary shard in the distributed file system to their respective target storage directories.
[0054] In this solution, using the cut interface of the distributed file system to implement the migration of the stored data corresponding to the migrating primary shard ensures the data migration speed.
[0055] The present invention also provides a data migration device, which is applied to an elastic search master node and includes:
[0056] A migration determination module, configured to determine the migrating primary shards in the primary shards of the old elastic search computing nodes after adding new elastic search computing nodes to the elastic search cluster;
[0057] A replica creation module, configured to control the new Elasticsearch computing node to create replica shards corresponding to the migrated primary shards; wherein, each of the replica shards stores the index structure data stored in the corresponding migrated primary shard;
[0058] A data migration module, configured to migrate the stored data in the storage directory corresponding to the migrated primary shard in the distributed file system to the corresponding target storage directory; wherein, the target storage directory is the storage directory of the replica shards in each of the new Elasticsearch computing nodes corresponding to the respective migrated primary shards in the distributed file system;
[0059] A replica upgrade module, configured to upgrade the replica shards corresponding to the migrated primary shards in the new Elasticsearch computing node to primary shards, and shut down the migrated primary shards.
[0060] The present invention also provides an electronic device, including:
[0061] A memory, configured to store a computer program;
[0062] A processor, configured to implement the steps of the data writing method or the data migration method as described above when executing the computer program.
[0063] In addition, the present invention also provides a readable storage medium, on which a computer program is stored, and the computer program implements the steps of the data writing method or the data migration method as described above when executed by a processor.
[0064] A data writing method provided by the present invention includes: a first primary shard in an Elasticsearch computing node receives primary shard write data sent by a target Elasticsearch computing node; wherein, the first primary shard is any primary shard in the Elasticsearch computing node; sending index structure data to a first replica shard; wherein, the index structure data is obtained by parsing the primary shard write data; storing the stored data in the storage directory corresponding to the first primary shard in the distributed file system; wherein, the first replica shard is the replica shard corresponding to the first primary shard, and the stored data includes the primary shard write data and the index structure data;
[0065] It can be seen that, in the present invention, by sending the index structure data to the first replica shard and storing the stored data in the storage directory corresponding to the first master shard in the distributed file system, the replica shard can utilize the received index structure data and the storage directory of the master shard in the distributed file system to provide query services, reducing the word segmentation and parsing process and separate data storage of the replica shard, reducing the CPU resource consumption and storage cost, so that the data balancing speed can be improved by using the distributed file system, and the expansion performance of the ES computing node can be enhanced. In addition, the present invention also provides a data writing device, a data migration method, a device, an electronic device and a readable storage medium, which also have the above beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0067] Figure 1 It is a flowchart of a data writing method provided by an embodiment of the present invention;
[0068] Figure 2 It is a structural block diagram of a data writing device provided by an embodiment of the present invention;
[0069] Figure 3 It is a flowchart of a data migration method provided by an embodiment of the present invention;
[0070] Figure 4 It is a structural block diagram of a data migration device provided by an embodiment of the present invention;
[0071] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention;
[0072] Figure 6 It is a specific structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0074] Please refer to Figure 1 , Figure 1 which is a flowchart of a data writing method provided by an embodiment of the present invention. The method may include:
[0075] Step 101: The first primary shard in the ES computing node receives the primary shard write data sent by the target ES computing node; wherein, the first primary shard is any primary shard in the ES computing node.
[0076] Wherein, the ES computing node in this embodiment may be a computing resource node in the ES cluster, and one or more shards (such as primary shards and replica shards) are set in each ES computing node. The first primary shard in this embodiment may be any primary shard in the ES computing node.
[0077] It can be understood that this embodiment takes the write request processing process (i.e., data writing process) of one primary shard (i.e., the first primary shard) in one ES computing node in the ES cluster as an example for display. For the write request processing processes of other primary shards in this ES computing node and other ES computing nodes in the ES cluster, they can be implemented in the same or similar manner as the method provided in this embodiment, and this embodiment does not make any restrictions on this.
[0078] Specifically, the system architecture of the data writing method provided in this embodiment may include: 1. The master node (Master) cluster in the top-level ES cluster. Only one master node is effective at the same time in the master node cluster as the ES master node; other master nodes in the master node cluster can all be standby master nodes and can back up the ES master node; the ES master node can manage all shards in all ES computing nodes in the ES cluster, including primary shards (Primary Shard) and replica shards (Replica Shard); the ES master node can save the metadata (Meta) information of all shards, such as whether each shard is a primary shard and the ES server instance (i.e., ES computing node) where each shard is located, etc. information. 2. The primary shard in the ES computing node can provide a complete full-text retrieval service, perform word segmentation and parsing on the incoming data, construct index structure data such as an inverted index, store the data, and complete the requirements of the query service. 3. The replica shard in the ES computing node can synchronize the index structure data from the corresponding primary shard and only provide query services. When the corresponding primary shard fails for various reasons (such as insufficient memory, server downtime), it can be quickly upgraded to a primary shard. 4. The underlying distributed file system can expose itself as a unified storage space, so that different ES servers can query the same storage directory. 5. The bottommost physical disk can provide physical storage resources for the distributed file system.
[0079] It should be noted that the target ES computing node in this step can be the ES computing node that sends the data to be written (i.e., the data to be written by the primary shard) required by the first primary shard to the first primary shard; the data to be written by the primary shard in this step can be the data to be written that needs to be written by the first primary shard and is distributed by the target ES computing node; for example, after the data to be written in the write request is written from the client device to the target ES computing node (such as an ES server), the target ES computing node can parse the written data to be written to determine which primary shard the data to be written belongs to; the target ES computing node can distribute the data to be written to the determined primary shard (such as the first primary shard), rather than the primary shard and the corresponding replica shards, that is, the data to be written will be used as the data to be written by the primary shard and will only be distributed to the corresponding primary shard, rather than the existing primary shard and replica shards, so as to reduce the CPU resource consumption for word segmentation parsing in the replica shards and the storage resource consumption for subsequent storage.
[0080] Correspondingly, the method provided in this embodiment may further include the distribution process of the ES computing node for the data to be written. For example, the ES computing node can receive the data to be written sent by the client device; parse the data to be written to determine the target primary shard; and send the data to be written to the target primary shard; where the target primary shard is the primary shard corresponding to the data to be written. For example, when the ES computing node in this step is the target computing node, after receiving the data to be written by the primary shard (i.e., the data to be written) sent by the client device, the ES computing node can parse the data to be written to determine the first primary shard corresponding to the data to be written, and then send the data to be written by the primary shard to the first primary shard to store the data to be written by the primary shard using the first primary shard.
[0081] Specifically, for the specific manner in which the above ES computing node parses the data to be written to determine the target primary shard, it can be set by the designer according to the practical scenario and user requirements. For example, it can be implemented in the same or similar manner as the method for determining the primary shard to which the data to be written belongs in the prior art. This embodiment does not impose any restrictions on this.
[0082] Step 102: Send the index structure data to the first replica shard; where the index structure data is obtained by parsing the data to be written by the primary shard, and the first replica shard is the replica shard corresponding to the first primary shard.
[0083] It can be understood that the index structure data in this step can be the data structure (i.e., the index structure data) that can be retrieved by full text for the ES computing node to parse the data written to the primary shard. Correspondingly, before this step, it can also include the first primary shard of the ES computing node parsing the data written to the primary shard to obtain the index structure data corresponding to the data written to the primary shard. For example, the ES computing node can use the first primary shard to perform word segmentation and parsing on the data written to the primary shard to generate the index structure data.
[0084] Specifically, for the specific method of the first primary shard of the above ES computing node parsing the data written to the primary shard to obtain the index structure data corresponding to the data written to the primary shard, it can be set by the designer according to the practical scenario and user requirements. For example, it can be implemented in the same or similar way as the word segmentation and parsing method of the shard for the written data in the prior art. As long as the first primary shard can obtain the index structure data (such as the inverted index) that can be retrieved by full text corresponding to the data written to the primary shard, this embodiment does not make any restrictions.
[0085] It should be noted that the first replica shard in this step can be the replica shard of the first primary shard. This embodiment does not limit the specific location of the first replica shard. For example, the first replica shard may not be on the same ES computing node as the first primary shard.
[0086] Among them, the first primary shard in the ES computing node in this step can use the network to synchronize the processed index structure data to the corresponding replica shard (i.e., the first replica shard), so that the first replica shard does not need to perform word segmentation and parsing on the data written to the primary shard by itself, reducing the CPU resource consumption of the ES cluster and improving the writing performance. And by storing the stored data in the storage directory corresponding to the first primary shard in the distributed file system and using the distributed file system to store the stored data, both the first primary shard and the first replica shard can use the index structure data saved in their respective memories and the stored data in the storage directory corresponding to the first primary shard in the distributed file system to provide query services.
[0087] That is to say, the first replica shard can update the index structure data (i.e., the replica index structure data) saved by itself according to the index structure data sent by the first primary shard; according to the obtained query request, use the replica index structure data saved by itself to query the data written to the primary shard in the storage directory corresponding to the first primary shard in the distributed file system to provide query services. Correspondingly, the first primary shard can use the index structure data saved by itself according to the obtained query request to query the data written to the primary shard in the storage directory corresponding to itself in the distributed file system to provide query services.
[0088] Correspondingly, when the ES computing node where the first primary shard is located in this embodiment includes a replica shard, the method provided in this embodiment may further include the process in which the ES computing node uses the replica shard to provide query services. For example, the second replica shard in the ES computing node updates the saved replica index structure data according to the index structure data sent by the second primary shard; uses the replica index structure data to query the data written to the primary shard in the storage directory corresponding to the second primary shard in the distributed file system; where the second replica shard is any replica shard in the ES computing node, and the second primary shard is the primary shard corresponding to the second replica shard.
[0089] Step 103: Store the stored data in the storage directory corresponding to the first primary shard in the distributed file system; where the first replica shard is the replica shard corresponding to the first primary shard, and the stored data includes the data written to the primary shard and the index structure data.
[0090] It should be noted that in this embodiment, by storing the stored data in the storage directory corresponding to the first primary shard in the distributed file system, the first replica shard corresponding to the first primary shard can directly use the index structure data saved in their respective memories and the stored data in the storage directory corresponding to the first primary shard in the distributed file system to provide query services; that is to say, the replica shards in the ES cluster can only provide query services, and when the corresponding primary shard fails for various reasons (such as insufficient memory, server downtime), it can be quickly upgraded to a primary shard to provide a complete full-text retrieval service. For example, the ES master node can, after a certain primary shard in the ES computing node fails, control a replica corresponding to the primary shard to be upgraded to a primary shard by outputting an upgrade instruction.
[0091] Specifically, in this embodiment, each primary shard can correspond to one or more replica shards, so that after the primary shard fails, one of its corresponding replica shards can be upgraded to a primary shard to continue to provide a complete full-text retrieval service instead of the failed primary shard.
[0092] Correspondingly, when the ES computing node where the first primary shard is located in this embodiment includes a replica shard, the method provided in this embodiment may further include the upgrade process of the replica shard in the ES computing node. For example, the second replica shard in the ES computing node is upgraded to a primary shard according to the upgrade instruction sent by the ES master node; where the second replica shard is any replica shard in the ES computing node; that is to say, the ES master node can, after detecting that the primary shard corresponding to the second replica shard fails, control the second replica shard to be upgraded to a primary shard by sending an upgrade instruction to the second replica shard and provide a complete full-text retrieval service.
[0093] Specifically, in this embodiment, after the first primary shard of the ES computing node sends the index structure data to the first replica shard and stores the stored data in the storage directory corresponding to the first primary shard in the distributed file system, it can return a write success message to the target ES computing node, so that the target ES computing node can return a write success message to the corresponding client device according to the received write success message.
[0094] Furthermore, the distributed file system in this embodiment can specifically be a distributed file system that uses erasure coding (such as 8+2 erasure coding or 12+4 erasure coding, etc.) for storage services, so as to use erasure coding to ensure the reliability of the stored data, and compared with the existing two-replica or three-replica solutions for ensuring data reliability, the storage cost is reduced; when using a distributed file system with 8+2 erasure coding storage service, if the single-replica storage cost is 1, the two-replica storage cost is 2, and the storage cost of 8+2 erasure coding is (8+2) / 8 = 1.25, and the storage cost is reduced by (2 - 1.25) / 2 * 100% = 37.5%; compared with the existing three-replica solution for ensuring data reliability, the storage cost can be reduced by (3 - 1.25) / 3 * 100% = 58%.
[0095] Moreover, when writing data to the ES cluster in this embodiment, a mechanism of only writing to the primary shard is used. Compared with double replicas, the CPU resources used for parsing the original written data can be reduced by 50%, improving the write performance; in the performance test of the same data three-node cluster, the write performance of this embodiment is 30% higher than that of the double-replica solution for the ES cluster; and it can avoid the situation that a single-replica ES cluster cannot guarantee reliability and cannot be used in the production environment, ensuring the reliability of the ES cluster.
[0096] In this embodiment, the embodiment of the present invention sends the index structure data to the first replica shard and stores the stored data in the storage directory corresponding to the first primary shard in the distributed file system, so that the replica shard can use the received index structure data and the storage directory of the primary shard in the distributed file system to provide query services, reducing the word segmentation parsing process and separate data storage of the replica shard, reducing the CPU resource consumption and storage cost, and thus being able to use the distributed file system to improve the data balancing speed and enhance the expansion performance of the ES computing node.
[0097] Based on the above embodiments, the data writing method provided in this embodiment may further include the data migration process of the ES cluster; when the ES computing node is the ES master node of the ES cluster, after adding a new ES computing node to the ES cluster, the ES computing node may determine the migration primary shard in the primary shards of the old ES computing node; control the new ES computing node to create a replica shard corresponding to the migration primary shard; wherein, the replica shards corresponding to the migration primary shard store the index structure data stored in their respective corresponding migration primary shards; migrate the stored data in the storage directory corresponding to the migration primary shard in the distributed file system to their respective corresponding target storage directories; wherein, the target storage directory is the storage directory of the replica shards in the respective new ES computing nodes corresponding to each migration primary shard in the distributed file system; upgrade the replica shards corresponding to the migration primary shard in the new ES computing node to primary shards, and close the migration primary shard, so as to complete the data migration corresponding to the primary shard migration when a new ES computing node is added to the ES cluster, and improve the data balancing speed.
[0098] Further, the process of the ES master node migrating the stored data in the storage directory corresponding to the migration primary shard in the distributed file system to their respective corresponding target storage directories may include: controlling the migration primary shard to stop processing their respective newly added write requests; after waiting for the current write requests of the migration primary shards to be processed, migrating the stored data in the storage directory corresponding to the migration primary shard in the distributed file system to their respective corresponding target storage directories; correspondingly, after the ES master node upgrades the replica shards corresponding to the migration primary shard in the new ES computing node to primary shards and closes the migration primary shard, it may also control the primary shards in the new ES computing node to resume processing their respective newly added write requests, so as to reduce the impact of the data migration of the migration primary shard on the write request processing by controlling the migration primary shard to pause processing their respective newly added write requests during the primary shard migration process.
[0099] Specifically, the ES master node may use the cut interface (mv interface) of the distributed file system to migrate the stored data in the storage directory corresponding to the migration primary shard in the distributed file system to their respective corresponding target storage directories, so as to use the cut interface of the distributed file system to implement the migration of the stored data corresponding to the migration primary shard, ensuring the data migration speed.
[0100] Correspondingly, when the ES computing node is not the ES master node of the ES cluster, the data writing method provided in this embodiment may further include the creation and upgrade process of the first primary shard of the ES computing node. For example, when a new ES computing node joins the ES cluster, it may create a replica shard corresponding to the original first primary shard of another ES cluster according to the replica shard creation instruction sent by the ES master node; and after the original first primary shard is closed, upgrade the replica shard corresponding to the original first primary shard to the primary shard according to the replica shard upgrade instruction sent by the ES master node to obtain the first primary shard.
[0101] Corresponding to the above method embodiment, an embodiment of the present invention further provides a data writing device, and a data writing device described below can be correspondingly referred to with a data writing method described above.
[0102] Please refer to Figure 2 , Figure 2 which is a structural block diagram of a data writing device provided in an embodiment of the present invention. The data writing device is applied to an ES computing node and may include:
[0103] A data receiving module 10, configured to receive the primary shard writing data sent by the target ES computing node by using the first primary shard; wherein, the first primary shard is any primary shard in the ES computing node;
[0104] A data sending module 20, configured to send the index structure data to the first replica shard; wherein, the index structure data is obtained by parsing the primary shard writing data, and the first replica shard is the replica shard corresponding to the first primary shard;
[0105] A data storage module 30, configured to store the storage data in the storage directory corresponding to the first primary shard in the distributed file system; wherein, the storage data includes the primary shard writing data and the index structure data.
[0106] Optionally, the data writing device may further include:
[0107] An index update module, configured to update the saved replica index structure data by using the second replica shard according to the index structure data sent by the second primary shard; wherein, the second replica shard is any replica shard in the ES computing node, and the second primary shard is the primary shard corresponding to the second replica shard;
[0108] A query service module, configured to query the primary shard writing data in the storage directory corresponding to the second primary shard in the distributed file system by using the replica index structure data.
[0109] Optionally, when the target ES computing node is an ES computing node, the data writing device may further include:
[0110] A write receiving module, configured to receive the main shard write data sent by a client device;
[0111] A write parsing module, configured to parse the main shard write data to determine a first main shard corresponding to the main shard write data;
[0112] A write distributing module, configured to send the main shard write data to the first main shard.
[0113] Optionally, the data writing device may further include:
[0114] A shard upgrading module, configured to upgrade a second replica shard in an ES computing node to a main shard according to an upgrade instruction sent by an ES master node; wherein, the second replica shard is any replica shard in the ES computing node.
[0115] Optionally, the distributed file system may specifically be a distributed file system adopting an erasure code storage service.
[0116] Optionally, when the ES computing node is the ES master node of an ES cluster, the data writing device may further include:
[0117] A determining module, configured to determine a migrating main shard in the main shards of an old ES computing node after a new ES computing node is added to the ES cluster;
[0118] A creating module, configured to control the new ES computing node to create a replica shard corresponding to the migrating main shard; wherein, the replica shard corresponding to the migrating main shard stores index structure data saved in its corresponding migrating main shard;
[0119] A migrating module, configured to migrate the stored data in the storage directory corresponding to the migrating main shard in the distributed file system to their respective target storage directories; wherein, the target storage directory is the storage directory of the replica shards in the new ES computing nodes corresponding to the respective migrating main shards in the distributed file system;
[0120] An upgrading module, configured to upgrade the replica shard corresponding to the migrating main shard in the new ES computing node to a main shard and shut down the migrating main shard.
[0121] Optionally, the migrating module may include:
[0122] A stopping sub-module, configured to control the migrating main shards to stop processing their respective newly added write requests;
[0123] A migrating sub-module, configured to wait until the current write requests of the migrating main shards are processed, and then migrate the stored data in the storage directory corresponding to the migrating main shard in the distributed file system to their respective target storage directories;
[0124] Correspondingly, the data writing device may further include:
[0125] A recovery module, configured to control the master shard in the new ES computing node to recover and process the newly added write requests respectively.
[0126] Optionally, the migration module may be specifically configured to use the cut interface of the distributed file system to migrate the stored data in the storage directory corresponding to the migrated master shard in the distributed file system to the corresponding target storage directory respectively.
[0127] In this embodiment, in the embodiment of the present invention, the index structure data is sent to the first replica shard through the data sending module 20, and the storage data is stored in the storage directory corresponding to the first master shard in the distributed file system through the data storage module 30, so that the replica shard can use the received index structure data and the storage directory of the master shard in the distributed file system to provide query services, reducing the word segmentation and parsing process and separate data storage of the replica shard, reducing the CPU resource consumption and storage cost, and thus being able to improve the data balancing speed and enhance the expansion performance of the ES computing node by using the distributed file system.
[0128] Based on the above embodiments, the embodiment of the present invention further provides a data migration method to implement the data migration of the newly added ES computing node with the distributed file system, enhancing the expansion performance of the ES computing node; a data migration method described below can be correspondingly referred to with a data writing method described above.
[0129] Please refer to Figure 3 , Figure 3 which is a flowchart of a data migration method provided by the embodiment of the present invention. The data migration method may include:
[0130] Step 201: After adding a new ES computing node to the ES cluster, the ES master node determines the migrated master shard in the master shards of the old ES computing nodes.
[0131] Wherein, the new ES computing node in this step may be the newly added ES computing node in the ES cluster; the old ES computing node in this step may be other normally working ES computing nodes outside the newly added ES computing node (i.e., the new ES computing node) in the ES cluster.
[0132] It can be understood that the migrated master shard in this step may be the master shard of the old ES computing node determined by the ES master node that needs to be migrated to the new ES computing node. That is to say, in this step, after the ES master node detects the addition of a new ES computing node (i.e., the new ES computing node) to the ES cluster, it triggers the rebalancing mechanism, calculates the shards (such as master shards and replica shards) that need to be transferred to the new ES computing node according to the scheduling policy, and uses the master shard in the calculated shards as the migrated master shard.
[0133] Specifically, for the specific method of the ES master node determining the migrated primary shard in the primary shards of the old ES computing nodes in this step, it can be set by the designer himself. For example, it can be implemented in the same or similar manner as the shard scheduling and balancing method of the ES master node in the prior art. This embodiment does not impose any restrictions on this.
[0134] It should be noted that this embodiment demonstrates the transfer of primary shards and data migration during the expansion process of ES computing nodes. This embodiment may also include the transfer of replica shards during the expansion process of ES computing nodes. For example, after adding a new ES computing node to the ES cluster, the ES master node determines the migrated replica shards in the replica shards of the old ES computing nodes; controls the new ES computing node to create replica shards corresponding to the migrated replica shards; and after the creation of the replica shards corresponding to the migrated replica shards in the new ES computing node is completed, closes the migrated replica shards.
[0135] Step 202: Control the new ES computing node to create replica shards corresponding to the migrated primary shards; wherein, the replica shards store the index structure data stored in their respective corresponding migrated primary shards.
[0136] It can be understood that the ES master node in this step can construct one replica shard corresponding to each migrated primary shard for the new ES computing node, and synchronize the index structure data in the corresponding migrated primary shards in the memory, so as to realize the synchronization between the migrated primary shards and one replica shard corresponding to each of them created in the new ES computing node.
[0137] Step 203: Migrate the stored data in the storage directory corresponding to the migrated primary shard in the distributed file system to their respective target storage directories; wherein, the target storage directory is the storage directory of the replica shards in the new ES computing nodes corresponding to each migrated primary shard in the distributed file system.
[0138] Specifically, in this step, the ES master node can control the distributed file system to migrate the stored data in the storage directory corresponding to the migrated primary shard to the storage directory of the replica shard corresponding to the new ES computing node, so as to use the distributed file system to realize the data migration of the migrated primary shard, improve the performance of shard rebalancing during the expansion of the ES cluster, and be able to reduce the time consumed for data balancing from 3 hours to less than 1 minute when expanding a computing node in a three-node cluster with 1TB of data; for example, the ES master node can use the cut interface (mv interface) of the distributed file system to migrate the stored data in the storage directory corresponding to the migrated primary shard in the distributed file system to their respective target storage directories.
[0139] Correspondingly, to avoid the impact of the data migration of the migrating primary shards on the processing of write requests, in this step, the ES master node can first control the migrating primary shards to stop processing their respective newly added write requests; after waiting for the processing of the current write requests of the migrating primary shards to be completed, migrate the stored data in the storage directories corresponding to the migrating primary shards in the distributed file system to their respective target storage directories; for example, the ES master node can stop the migrating primary shards from processing newly added write requests after completing the synchronization of one replica shard corresponding to each of the migrating primary shards and the new ES computing nodes in step 202, wait for the processing of their current write requests to be completed, that is, after the write data of their respective primary shards is written, and use the cut interface of the distributed file system to migrate the stored data in the storage directories corresponding to the migrating primary shards in the distributed file system to their respective target storage directories.
[0140] Correspondingly, after step 204, the ES master node can control the upgraded primary shards in the new ES computing nodes to resume processing their respective newly added write requests to resume the services of the migrated primary shards.
[0141] Step 204: Upgrade the replica shards corresponding to the migrating primary shards in the new ES computing nodes to primary shards and close the migrating primary shards.
[0142] It can be understood that in this step, the ES master node can control the replica shards corresponding to the migrating primary shards in the new ES computing nodes to be upgraded to primary shards and close the original migrating primary shards, so as to realize the migration of the primary shards to be migrated in the old ES computing nodes to the new ES computing nodes and realize the migration of the primary shards during the expansion of the ES computing nodes.
[0143] In this embodiment, the embodiment of the present invention migrates the stored data in the storage directories corresponding to the migrating primary shards in the distributed file system to their respective target storage directories, uses the distributed file system to realize the migration of the stored data corresponding to the primary shards, and improves the data balancing speed.
[0144] Corresponding to the above method embodiment, the embodiment of the present invention further provides a data migration device, and a data migration device described below can be correspondingly referred to the data migration method described above.
[0145] Please refer to Figure 4 , Figure 4 which is a structural block diagram of a data migration device provided by the embodiment of the present invention. The data migration device is applied to the ES master node and may include:
[0146] A migration determination module 40, configured to determine the migrating primary shards in the primary shards of the old ES computing nodes after adding new ES computing nodes to the ES cluster;
[0147] A replica creation module 50, configured to control the creation of replica shards corresponding to the migrated primary shards on new ES computing nodes; wherein, each replica shard stores the index structure data stored in the corresponding migrated primary shard.
[0148] A data migration module 60, configured to migrate the stored data in the storage directory corresponding to the migrated primary shard in the distributed file system to their respective target storage directories; wherein, the target storage directory is the storage directory of the replica shards in the new ES computing nodes corresponding to each migrated primary shard in the distributed file system.
[0149] A replica upgrade module 70, configured to upgrade the replica shards corresponding to the migrated primary shards in the new ES computing nodes to primary shards, and shut down the migrated primary shards.
[0150] Optionally, the data migration module 60 may include:
[0151] A write pause sub-module, configured to control the migrated primary shards to stop processing their respective newly added write requests.
[0152] A data migration sub-module, configured to wait for the processing of the current write requests of the migrated primary shards to be completed, and then migrate the stored data in the storage directory corresponding to the migrated primary shard in the distributed file system to their respective target storage directories.
[0153] Correspondingly, the data migration device may further include:
[0154] A write recovery module, configured to control the primary shards in the new ES computing nodes to resume processing their respective newly added write requests.
[0155] Optionally, the data migration module 60 may be specifically configured to use the cut interface of the distributed file system to migrate the stored data in the storage directory corresponding to the migrated primary shard in the distributed file system to their respective target storage directories.
[0156] In this embodiment, in the embodiment of the present invention, the data migration module 60 migrates the stored data in the storage directory corresponding to the migrated primary shard in the distributed file system to their respective target storage directories, and uses the distributed file system to implement the migration of the stored data corresponding to the primary shard, thereby improving the data balancing speed.
[0157] Corresponding to the above method embodiment, the embodiment of the present invention further provides an electronic device, and an electronic device described below can be mutually referred to with a data writing method and a data migration method described above.
[0158] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an electronic device provided by the embodiment of the present invention. The electronic device may include:
[0159] A memory D1 for storing a computer program;
[0160] A processor D2 for implementing the steps of the data writing method or the data migration method provided by the above method embodiments when executing the computer program.
[0161] Specifically, please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPU) 322 (for example, one or more processors) and a memory 332, and one or more storage media 330 for storing application programs 342 or data 344 (for example, one or more mass storage devices). Among them, the memory 332 and the storage medium 330 may be transient storage or persistent storage. The program stored in the storage medium 330 may include one or more units (not shown in the figure), and each unit may include a series of instruction operations on the electronic device. Further, the central processor 322 may be configured to communicate with the storage medium 330 and execute a series of instruction operations in the storage medium 330 on the electronic device 310.
[0162] The electronic device 310 may further include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341. For example, Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0163] Among them, the electronic device 310 may specifically be a server of an ES cluster (i.e., an ES server).
[0164] The steps in the data writing method or the data migration method described above may be implemented by the structure of the electronic device.
[0165] Corresponding to the above method embodiments, an embodiment of the present invention further provides a readable storage medium, and the readable storage medium described below can be correspondingly referred to with the data writing method and the data migration method described above.
[0166] A readable storage medium has a computer program stored thereon, and when the computer program is executed by a processor, the steps of the data writing method or the data migration method of the above method embodiments are implemented.
[0167] Specifically, the readable storage medium may be a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, or other readable storage media that can store program codes.
[0168] The embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.
[0169] The above has introduced in detail a data writing method, a data migration method, a device, an electronic device, and a readable storage medium provided by the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A data writing method, characterized in that, include: The first primary shard in the elastic search computing node receives the primary shard write data sent by the target elastic search computing node; wherein the first primary shard is any primary shard in the elastic search computing node; the target elastic search computing node distributes the data to be written to the determined primary shard; Sending the index structure data to the first replica shard; wherein the index structure data is the full-text searchable index structure data obtained by the first primary shard by parsing the data written by the primary shard, and the first replica shard is the replica shard corresponding to the first primary shard; The storage data is stored in a storage directory corresponding to the first primary shard in the distributed file system; wherein the storage data includes the primary shard write data and the index structure data.
2. The data writing method according to claim 1, characterized in that, Also includes: The second replica shard in the elastic search computing node updates the stored replica index structure data according to the index structure data sent by the second primary shard; wherein the second replica shard is any replica shard in the elastic search computing node, and the second primary shard is the primary shard corresponding to the second replica shard; The replica index structure data is used to query the primary shard write data under the storage directory corresponding to the second primary shard in the distributed file system.
3. The data writing method according to claim 1, characterized in that, When the target elastic search computing node is the elastic search computing node, before the first primary shard in the elastic search computing node receives the primary shard write data sent by the target elastic search computing node, the method further includes: The elastic search computing node receives the primary shard write data sent by the client device; Parsing the data written to the primary shard to determine a first primary shard corresponding to the data written to the primary shard; The data written to the primary shard is sent to the first primary shard.
4. The data writing method according to claim 1, characterized in that, Also includes: The second replica shard in the elastic search computing node is upgraded to the primary shard according to the upgrade instruction sent by the elastic search master node; wherein the second replica shard is any replica shard in the elastic search computing node.
5. The data writing method according to claim 1, characterized in that, The distributed file system is specifically a distributed file system that adopts erasure code storage service.
6. The data writing method according to any one of claims 1 to 5, characterized in that, When the elastic search computing node is an elastic search master node of an elastic search cluster, it also includes: After adding a new elastic search computing node to the elastic search cluster, determining a migration primary shard in the primary shard of the old elastic search computing node; Control the new elastic search computing node to create a replica shard corresponding to the migrated primary shard; wherein the replica shard corresponding to the migrated primary shard stores the index structure data stored in the respective corresponding migrated primary shards; Migrate the storage data in the storage directory corresponding to the migrated primary shard in the distributed file system to the corresponding target storage directory; wherein the target storage directory is the storage directory of the replica shard in the new elastic search computing node corresponding to each of the migrated primary shards in the distributed file system; The replica shard corresponding to the migrated primary shard in the new elastic search computing node is upgraded to the primary shard, and the migrated primary shard is closed.
7. The data writing method according to claim 6, characterized in that, Migrating the stored data in the storage directory corresponding to the migration master shard in the distributed file system to their respective target storage directories includes: Controlling the migration master shards to stop processing their respective newly added write requests; After waiting for the processing of the current write requests of the migration master shards to be completed, migrating the stored data in the storage directory corresponding to the migration master shards in the distributed file system to their respective target storage directories; Correspondingly, after upgrading the replica shards corresponding to the migration master shards in the new Elasticsearch computing nodes to master shards and shutting down the migration master shards, it further includes: Controlling the master shards in the new Elasticsearch computing nodes to resume processing their respective newly added write requests.
8. A data writing device, characterized in that, Applied to Elasticsearch computing nodes, it includes: A data receiving module, configured to use a first master shard to receive the master shard write data sent by a target Elasticsearch computing node; wherein, the first master shard is any master shard in the Elasticsearch computing node; the target Elasticsearch computing node distributes the data to be written to the determined master shard; A data sending module, configured to send the index structure data to a first replica shard; wherein, the index structure data is the index structure data that can be retrieved by full text obtained by the first master shard through parsing the master shard write data, and the first replica shard is the replica shard corresponding to the first master shard; A data storage module, configured to store the stored data in the storage directory corresponding to the first master shard in the distributed file system; wherein, the stored data includes the master shard write data and the index structure data.
9. A data migration method, characterized in that, It includes: After adding a new Elasticsearch computing node to the Elasticsearch cluster, the Elasticsearch master node determines the migration master shards in the master shards of the old Elasticsearch computing nodes; Controlling the new Elasticsearch computing node to create replica shards corresponding to the migration master shards; wherein, the replica shards store the index structure data stored in their respective corresponding migration master shards; the index structure data is the index structure data that can be retrieved by full text obtained by the migration master shard through parsing the master shard write data; Migrating the stored data in the storage directory corresponding to the migration master shards in the distributed file system to their respective target storage directories; wherein, the target storage directory is the storage directory of the replica shards in the respective new Elasticsearch computing nodes corresponding to each of the migration master shards in the distributed file system; Upgrading the replica shards corresponding to the migration master shards in the new Elasticsearch computing nodes to master shards and shutting down the migration master shards.
10. The data migration method according to claim 9, wherein, The migrating the stored data in the storage directory corresponding to the migration master shards in the distributed file system to their respective target storage directories includes: Controlling the migration master shards to stop processing their respective newly added write requests; After waiting for the processing of the current write requests of the migration master shards to be completed, migrating the stored data in the storage directory corresponding to the migration master shards in the distributed file system to their respective target storage directories; Correspondingly, after upgrading the replica shard corresponding to the migrated primary shard in the new Elasticsearch computing node to the primary shard and closing the migrated primary shard, it further includes: Controlling the primary shards in the new Elasticsearch computing node to resume processing their respective newly added write requests.
11. The data migration method according to claim 9, wherein, The migrating the stored data in the storage directory corresponding to the migrated primary shard in the distributed file system to their respective target storage directories includes: Using the cut interface of the distributed file system to migrate the stored data in the storage directory corresponding to the migrated primary shard in the distributed file system to their respective target storage directories.
12. A data migration device, wherein, Applied to the Elasticsearch master node, it includes: A migration determination module, configured to determine the migrated primary shards in the primary shards of the old Elasticsearch computing nodes after adding a new Elasticsearch computing node to the Elasticsearch cluster; A replica creation module, configured to control the new Elasticsearch computing node to create replica shards corresponding to the migrated primary shards; wherein, the replica shards store the index structure data saved in their respective corresponding migrated primary shards; the index structure data is the index structure data that can be retrieved by full text obtained by the migrated primary shard by parsing the data written to the primary shard; A data migration module, configured to migrate the stored data in the storage directory corresponding to the migrated primary shard in the distributed file system to their respective target storage directories; wherein, the target storage directory is the storage directory of the replica shards in the respective corresponding new Elasticsearch computing nodes of each of the migrated primary shards in the distributed file system; A replica upgrade module, configured to upgrade the replica shard corresponding to the migrated primary shard in the new Elasticsearch computing node to the primary shard and close the migrated primary shard.
13. An electronic device, wherein, It includes: A memory, configured to store a computer program; A processor, configured to implement the steps of the data writing method according to any one of claims 1 to 7 or the data migration method according to any one of claims 9 to 11 when executing the computer program.
14. A readable storage medium, wherein, A computer program is stored on the readable storage medium, and when the computer program is executed by the processor, it implements the steps of the data writing method according to any one of claims 1 to 7 or the data migration method according to any one of claims 9 to 11.
Citation Information
Patent Citations
Index allocation method and device for improving performance of ES-based log system
CN110990366A
Data migration method and device of fragmentation cluster and fragmentation cluster system
CN111708763A
ElasticSearch-based data indexing method and device, computer equipment and storage medium
CN111797096A