Data synchronization method based on cloud search service, medium, electronic equipment and product
By calculating the target step size and index number during primary shard migration in the cloud search service, the problem of index file name conflicts is resolved, the index files can be smoothly uploaded to remote storage, and the storage-computing separation architecture is supported.
Patent Information
- Application Number
- CN202511151638.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-08-15
AI Technical Summary
In the Cloud Search service, when a new primary shard creates an index file, the file name may conflict with the index file name of the old primary shard, causing serious problems, especially in remote storage scenarios.
By determining the migration method of the first primary shard, calculating the target step size, and determining the second index number of the second primary shard based on the target step size and the first index number, the index file name of the second primary shard is ensured to be different from the index file name of the first primary shard to avoid conflicts.
During the primary shard migration process, a larger second index number is used to ensure that the index file names in the remote storage do not conflict, ensure that the index files are uploaded smoothly, and support the storage and computing separation architecture.
Smart Images

Figure CN120804045A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, in particular, to a data synchronization method based on cloud search service, medium, electronic device and product. BACKGROUND
[0002] When the primary shard is migrated, the new primary shard creates a new index file, at this time, the file name of the new index file may conflict with the file name of the index file of the old primary shard. Especially in the scenario of using remote storage, the new index file created by the new primary shard needs to be uploaded to the remote storage. If the file name of an index file is the same as the file name of another index file, but the data contents of the two index files are different, this will cause serious problems. Therefore, how to ensure that the file name of the index file does not conflict becomes a technical problem to be solved. SUMMARY
[0003] This summary is provided to introduce a selection of concepts, which will be described with greater specificity in the detailed description section. This summary is not intended to identify key or essential features of the claimed technology, nor is it intended to limit the scope of the claimed technology.
[0004] In a first aspect, the present disclosure provides a data synchronization method based on a cloud search service, comprising: In response to a first primary shard of the cloud search service needing to be migrated, determining a first index number currently used by the first primary shard, the first index number being used by the first primary shard to determine a file name corresponding to an index file written into the first primary shard; According to a migration mode of the first primary shard, determining a target step, and according to the target step and the first index number, determining a second index number corresponding to a second primary shard for replacing the first primary shard, the target step being used to make the second index number greater than the first index number; During the migration of the first primary shard to the second primary shard, based on the second index number, determining a file name corresponding to an index file written into the second primary shard, so that the file name corresponding to the index file written into the second primary shard is different from the file name of the index file of the first primary shard.
[0005] In a second aspect, the present disclosure provides a data synchronization device based on a cloud search service, comprising: The first determining module is configured to determine a first index number currently used by a first master shard in response to the first master shard needing to be migrated, the first index number being used by the first master shard to determine a file name corresponding to an index file written into the first master shard; The second determining module is configured to determine a target step according to a migration mode of the first master shard, and determine a second index number corresponding to a second master shard replacing the first master shard according to the target step and the first index number, the target step being used to make the second index number greater than the first index number. The third determining module is configured to determine, in a process of migrating the first master shard to the second master shard, a file name corresponding to an index file written into the second master shard based on the second index number, so that the file name corresponding to the index file written into the second master shard is different from a file name of an index file of the first master shard.
[0006] In a third aspect, the present disclosure provides a computer readable medium having a computer program stored thereon, the computer program being executed by a processing device to implement the steps of the method of the first aspect.
[0007] In a fourth aspect, the present disclosure provides an electronic device, comprising: a storage device having a computer program stored thereon; a processing device configured to execute the computer program in the storage device to implement the steps of the method of the first aspect.
[0008] In a fifth aspect, the present disclosure provides a computer program product comprising a computer program, the computer program being executed by a processor to implement the steps of the method of the first aspect.
[0009] Based on the above technical scheme, by responding to the migration of the master shard of the cloud search service, the first index number currently used by the first master shard is determined, the target step is determined according to the migration mode of the first master shard, and the second index number corresponding to the second master shard is determined according to the target step and the first index number. Then, in the process of migrating the first master shard to the second master shard, based on the second index number, the file name corresponding to the index file written to the second master shard is determined, so that the file name corresponding to the index file written to the second master shard is different from the file name of the index file of the first master shard. Not only can the file name of the index file written to the second master shard be avoided from conflicting with the file name of the index file in the first master shard, but also in the scenario of using remote storage, the index file created by the second master shard needs to be uploaded to the remote storage. By using a larger second index number to determine the file name, it can be ensured that the file name in the remote storage will not conflict, so that the index file created by the second master shard can be smoothly uploaded to the remote storage.
[0010] Other features and advantages of the present disclosure will be described in detail in the following detailed description section. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and other features, advantages and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments with reference to the attached drawings. The same or similar elements are denoted by the same or similar reference numerals throughout the drawings. It is to be understood that the drawings are schematic, and the original and elements are not necessarily drawn to scale. In the drawings: Figure 1 is an architecture diagram of a cloud search service according to some embodiments.
[0012] Figure 2 is a flowchart of a data synchronization method based on a cloud search service according to some embodiments.
[0013] Figure 3 is a schematic diagram of pre-copying according to some embodiments.
[0014] Figure 4 is a schematic diagram of starting a second master shard by a read-write engine according to some embodiments.
[0015] Figure 5 is a schematic diagram of starting a replica shard by a read-only engine according to some embodiments.
[0016] Figure 6 is a structural schematic diagram of a data synchronization apparatus based on a cloud search service according to some embodiments.
[0017] Figure 7 is a structural schematic diagram of an electronic device according to some embodiments. DETAILED DESCRIPTION
[0018] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0019] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0020] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0021] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0022] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0023] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0024] The data synchronization method based on cloud search service provided by the present disclosure is applicable to cloud search service. Among them, the cloud search service referred to in the embodiment of the present disclosure can be understood as a one-stop information retrieval and analysis platform, which can be applied to business scenarios such as full-text search, vector search, hybrid search, AI search, spatiotemporal retrieval, etc. For example, in a log analysis system, a cloud search service can be used to store and query log data. In an e-commerce platform, a cloud search service can be used to provide a product search function and provide a fast and accurate search experience. In a big data monitoring system, a cloud search service can be used to store and analyze monitoring data to help users discover potential problems in a timely manner.
[0025] In practical applications, the cloud search service can be distributed deployed on multiple nodes (such as cloud servers), and specifically, the cloud search service can be deployed on multiple nodes in a storage-computing separation architecture. The following will combine the accompanying drawings to describe the cloud search service in detail. Figure 1 The storage-computing separation architecture of the cloud search service is described in detail.
[0026] Figure 1 is an architecture diagram of the cloud search service according to some embodiments. As shown in Figure 1 , the cloud search service includes multiple computing nodes and storage nodes, which work together to provide distributed search and analysis functions. Among them, the computing nodes are used to provide data computing capabilities, and the storage nodes are used to deploy remote storage to provide data storage capabilities. As shown in Figure 1 , the computing node 1, the computing node 2, and the computing node N are used as computing nodes to deploy master shards and replica shards, P0, P1, …, PN represent the master shards deployed in a computing node, and R0, R1, …, RN represent the replica shards corresponding to the master shards. It should be noted that the master shards and the corresponding replica shards can be deployed on different nodes. Moreover, the number of computing nodes and the number of storage nodes can be elastically expanded or contracted according to the needs of the cloud search service.
[0027] In the cloud search service, an index is a data structure used to store and search documents, and a single index can store a large amount of data beyond the hardware limit of a single node. For example, an index with 1 billion document data needs to occupy 1 TB of disk space, and any node may not have such a large disk space, or it is too slow to process search requests through a single node. Therefore, in the cloud search service, an index is divided into multiple shards, as shown in Figure 1 , the index can be divided into P0, P1, …, PN, etc. N shards. Each shard saves part of the data in all the data in the index, and each shard itself is equivalent to a functional and independent "index", and each shard is deployed on different nodes.
[0028] Each shard can include a master shard and a replica shard, the master shard is used to store the actual index file and is responsible for reading and writing the index file, the number of master shards is defined when the index is created, and the number of master shards cannot be changed after creation. Each master shard can have zero or more replica shards, and the replica shard is used to copy the index file stored by the master shard and is responsible for reading the index file. Through the replica shard, the high availability, fault tolerance, and query performance of the cloud search service can be improved. The replica shard copies the index file of the corresponding master shard to ensure the redundancy and safety of the index file, and when the master shard fails, the corresponding replica shard can be promoted to a new master shard to continue providing search services.
[0029] In the storage-computation separation architecture shown in Figure 1 In the storage-computation separation architecture shown in
[0030] Therefore, in the storage-computation separation architecture shown in Figure 1 The computing node shown in is responsible for data computation and does not actually store the index file, while the storage node is responsible for actually storing the index file, so that the cloud search service can realize separation of computation and storage, thereby reducing the local storage cost of the computing node.
[0031] The client can send a write request to any computing node in the cloud search service. The computing node receiving the write request acts as a coordination node, which forwards the write request to the computing node of the corresponding master shard. The master shard processes the write request, writes the corresponding index file to the master shard, and synchronizes the index file to the corresponding replica shard. The coordination node returns a response to the client after confirming that the write is successful. Of course, when the master shard processes the write request, the master shard can directly write the index file corresponding to the write request to the remote storage and maintain meta information on the master shard indicating the storage path of the index file in the remote storage. The master shard can also synchronize the meta information to the corresponding replica shard. When the master shard or the replica shard receives a query request sent by the client, the master shard or the replica shard reads the required field value from the remote storage through the meta information, so that the cloud search service can realize separation of storage and computation.
[0032] The cloud search service-based data synchronization method provided by the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0033] Figure 2 is a flowchart of a cloud search service-based data synchronization method according to some embodiments. As shown in Figure 2 The cloud search service-based data synchronization method provided by the embodiments of the present disclosure can be specifically executed by a cloud search service-based data synchronization device. The device can be implemented in the form of software and / or hardware. As shown in Figure 2 The method can include the following steps.
[0034] In step 210, in response to the first master shard of the cloud search service needing to be migrated, a first index number currently used by the first master shard is determined, the first index number being used by the first master shard to determine a file name of an index file written into the first master shard.
[0035] Here, the first master shard of the cloud search service can refer to any one master shard in the cloud search service. The first master shard of the cloud search service needing to be migrated can be triggered in a case that a node where the first master shard is located is unavailable, load balancing is performed again, and the like. For example, the first master shard needing to be migrated can be that the node where the first master shard is located exits, and then a replica shard corresponding to the first master shard is upgraded to a second master shard. Of course, for another example, when the disk space of the node where the first master shard is located is insufficient or load balancing is needed, the first master shard is migrated from one node to another node.
[0036] During the migration of the first master shard, the first master shard is in a state of "being migrated" or "partially available", the first master shard still exists, but new write operations are blocked and need to wait for the migration to be completed. The second master shard is a master shard created on a new node and used to receive data migrated from the first master shard. During the migration of the first master shard, the second master shard starts from a blank state or a state of a replica shard, then receives data from the first master shard, and once the data replication is completed, the second master shard is marked as "active" and starts to process read and write requests. Therefore, when the first master shard needs to be migrated, the second master shard can be understood as a new master shard formed by migration of the old first master shard.
[0037] That is, in the embodiment of the disclosure, the second master shard is a master shard taking over the role of the old first master shard, the second master shard becomes a master shard of the shard corresponding to the old first master shard, the second master shard is used to replace the working logic of the first master shard and is responsible for processing read and write requests, and the old first master shard is deleted after being verified.
[0038] The first index number currently used by the first master shard refers to an index number currently used by the first master shard to determine a file name of an index file written into the first master shard. In the cloud search service, the first master shard maintains a segment counter, which is used to record the number of changes of segment files in the first master shard. Whenever a new segment file is created in the first master shard, the count of the segment counter is increased accordingly. Correspondingly, the first index number currently used by the first master shard can refer to the count value of the segment counter when the first master shard needs to be migrated.
[0039] It should be noted that in the cloud search service, the index data is split into multiple segment files, each of which is a complete, immutable index structure, and has its own dictionary, inverted index and document storage data structure. Therefore, in the embodiments of the present disclosure, the index file can refer to each segment file.
[0040] In step 220, according to the migration mode of the first primary shard, a target step is determined, and according to the target step and the first index number, a second index number corresponding to the second primary shard for replacing the first primary shard is determined, and the target step is used to make the second index number greater than the first index number.
[0041] Here, generally speaking, when a second primary shard needs to be started, the second primary shard is initialized and run with the help of an engine. Therefore, as some examples, the second index number corresponding to the second primary shard can be determined by the engine for starting the second primary shard, according to the first index number currently used by the first primary shard and the target step. The target step is related to the migration mode of the primary shard, and the target step is used to make the second index number greater than the first index number.
[0042] Illustratively, the sum of the first index number and the target step can be determined as the second index number. For example, assuming that the first index number is 100 and the target step is 200000, the second index number can be 200100. In the embodiments of the present disclosure, when the first primary shard needs to be migrated, the second primary shard uses a larger second index number to determine the file name of the index file, thereby avoiding the file name of the index file written to the second primary shard from conflicting with the file name of the index file on the old first primary shard.
[0043] Illustratively, the engine for starting the second primary shard can be a read-write engine. In the embodiments of the present disclosure, the second index number used to determine the file name of the index file can be controlled at the engine level, thereby avoiding the generation of index files with the same file name but different file contents.
[0044] It should be understood that the target step is a dynamic step, and the migration mode of the first primary shard is different, and the target step is also different. In the embodiments of the present disclosure, the migration mode of the first primary shard includes two kinds, one is that the replica shard corresponding to the first primary shard is upgraded to the second primary shard, and the other is that the first primary shard is directly migrated to another node.
[0045] For example, in a case where the migration manner of the first primary shard represents that a replica shard corresponding to the first primary shard is upgraded to a second primary shard, the target step can be calculated by the term of the second primary shard. In a case where the migration manner of the first primary shard represents that the first primary shard is directly migrated to another node, the target step can be determined by the number of migrations of the first primary shard. By the target step related to the migration manner of the first primary shard, the second index number can be dynamically adjusted to use a larger second index number to define the file name of the index file in the second primary shard when the first primary shard is migrated, so as to avoid the file name of the index file in the second primary shard from conflicting with the file name of the index file in the old first primary shard.
[0046] It is worth noting that, as shown in Figure 1 In a case where the cloud search service includes remote storage, the first primary shard uploads the index file to the remote storage for storage. When the first primary shard is migrated, the second primary shard generates a new index file and uploads the new index file to the remote storage. At this time, since the old first primary shard will continue to write data in the process of the second primary shard taking over the old first primary shard, if the first index number is directly used as the second index number corresponding to the second primary shard, there will be index files with the same file name in the remote storage. By using the second index number greater than the first index number, the index files with the same file name in the remote storage can be avoided, so that the index file of the second primary shard can be successfully uploaded to the remote storage for storage when the first primary shard is migrated, thereby supporting the separation of storage and computing architecture.
[0047] In step 230, in the process of migrating the first primary shard to the second primary shard, based on the second index number, the file name corresponding to the index file written to the second primary shard is determined, so that the file name corresponding to the index file written to the second primary shard is different from the file name of the index file of the first primary shard.
[0048] Here, after the second index number is determined, the file name of the index file written to the second primary shard can be determined by the second index number. Exemplarily, the second index number can be used as the initial number of the segment counter maintained by the second primary shard. Whenever a new index file is created in the second primary shard, the file name of the index file is determined based on the number recorded by the segment counter. Moreover, the count of the segment counter is correspondingly increased. Correspondingly, the file name of the index file written to the second primary shard subsequently is also determined by a larger second index number.
[0049] It is worth mentioning that since the second index number is determined by the engine for starting the second primary shard, for each second primary shard, its own second index number is determined by the engine corresponding to the second primary shard, so that the second index number is determined in a decentralized manner, and since the second index number is greater than the first index number, the file name of the index file written to the second primary shard will be different from the file name of the index file in the old first primary shard.
[0050] Therefore, by responding to the need for migration of the primary shard of the cloud search service, the first index number currently used by the first primary shard is determined, the target step is determined according to the migration mode of the first primary shard, and the second index number corresponding to the second primary shard is determined according to the target step and the first index number. Then, during the migration of the first primary shard to the second primary shard, the file name corresponding to the index file written to the second primary shard is determined based on the second index number, so that the file name corresponding to the index file written to the second primary shard is different from the file name of the index file of the first primary shard. Not only can it avoid conflicts between the file name of the index file written to the second primary shard and the file name of the index file in the first primary shard, but in the scenario of using remote storage, the index file created by the second primary shard needs to be uploaded to the remote storage. By using a larger second index number to determine the file name, it can be ensured that the file name in the remote storage will not conflict, thereby ensuring that the index file created by the second primary shard can be successfully uploaded to the remote storage.
[0051] In some implementable embodiments, in step 220, in the case where the migration mode represents that the replica shard corresponding to the first primary shard is upgraded to the second primary shard, the target step is determined according to the term of the second primary shard, and the size of the target step is positively correlated with the term of the second primary shard.
[0052] Here, the migration mode of the first primary shard can include the replica shard corresponding to the first primary shard being upgraded to the second primary shard. Accordingly, in the case where the replica shard corresponding to the first primary shard is upgraded to the second primary shard, the engine for starting the second primary shard can be used to determine the target step according to the term of the second primary shard, and then determine the second index number corresponding to the second primary shard according to the first index number and the target step.
[0053] The size of the target step is positively correlated with the term of the second primary shard. That is, in the case where the migration mode of the first primary shard is that the replica shard corresponding to the first primary shard is upgraded to the second primary shard, the target step can be determined by the term of the second primary shard, and the larger the term of the second primary shard, the larger the target step.
[0054] When the first primary shard fails, the cloud search service elects a second primary shard from the replica shards of the first primary shard through an election algorithm. Once a replica shard wins the election and becomes the second primary shard, the term of the second primary shard begins. During the term, the second primary shard undertakes key responsibilities, such as processing read and write requests from clients, coordinating write requests, ensuring data consistency, and synchronizing write operations and the like to corresponding replica shards to ensure that the data of the replica shards can be updated in a timely manner and consistent with the second primary shard.
[0055] In the case where the replica shard corresponding to the first primary shard is upgraded to the second primary shard, the term of the second primary shard upgraded from the replica shard is different from the term of the first primary shard. For example, assume that shard A is the first primary shard, the term of the first primary shard is 1, shard B is the replica shard of the first primary shard, when shard A hangs up, shard B is upgraded to the second primary shard, at this time, the term of shard B is changed to 2. If shard B hangs up after being upgraded to the second primary shard, shard A is switched to the second primary shard again, and the term of shard A is changed to 3.
[0056] In some embodiments, the sum of the first index number and the target step length can be determined as the second index number.
[0057] Exemplarily, the second index number can be obtained through a first calculation formula, the first calculation formula being:
[0058] wherein A1 is the second index number, B1 is the first index number, is a reference step length, C1 is the term, is a target step length.
[0059] It should be noted that the reference step length can be a constant. Exemplarily, the reference step length can be 100000. Of course, the reference step length can also be other values, which can be set according to actual conditions.
[0060] Assume that shard A is the first primary shard, the term of the first primary shard is 1, shard B is the replica shard of the first primary shard, the first index numbers corresponding to the first primary shard and the replica shard are both 100, when shard A hangs up, shard B is upgraded to the second primary shard, the term of shard B is changed to 2, based on the first calculation formula, the second index number corresponding to shard B is If shard B goes down after upgrading to the second primary shard, and shard A switches to the second primary shard again, the term of shard A will change to 3, and based on the first calculation formula, the second index number corresponding to shard A is .
[0061] It should be understood that the first calculation formula provided by the above embodiment is used as an example of calculating the second index number, and in other embodiments, other ways can be used to calculate the second index number, for example, the product between the first index number and the target step can be used as the second index number.
[0062] Therefore, in the case where the migration manner represents that the replica shard corresponding to the first primary shard is upgraded to the second primary shard, by dynamically adjusting the target step through the term of the second primary shard, the second index number corresponding to the second primary shard can be much larger than the first index number, so as to avoid the file name of the index file of the second primary shard from conflicting with the file name of the index file in the old first primary shard.
[0063] In some implementable embodiments, in step 220, in the case where the migration manner represents that the first primary shard is migrated to another node, the target step is determined according to the migrated number of times of the first primary shard to the another node, and the size of the target step is positively correlated with the migrated number of times of the first primary shard.
[0064] Here, the migration manner of the first primary shard can include that the primary shard is migrated to another node. The another node can be a node that does not deploy the replica shard corresponding to the first primary shard. Accordingly, in the case where the first primary shard is migrated to the another node, the target step can be determined according to the migrated number of times of the first primary shard to the another node through the engine used to start the second primary shard, and then the second index number corresponding to the second primary shard is determined according to the first index number, the number of log files corresponding to the transaction log included in the first primary shard, and the target step.
[0065] The size of the target step is positively correlated with the migrated number of times of the first primary shard. That is, in the case where the migration manner of the first primary shard represents that the first primary shard is migrated to another node, the target step can be determined through the migrated number of times of the first primary shard, and the greater the migrated number of times of the first primary shard, the greater the target step.
[0066] It should be noted that in the case where the first primary shard is migrated to another node, the second primary shard is formed by the migration of the old first primary shard, so the term of the second primary shard does not change and remains consistent with the term of the old first primary shard. Therefore, the target step can be dynamically adjusted through the migrated number of times of the first primary shard.
[0067] The migrated number of the first primary shard can refer to the number of times that the first primary shard is migrated to another node. It should be noted that in the process of migrating the first primary shard to another node, the migration may fail, at which time the process of migrating the first primary shard to another node is terminated, and then the process of migrating the first primary shard to another node is restarted until the first primary shard is successfully migrated to another node. Therefore, in this process, the number of times that the first primary shard is migrated to another node is the migrated number of the first primary shard.
[0068] Exemplarily, when the first primary shard is migrated for the first time, the migrated number of the first primary shard can be 0. When the first primary shard fails to be migrated to another node for the first time, the process of migrating the first primary shard to another node for the second time is started, and when the first primary shard is migrated to another node for the second time, the migrated number of the first primary shard can be 1. Similarly, as the migrated number of the first primary shard increases, the target step size also increases.
[0069] The transaction log (Translog) is an important part of the first primary shard, which is used to record all change operations on index data. When the first primary shard is migrated, the transaction log needs to be recovered during peer recovery (which refers to the process of recovering a replica shard from a primary shard). Therefore, when determining the second index number, the number of logs of the transaction log needs to be considered.
[0070] In some embodiments, the sum of the first index number, the number of logs, and the target step size can be determined as the second index number.
[0071] Exemplarily, the second index number can be obtained by a second calculation formula, which is:
[0072] wherein A2 is the second index number, B2 is the first index number, C2 is the number of logs, is a reference step size, is a migrated number, is a target step size.
[0073] It should be noted that the reference step size can be a constant. Exemplarily, the reference step size can be 100000. Of course, the reference step size can also be other values, which can be set according to actual conditions.
[0074] Of course, the second calculation formula provided in the above embodiment is an example of calculating the second index number, and in other embodiments, other ways can also be used to calculate the second index number, for example, the product of the first index number, the number of logs, and the target step size can be used as the second index number.
[0075] In some embodiments, during the migration of the first master shard to another node, if the file name of the index file written to the second master shard is the same as the file name of the index file in the first master shard, the process of migrating the first master shard to another node is terminated, the migration times of the first master shard is increased, and then the first master shard is migrated to another node again based on the increased migration times.
[0076] In some embodiments, during the migration of the first master shard to another node, if the file name of the index file written to the second master shard is the same as the file name of the index file in the first master shard, the process of migrating the first master shard to another node is terminated, the migration times of the first master shard is increased, and then the first master shard is migrated to another node again based on the increased migration times.
[0077] It should be understood that a counter can be maintained to record the migration times of the first master shard, and the initial value of the counter can be 0. When it is detected that the file name of the index file written to the second master shard is the same as the file name of the index file in the first master shard, the process of migrating the first master shard to another node is terminated, the count value of the counter is increased by 1, and then the process of migrating the first master shard to another node is restarted.
[0078] Since the migration times is increased by 1 when the first master shard is migrated to another node next time, the second index number used by the second master shard will be larger, which can significantly reduce the possibility of file name conflict when the first master shard is migrated to another node next time. Of course, if the file name of the index file written to the second master shard is the same as the file name of the index file in the first master shard, the process of migrating the first master shard to another node is terminated, the migration times of the first master shard is increased, and the next migration is performed until the migration is successful.
[0079] Therefore, in the case of migrating the first master shard to another node, the target step size is dynamically adjusted by the migration times of the first master shard, which can make the second index number corresponding to the second master shard much larger than the first index number, thereby avoiding the file name conflict of the index file of the second master shard and the index file of the first master shard.
[0080] In some implementable embodiments, in a case where the first primary shard performs merging of multiple segment files to obtain a target segment file, the target segment file is copied to a replica shard corresponding to the first primary shard through a pre-copy mechanism, and the target segment file is controlled to be visible to the first primary shard in a case where the target segment file has been copied to the replica shard.
[0081] Here, segment file replication is a new data replication strategy, which improves index throughput and improves resource utilization by copying segment files in the index to replica shards. When a new segment file is generated or old segment files are merged in the first primary shard, the new segment file needs to be synchronized to the replica shard, and when a node fails, a newly added replica shard needs to rebuild data by replicating segment files of the first primary shard. Compared with document replication synchronization, the synchronization mode of segment file replication can significantly reduce the memory and central processing unit overhead.
[0082] In a cloud search service, two types of segment files are included, one is a target segment file generated by merging segment files, and the other is a segment file constructed by Refresh (refreshing, which refers to refreshing a segment file in memory to the index) for incremental indexing. For the target segment file, assuming that a target segment file is 1 GB and the transmission bandwidth is 50 MB / s, the replication time of the target segment file needs at least 20 s, and for larger target segment files, the replication time is longer, resulting in a large visibility delay between the first primary shard and the replica shard.
[0083] In the embodiments of the present disclosure, in a case where the first primary shard performs merging of multiple segment files to obtain a target segment file, the target segment file is copied to a replica shard corresponding to the first primary shard through a pre-copy mechanism. And, in a case where the target segment file has been copied to the replica shard, the target segment file is controlled to be visible to the first primary shard, thereby reducing the visibility delay between the first primary shard and the replica shard. For the segment file generated by Refresh, the segment replication process is performed normally.
[0084] The pre-copy mechanism can be an IndexWriter.IndexReaderWarmer (an interface that allows pre-warming of newly merged segment files before they are committed to the index) pre-copy mechanism of Lucene. The target segment file being visible to the first primary shard means that the first primary shard can use the target segment file to process read and write requests.
[0085] It should be understood that the replica shard creates a listener on the first primary shard for detecting whether the first primary shard generates a new segment file, and when a target segment file is visible to the first primary shard, the replica shard senses that the first primary shard has a target segment file, at which time the replica shard needs to copy the target segment file. However, in the embodiment of the present disclosure, the target segment file has been pre-copied to the replica shard before the target segment file is visible to the first primary shard, and therefore when the replica shard senses that the first primary shard has the target segment file, the replica shard can directly load the target segment file without copying. Therefore, the visibility delay of the target segment file in the first primary shard and the replica shard can be reduced to the millisecond level.
[0086] Figure 3 FIG. 1 is a schematic diagram of pre-copying according to some embodiments. As shown in FIG. 1, the first primary shard processes a write stream through an engine. In the first primary shard, there are a segment file 1, a segment file 2, a segment file 3 generated through refresh, and a segment file 4 (target segment file) generated by merging the segment file 1 and the segment file 2. For the segment file 4, the segment file 4 is copied to the replica shard through a pre-copying mechanism, and the segment file 4 is visible to the first primary shard after the segment file 4 is completely copied to the replica shard. For the segment file 3, the segment file 3 is copied to the replica shard through a normal segment replication process. In the replica shard, the replica shard loads the segment file 3 and the segment file 4 into the engine. Figure 3
[0087] Since the segment file 4 is visible to the first primary shard after the segment file 4 is completely copied to the replica shard, at this time the replica shard senses that the first primary shard has the segment file 4, and since the segment file 4 already exists in the local of the replica shard, the replica shard does not need to re-copy the segment file 4, but directly controls the segment file 4 to be visible to the replica shard, thereby reducing the visibility delay of the segment file 4 between the first primary shard and the replica shard.
[0088] Thus, through the pre-copying mechanism, the visibility delay of the target segment file generated by merging between the first primary shard and the replica shard can be reduced.
[0089] In some implementable embodiments, when the first primary shard needs to be migrated, a second primary shard is started through a read-write engine, so that the second primary shard is in a readable and writable state during the migration of the first primary shard to the second primary shard.
[0090] Here, the read-write engine can be an engine supporting read-write function, and exemplarily, the read-write engine can be InternalEngine. In the case that the first primary shard needs to be migrated, the second primary shard is started by the read-write engine, and then the second primary shard can be in a read-write state during the migration of the first primary shard. In this way, when the first primary shard and the second primary shard perform handoff, the process of handoff is very short, and almost does not affect the real-time write of the user.
[0091] It should be noted that if the second primary shard is started by the read-only engine, the second primary shard will first create a read-only engine, which is only used to receive the segment file and the translog file sent by the first primary shard, and does not play back the translog file. Until the handoff stage of the first primary shard and the second primary shard, the first primary shard will initiate a new round of forced segment replication, and then reset the read-only engine to the read-write engine. At this time, the first primary shard will receive the new write operation. However, in the handoff stage, the user write data will be interrupted, and if the user writes a large amount of data during the migration process, the handoff time will be too long, causing the data write interruption time to be too long, and even causing the user write to fail.
[0092] Figure 4 FIG. 1 is a schematic diagram of starting the second primary shard by the read-write engine according to some embodiments. As shown in FIG. 1, the read-write engine of the first primary shard is responsible for processing the write stream, writing the data of the write stream into the memory and recording in the transaction log, and the read-write engine of the first primary shard will perform a flush operation at regular intervals to write the data in the memory to the disk to form a segment file (such as segment file 1, segment file 2). Figure 4
[0093] In the case that the first primary shard needs to be migrated, first, a segmented copy operation is performed to copy the segmented files (e.g., segmented file 1, segmented file 2) to the second primary shard, the second primary shard loads the segmented files obtained to the read-write engine of the second primary shard, and updates the segmented file information. Then, the first primary shard creates a transaction log snapshot and sends the transaction log snapshot to the second primary shard, the second primary shard receives the transaction log snapshot and replays the transaction log snapshot to apply the transaction log snapshot to the transaction log of the second primary shard. In this way, the second primary shard has the same uncommitted transaction records as the first primary shard. Finally, a handoff operation is performed, specifically, a completeRelocation operation is performed on the first primary shard, the first primary shard hands over the control to the new second primary shard, and the second primary shard also performs the completeRelocation operation after receiving the control to confirm that the second primary shard is ready to work independently. In the handoff stage, the first primary shard blocks new write operations to ensure that all uncommitted transactions are synchronized to the second primary shard.
[0094] That is, in the case that the first primary shard needs to be migrated, the second primary shard is started through the read-write engine, compared to starting the second primary shard through the read-only engine, there is no need to perform forced segmented copy, only a simple completeRelocation operation is needed, so that the handoff process can be extremely short and almost does not affect the real-time write of the user.
[0095] Therefore, through the above embodiments, it can be ensured that the first primary shard can be smoothly migrated without blocking the real-time write of the user.
[0096] In some implementable embodiments, in the case that the first primary shard replicates the index file to the replica shard corresponding to the first primary shard, the replica shard is started through the read-only engine to make the replica shard in a read-only state.
[0097] Here, when the first primary shard replicates the index file to the replica shard corresponding to the first primary shard, the replica shard is started through the read-only engine.
[0098] It should be understood that starting the replica shard through the read-only engine to make the replica shard in a read-only state can avoid the index operation (i.e., the replica shard participates in the document write process) of the replica shard.
[0099] Figure 5 is a schematic diagram of starting the replica shard through the read-only engine according to some embodiments. As Figure 5 shown, when the first primary shard replicates the segmented files to the replica shard, the replica shard is started through the read-only engine, and the segmented files (e.g., segmented file 1, segmented file 2, segmented file 3) are copied to the replica shard.
[0100] In some implementable embodiments, the first primary shard can be controlled to upload the index file written in the first primary shard to a remote storage, and the remote storage is used for the first primary shard or a replica shard corresponding to the first primary shard to read the required index file from the remote storage.
[0101] Here, as shown in Figure 1 the cloud search service includes a remote storage, the first primary shard can upload the index file to the remote storage, and maintain the meta information corresponding to the index file in the first primary shard, where the meta information is used to indicate the storage path of the index data in the remote storage.
[0102] It should be noted that, since the first primary shard uploads the index file to the remote storage for storage, the computing node where the first primary shard is located does not need to mount a local disk for data storage, thereby reducing the cost of local storage. For the computing node, a small amount of meta information is maintained, and when receiving a query request, the first primary shard or the replica shard determines the storage path of the data to be queried by the query request in the remote storage through the meta information, and the first primary shard or the replica shard reads the data to be queried by the query request from the remote storage through the storage path. For the storage node, the data can be stored through the remote storage, and the part of data calculation is responsible by the computing node, therefore, the cloud search service can realize the separation of computing and storage.
[0103] The first primary shard can also synchronize the meta information to the replica shard corresponding to the primary shard through a segment replication process. The client can send a query request to any node in the cloud search service, and the node receiving the query request acts as a coordination node, and randomly forwards the query request to the corresponding first primary shard or replica shard. When the first primary shard or the replica shard receives the query request, the first primary shard or the replica shard accesses the index file in the remote storage according to the meta information maintained by the first primary shard or the replica shard, and determines the field value corresponding to the query request from the index file.
[0104] Wherein, the field is a basic data unit in the document, used to store and index specific document data. Each field has a name and one or more values. The field is used to define the structure and content of the document, facilitating search, indexing and query. The field value refers to the specific data value stored in each field of the document. For example, in the user information document, the value of the "name" field can be a specific name, the value of the "age" field can be an age value, etc.
[0105] In the embodiments of the present disclosure, the first primary shard can write the index file to the remote storage, and the first primary shard and the replica shard both support direct access to the remote storage to read the field value corresponding to the query request, thereby greatly reducing the cost of local storage.
[0106] It is worth mentioning that since the second index number is determined by the engine for starting the second primary shard, for each second primary shard, its own second index number is determined by the engine corresponding to the second primary shard, so that the second index number is determined in a decentralized manner, and since the second index number is greater than the first index number, the file name of the index file written by the second primary shard to the remote storage will be different from the file name of the index file written by the old first primary shard to the remote storage.
[0107] Since the file name of the index data written by the second primary shard to the remote storage will be different from the file name of the index file written by the old first primary shard to the remote storage, the first primary shard and the replica shard corresponding to the first primary shard can directly query the remote storage index file to obtain the field value.
[0108] Figure 6 Figure 1 is a structural schematic diagram of a cloud search service-based data synchronization device according to some embodiments. As shown in Figure 6 The cloud search service-based data synchronization device 500 provided by the embodiment of the present disclosure comprises: A first determination module 501 configured to determine a first index number currently used by a first primary shard of a cloud search service in response to the first primary shard needing to be migrated, the first index number being used by the first primary shard to determine a file name corresponding to an index file written by the first primary shard; A second determination module 502 configured to determine a target step according to a migration mode of the first primary shard, and determine a second index number corresponding to a second primary shard replacing the first primary shard according to the target step and the first index number, the target step being used to make the second index number greater than the first index number; A third determination module 503 configured to determine a file name corresponding to an index file written by the second primary shard based on the second index number during migration of the first primary shard to the second primary shard, so that the file name corresponding to the index file written by the second primary shard is different from the file name of the index file of the first primary shard.
[0109] Optionally, the second determination module 502 is specifically configured to: In a case where the migration mode represents that a replica shard corresponding to the first primary shard is upgraded to the second primary shard, determine the target step according to a term of the second primary shard, the size of the target step being positively correlated with the term of the second primary shard.
[0110] Optionally, the second determination module 502 is specifically configured to: In the migration mode, the first master shard is migrated to another node, and the target step is determined according to the number of migrations of the first master shard to the another node, and the size of the target step is positively correlated with the number of migrations of the first master shard.
[0111] Optionally, the cloud search service-based data synchronization apparatus 500 further comprises: The termination module is configured to, in the process of migration of the first master shard to the another node, if the file name of the index file written in the second master shard is the same as the file name of the index file in the first master shard, terminate the process of migration of the first master shard to the another node, and increase the number of migrations of the first master shard; The re-migration module is configured to, based on the increased number of migrations, re-migrate the first master shard to the another node.
[0112] Optionally, the cloud search service-based data synchronization apparatus 500 further comprises: The copy module is configured to, in the case that the first master shard performs merging of a plurality of segment files to obtain a target segment file, copy the target segment file to the replica shard corresponding to the first master shard through a pre-copy mechanism, and control the target segment file to be visible to the first master shard in the case that the target segment file has been copied to the replica shard.
[0113] Optionally, the cloud search service-based data synchronization apparatus 500 further comprises: The first start module is configured to, in the case that the first master shard needs to be migrated, start the second master shard through a read-write engine, so that the second master shard is in a readable and writable state in the process of migration of the first master shard to the second master shard. The second start module is configured to, in the case that the first master shard replicates an index file to the replica shard corresponding to the first master shard, start the replica shard through a read-only engine, so that the replica shard is in a read-only state.
[0114] Optionally, the cloud search service-based data synchronization apparatus 500 further comprises: The upload module is configured to control the first master shard to upload the index file written in the first master shard to a remote storage, and the remote storage is used for the first master shard or the replica shard corresponding to the first master shard to read the required index file from the remote storage.
[0115] The function logic performed by each functional module in the above-mentioned data synchronization apparatus 500 based on the cloud search service has been described in detail in the section about the method, and will not be repeated here.
[0116] Reference is made below Figure 7 which shows a structural schematic diagram of an electronic device (e.g., a server) 600 suitable for use to implement embodiments of the present disclosure. Figure 7 The electronic device shown is merely an example and should not bring any limitation to the function and use range of embodiments of the present disclosure.
[0117] As Figure 7 shown, the electronic device 600 can include a processing device (e.g., a central processor, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0118] Generally, the following devices can be connected to the I / O interface 605: input devices 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 608 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 609. The communication devices 609 can allow the electronic device 600 to communicate with other devices wirelessly or through wires to exchange data. Although Figure 7 The electronic device 600 is shown with various devices, but it should be understood that all the shown devices are not required to be implemented or possessed. More or fewer devices can be alternatively implemented or possessed.
[0119] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 609, or installed from the storage devices 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the methods of embodiments of the present disclosure are performed.
[0120] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus or device. In the disclosure, the computer-readable signal medium can include a data signal that propagates in a baseband or as part of a carrier wave, carrying computer-readable program code. Such a propagated data signal can take many forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that can send, propagate, or transfer program code for use by or in connection with an instruction execution system, apparatus or device. Program code contained in a computer-readable medium can be transmitted using any suitable medium, including but not limited to wire, cable, optical fiber, RF, etc., or any suitable combination of the foregoing.
[0121] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0122] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and not be assembled into the electronic device.
[0123] The computer-readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: In response to a first primary shard of a cloud search service needing to be migrated, a first index number currently used by the first primary shard is determined, the first index number being used by the first primary shard to determine a file name corresponding to an index file written to the first primary shard; According to a migration manner of the first primary shard, a target step is determined, and according to the target step and the first index number, a second index number corresponding to a second primary shard replacing the first primary shard is determined, the target step being used to make the second index number greater than the first index number; In a process in which the first primary shard is migrated to the second primary shard, based on the second index number, a file name corresponding to an index file written to the second primary shard is determined, so that the file name corresponding to the index file written to the second primary shard is different from a file name of an index file of the first primary shard.
[0124] Computer program code for carrying out operations of the present disclosure can be written in any of one or more programming languages, including object oriented programming languages such as Java, Smalltalk, C++, or conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0125] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or in the reverse order, depending on the functionality involved. It is also noted that each block of the block diagrams and / or flow diagrams and combinations of blocks in the block diagrams and / or flow diagrams can be implemented by special purpose hardware-based systems which perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0126] The modules involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a module does not necessarily limit the module itself.
[0127] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0128] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0129] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the present disclosure is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the scope of the above disclosure. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0130] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0131] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims. With respect to the devices in the above-described embodiments, in which various modules perform operations, the specific manner in which the operations are performed by the various modules has been described in detail in the embodiments relating to the method. No further elaboration will be made here.
Claims
1. A data synchronization method based on cloud search service, characterized in that: include: In response to a need to migrate a first primary shard of a cloud search service, determining a first index number currently used by the first primary shard, where the first index number is used by the first primary shard to determine a file name corresponding to an index file written to the first primary shard; Determine a target step size according to the migration mode of the first primary shard, and determine a second index number corresponding to a second primary shard that replaces the first primary shard based on the target step size and the first index number, where the target step size is set so that the second index number is greater than the first index number; During the process of migrating the first primary shard to the second primary shard, the file name corresponding to the index file written to the second primary shard is determined based on the second index number, so that the file name corresponding to the index file written to the second primary shard is different from the file name of the index file of the first primary shard.
2. The method according to claim 1, characterized in that The determining of the target step size according to the migration mode of the first primary shard includes: When the migration method indicates that the replica shard corresponding to the first primary shard is upgraded to the second primary shard, the target step size is determined according to the term of the second primary shard, and the size of the target step size is positively correlated with the term of the second primary shard.
3. The method according to claim 1, characterized in that The determining of the target step size according to the migration mode of the first primary shard includes: The migration method represents the migration of the first primary shard to another node. The target step size is determined based on the number of times the first primary shard has been migrated to the other node. The size of the target step size is positively correlated with the number of times the first primary shard has been migrated.
4. The method according to claim 3, characterized in that The method further comprises: During the migration of the first primary shard to the other node, if the file name of the index file written to the second primary shard is the same as the file name of the index file in the first primary shard, the migration of the first primary shard to the other node is terminated, and the migration count of the first primary shard is increased; Based on the increased number of migrations, the first primary shard is re-migrated to the other node.
5. The method according to any one of claims 1 to 4, characterized in that The index file is a segmented file, and the method further includes: When the first primary shard executes a merge of multiple segment files to obtain a target segment file, the target segment file is copied to the replica shard corresponding to the first primary shard through a pre-copy mechanism, and when the target segment file has been copied to the replica shard, the target segment file is controlled to be visible to the first primary shard.
6. The method according to any one of claims 1 to 4, characterized in that The method further comprises: When the first primary shard needs to be migrated, the second primary shard is started by the read / write engine, so that the second primary shard is in a readable and writable state during the migration process from the first primary shard to the second primary shard; In a case where the first primary shard copies the index file to the replica shard corresponding to the first primary shard, the replica shard is started by a read-only engine to put the replica shard in a read-only state.
7. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Control the first primary shard to upload the index file written in the first primary shard to a remote storage, where the remote storage is used for the first primary shard or the replica shard corresponding to the first primary shard to read the required index file from the remote storage.
8. A data synchronization device based on cloud search service, characterized in that: include: a first determining module configured to, in response to a first primary shard of a cloud search service needing to be migrated, determine a first index number currently used by the first primary shard, the first index number being used by the first primary shard to determine a file name corresponding to an index file written to the first primary shard; A second determining module is configured to determine a target step size according to the migration mode of the first primary shard, and determine a second index number corresponding to a second primary shard that is to replace the first primary shard based on the target step size and the first index number, wherein the target step size is used to make the second index number greater than the first index number; The third determination module is configured to determine the file name corresponding to the index file written to the second primary shard based on the second index number during the process of migrating the first primary shard to the second primary shard, so that the file name corresponding to the index file written to the second primary shard is different from the file name of the index file of the first primary shard.
9. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processing device, the steps of the method according to any one of claims 1 to 7 are implemented.
10. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 7.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Data writing method and device, data migration method and device and electronic equipment
CN113986878A
Index fragment merging method and device
CN114579562A
Data synchronization method, device and equipment and readable storage medium
CN117216160A
Chinese assignment divergence retrieval system based on Elasticsearch
CN119474256A
Data query method based on cloud search service, medium, equipment and product
CN121116971A