Multi-cluster index updating method, device, equipment, medium and product
By smoothly migrating business traffic to other clusters in a multi-cluster environment and performing phased index updates, the problem of interface instability caused by simultaneous updates across multiple clusters was solved, and an efficient and stable data retrieval service was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD
- Filing Date
- 2026-03-06
- Publication Date
- 2026-07-21
AI Technical Summary
Updating the index simultaneously across multiple clusters can easily lead to unstable cluster interface responses, especially in high-concurrency scenarios, affecting the performance and stability of the retrieval system.
By smoothly migrating the business traffic of the target retrieval cluster to the remaining retrieval clusters, and updating the data index of the target retrieval cluster after the traffic migration is completed, a multi-stage traffic migration strategy and a data volume floating algorithm are adopted to ensure interface response efficiency. Grouped datasets are imported in parallel and index aliases are switched.
This effectively avoids the instability of interface responses caused by simultaneous index updates across the entire cluster, improves the overall performance and stability of the retrieval system, and enhances the user experience.
Smart Images

Figure CN122432388A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data processing technology, and in particular to a method, apparatus, device, medium and product for updating multi-cluster indexes. Background Technology
[0002] Retrieval systems are designed to meet the query needs of big data and can be applied to scenarios such as e-commerce product search, news search, and user tag search. With the continuous growth of data volume and the increasing complexity of business scenarios, single-cluster retrieval systems are gradually becoming insufficient to meet the demands for efficient and stable data retrieval. Multi-cluster architectures have emerged to address this need, effectively improving the overall performance of the system by storing and processing data across multiple clusters.
[0003] Indexes are the core data structure of a retrieval system. When big data is updated, indexes also need to be updated accordingly to ensure the accuracy of retrieval results. In multi-cluster environments, multiple clusters typically update the index simultaneously. However, this approach can easily lead to unstable cluster interface responses, especially in high-concurrency scenarios. Some clusters may be unable to respond in a timely manner due to excessive pressure, thus affecting the performance and stability of the entire retrieval system and consequently impacting the user experience. Summary of the Invention
[0004] This application provides a method, apparatus, device, medium, and product for updating indexes across multiple clusters, in order to solve the problem that the existing technology of updating indexes across multiple clusters simultaneously can easily lead to unstable cluster interface responses, thereby affecting query performance.
[0005] Firstly, this application provides a multi-cluster index update method, including: For each retrieval cluster in the retrieval cluster set, the retrieval cluster is determined as the target retrieval cluster for the current index update, and the business traffic of the target retrieval cluster is smoothly migrated to the remaining retrieval clusters, wherein the remaining retrieval clusters are the retrieval clusters in the retrieval cluster set other than the target retrieval cluster; after the business traffic diversion of the target retrieval cluster is completed, the data index of the target retrieval cluster is updated. Once each retrieval cluster has completed its data index update, corresponding business traffic is allocated to each retrieval cluster.
[0006] In one embodiment, smoothly migrating the service traffic of the target retrieval cluster to the remaining retrieval clusters includes: Based on the performance metrics of the remaining retrieval clusters, the traffic flow ratio of the remaining retrieval clusters is determined; the traffic flow ratio is used to indicate the proportion of business traffic migrated to the remaining retrieval clusters. Based on a multi-stage traffic migration strategy, the business traffic of the target retrieval cluster is migrated to the remaining retrieval clusters in stages according to the traffic flow ratio.
[0007] In one embodiment, the step of migrating the service traffic of the target retrieval cluster to the remaining retrieval clusters in stages according to the traffic flow ratio based on a multi-stage traffic migration strategy includes: At the current stage, based on a multi-stage traffic migration strategy, the business traffic of the target retrieval cluster is migrated to the other retrieval clusters according to the traffic flow ratio; Determine the interface response efficiency of the remaining retrieval clusters; If the interface response efficiency is less than a preset response efficiency threshold, the multi-stage traffic migration strategy is adjusted to obtain an adjusted multi-stage traffic migration strategy. The adjusted multi-stage traffic migration strategy is determined as the multi-stage traffic migration strategy, and the next stage is determined as the current stage. The steps of migrating the business traffic of the target retrieval cluster to the other retrieval clusters according to the traffic flow ratio are iteratively executed in the current stage based on the multi-stage traffic migration strategy until the last stage is executed. If the interface response efficiency is greater than or equal to the preset response efficiency threshold, the multi-stage traffic migration strategy is iteratively executed. In the current stage, the business traffic of the target retrieval cluster is migrated to the other retrieval clusters according to the traffic flow ratio, until the last stage is executed.
[0008] In one embodiment, updating the data index of the target retrieval cluster includes: Import the dataset to be imported into the target retrieval cluster; Obtain the total current imported data volume and the total historical imported data volume of the target retrieval cluster; Based on the data volume fluctuation algorithm and the total amount of historical imported data, the normal range of the total amount of imported data is determined; If the current total amount of imported data does not meet the normal range of the total amount of imported data, then return to the step of importing the dataset to be imported into the target retrieval cluster; If the total amount of imported data meets the normal range of the total amount of imported data, then the first index alias will be switched from the index of the original dataset to the index of the newly imported dataset.
[0009] In one embodiment, before updating the data index for each of the retrieval clusters, the method further includes: The dataset to be imported is divided according to its data type, resulting in multiple grouped datasets; For each grouped dataset, a group index is constructed in each of the retrieval clusters based on the group identifier and the current date information.
[0010] In one embodiment, updating the data index of the target retrieval cluster includes: Based on each of the grouping indexes, the grouping datasets are imported into the target retrieval cluster in parallel; For each grouped dataset, switch the second index alias from the group index of the original grouped dataset to the group index of the newly imported grouped dataset.
[0011] Secondly, this application also provides a multi-cluster index update apparatus, comprising: The multi-cluster index update module is used to determine each retrieval cluster in the retrieval cluster set as the target retrieval cluster for the current index update, and smoothly migrate the business traffic of the target retrieval cluster to the remaining retrieval clusters, wherein the remaining retrieval clusters are the retrieval clusters in the retrieval cluster set other than the target retrieval cluster; after the business traffic diversion of the target retrieval cluster is completed, the data index of the target retrieval cluster is updated. The traffic configuration module is used to allocate corresponding business traffic to each of the search clusters after each search cluster has completed the data index update.
[0012] Thirdly, this application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of any of the above-described multi-cluster index update methods.
[0013] Fourthly, this application also provides a non-transitory computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the above-described multi-cluster index update methods.
[0014] Fifthly, this application also provides a computer program product, the computer program product including a computer program that can be stored on a non-transitory computer-readable storage medium, and the computer program, when executed by the processor, implements the steps of any of the above-described multi-cluster index update methods.
[0015] The multi-cluster index update method, apparatus, device, medium, and product provided in this application smoothly migrate business traffic to other clusters before updating a single cluster, and perform independent updates for each cluster after the full traffic migration is completed. This avoids the instability of interface response caused by updating the index of the entire cluster at the same time. Traffic allocation is then restored uniformly after all clusters have been updated. This keeps the business impact generated during the index update iteration process at an extremely low level, greatly improves the overall performance and stability of the retrieval system, and thus enhances the user experience. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is one of the flowcharts illustrating the multi-cluster index update method provided in this application.
[0018] Figure 2 This is a schematic diagram of traffic allocation under normal business conditions provided in this application.
[0019] Figure 3 This is one of the business traffic allocation diagrams provided in this application under the condition of index update.
[0020] Figure 4 This is the second schematic diagram of business traffic allocation under the index update scenario provided in this application.
[0021] Figure 5 This is the second flowchart of the multi-cluster index update method provided in this application.
[0022] Figure 6 This is a schematic diagram of the structure of the multi-cluster index update device provided in this application.
[0023] Figure 7 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0025] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein.
[0026] This application presents the following embodiments, and in conjunction with Figures 1-5 Describe the multi-cluster index update method provided in this application.
[0027] The multi-cluster index update method provided in this application embodiment can be implemented based on a multi-cluster index update device. Therefore, this application embodiment uses a multi-cluster index update device as the execution subject to describe the multi-cluster index update method.
[0028] Figure 1 This is one of the flowcharts illustrating the multi-cluster index update method provided in this application.
[0029] like Figure 1 As shown, the multi-cluster index update method includes the following steps: Step 101: For each retrieval cluster in the retrieval cluster set, determine the retrieval cluster as the target retrieval cluster for the current index update, and smoothly migrate the business traffic of the target retrieval cluster to the remaining retrieval clusters, wherein the remaining retrieval clusters are the retrieval clusters in the retrieval cluster set other than the target retrieval cluster; after the business traffic diversion of the target retrieval cluster is completed, update the data index of the target retrieval cluster.
[0030] Specifically, the retrieval system adopts a multi-cluster architecture, which includes multiple retrieval clusters that collaborate to complete big data retrieval tasks. Each retrieval cluster possesses certain data processing and indexing capabilities, and the efficient and stable operation of the retrieval system can be ensured through reasonable allocation of business traffic. The retrieval system includes, but is not limited to, open-source retrieval engines such as Elasticsearch and Apache Solr, which are highly efficient and scalable when handling big data retrieval tasks. This application's embodiment uses Elasticsearch as an example; in the multi-cluster architecture, the Elasticsearch system includes multiple Elasticsearch clusters, referred to as es clusters.
[0031] Figure 2 This is a diagram illustrating the traffic allocation under normal business conditions provided in this application. For example... Figure 2As shown, assuming the Elasticsearch system includes two Elasticsearch clusters, ES cluster 1 and ES cluster 2, under normal business conditions, the traffic control tool evenly distributes client traffic across ES cluster 1 and ES cluster 2. Specifically, it allocates 50% of client traffic to ES cluster 1 and 50% to ES cluster 2, achieving balanced traffic distribution between the clusters. Of course, traffic can also be dynamically distributed between ES cluster 1 and ES cluster 2. For example, based on the real-time load of ES cluster 1 and ES cluster 2, more traffic can be allocated to the higher-performing cluster (e.g., the lower-loaded cluster), while less traffic can be allocated to the lower-performing cluster (e.g., the higher-loaded cluster). This ensures that both clusters operate efficiently, preventing one cluster from being overloaded and affecting overall retrieval performance. It also fully utilizes the resources of each cluster, improving the overall processing capacity and stability of the Elasticsearch system.
[0032] Big data is stored on the Elasticsearch system, and a unified alias is used during queries. Big data is dynamically updated, with the frequency and method of updates depending on application and business requirements. Elasticsearch is a dynamic, continuously running engine that needs to periodically import new data and update its internal indexes accordingly to provide accurate, reliable, and efficient services for retrieval and business analysis. Because this application enables business traffic switching, each Elasticsearch cluster needs to import the same data to ensure data consistency across clusters, thereby guaranteeing the accuracy and consistency of search results.
[0033] To avoid instability caused by multiple clusters updating the index simultaneously, an isolated index update approach is adopted, updating the index one cluster at a time. Each Elasticsearch cluster currently performing an index update process is referred to as the target retrieval cluster, and the Elasticsearch clusters other than the target retrieval cluster are referred to as the remaining retrieval clusters, which include one or more.
[0034] Before index updates, traffic control tools should be used to smoothly migrate the business traffic of the target retrieval cluster to the remaining retrieval clusters. The traffic migration process must be smooth to avoid sudden increases or decreases in traffic that could impact the business of other retrieval clusters and thus affect the stability of the entire retrieval system.
[0035] After the business traffic of the target retrieval cluster is redirected, the Elasticsearch cluster no longer carries retrieval business requests. At this time, the data index update operation can be performed on the target retrieval cluster. First, the dataset to be imported is imported into the target retrieval cluster. After the import is completed, the index structure is switched to the newly imported dataset.
[0036] The above index update process is executed on each Elasticsearch cluster to achieve seamless data index updates across multiple clusters, ensuring that the retrieval system can continuously provide stable and accurate retrieval services during the update process.
[0037] Combination Figure 3 and Figure 4 , Figure 3 This is one of the business traffic allocation diagrams provided in this application during index update. Figure 4 This is the second diagram illustrating traffic allocation during index updates provided in this application. Assume the Elasticsearch system includes Elasticsearch cluster 1 and Elasticsearch cluster 2. When an index update is needed for Elasticsearch cluster 1, the traffic control tool smoothly migrates traffic from Elasticsearch cluster 1 to Elasticsearch cluster 2 until all traffic from Elasticsearch cluster 1 has been migrated to Elasticsearch cluster 2. Then, Elasticsearch cluster 1 undergoes a data index update. Similarly, when an index update is needed for Elasticsearch cluster 2, the traffic control tool smoothly migrates traffic from Elasticsearch cluster 2 to Elasticsearch cluster 1 until all traffic from Elasticsearch cluster 2 has been migrated to Elasticsearch cluster 1. Then, Elasticsearch cluster 2 undergoes a data index update. Once both Elasticsearch clusters 1 and 2 have completed their data index updates, the corresponding traffic is allocated to each Elasticsearch cluster.
[0038] Step 102: After each retrieval cluster has completed the data index update, allocate corresponding business traffic to each retrieval cluster.
[0039] Specifically, after all Elasticsearch clusters have completed data index updates, it is necessary to reallocate client traffic across these clusters to restore the entire retrieval system to normal traffic distribution and continue providing services for retrieval. At this point, traffic control tools can either distribute traffic evenly across multiple clusters or dynamically across them.
[0040] In a balanced distribution scenario, traffic control tools need to distribute business traffic evenly across each Elasticsearch cluster in the same proportion. This is typically a strategy used when multiple clusters have comparable performance and business demands are relatively stable.
[0041] In a dynamic allocation scenario, the traffic control tool needs to dynamically adjust the traffic allocation ratio based on the real-time performance metrics of each Elasticsearch cluster (such as CPU utilization, memory usage, load, and index query response time), and then distribute business traffic to each Elasticsearch cluster according to the traffic allocation ratio. For example, a higher proportion of business traffic can be allocated to clusters with better performance, while a lower proportion can be allocated to clusters with slightly weaker performance, thereby achieving load balancing among clusters. This dynamic allocation strategy can flexibly adjust traffic allocation according to the actual operating conditions of the clusters.
[0042] The multi-cluster index update method provided in this application smoothly migrates business traffic to other clusters before updating a single cluster, and performs independent updates for each cluster after the full traffic migration is completed. This avoids the instability of interface response caused by updating the index across all clusters at the same time. Traffic allocation is then restored uniformly after all clusters have been updated. This keeps the business impact generated during the index update iteration process at a very low level, greatly improving the overall performance and stability of the retrieval system, and thus enhancing the user experience.
[0043] In one embodiment, the step of smoothly migrating the service traffic of the target retrieval cluster to the remaining retrieval clusters includes: Based on the performance metrics of the remaining retrieval clusters, the traffic flow ratio of the remaining retrieval clusters is determined; the traffic flow ratio is used to indicate the proportion of business traffic migrated to the remaining retrieval clusters. Based on a multi-stage traffic migration strategy, the business traffic of the target retrieval cluster is migrated to the remaining retrieval clusters in stages according to the traffic flow ratio.
[0044] Specifically, when there are multiple other retrieval clusters, a comprehensive evaluation is conducted based on the real-time performance metrics of these clusters, such as CPU utilization, memory usage, load, and index query response time. These performance metrics comprehensively reflect the current processing capacity and load status of the clusters, providing a basis for determining a reasonable traffic flow ratio.
[0045] Then, based on the comprehensive evaluation results, the traffic flow ratio of these remaining retrieval clusters is determined. This ratio is used to indicate the proportion of business traffic migrated to each of the remaining retrieval clusters. This traffic ratio is determined based on the relative performance of the remaining retrieval clusters. Clusters with better performance will receive a higher proportion of migrated traffic to handle more business traffic.
[0046] If there is only one other retrieval cluster, the traffic migrated to that cluster will account for 100%, and all business traffic needs to be migrated to that cluster.
[0047] After determining the traffic flow ratio for the remaining search clusters, a multi-stage traffic migration strategy is adopted to migrate the business traffic of the target search cluster to the remaining search clusters in stages according to the determined traffic flow ratio. This staged migration method can ensure a smooth traffic migration process and avoid significant business impact (such as timeout rates) on the remaining search clusters. Simultaneously, by monitoring performance metric changes during the migration process in real time, the traffic flow ratio for the remaining search clusters can be adjusted promptly to ensure a smooth migration process.
[0048] For example, a multi-stage traffic migration strategy can adopt a phased, gradual migration approach, progressively increasing the proportion of traffic migrated to other retrieval clusters. If five stages are set, the traffic migration proportions for these five stages can be set using the same increment, such as 20%, 40%, 60%, 80%, and 100%, or not using the same increment, such as 10%, 30%, 50%, 80%, and 100%. The specific traffic migration proportions can be flexibly adjusted according to the actual situation.
[0049] Traffic migration control can be implemented using user identifiers, unique query identifiers, etc. Specifically, during traffic migration, each query request carries a user identifier or a unique query identifier. The retrieval cluster uses these identifiers to identify the source of the request and the data ownership. For query requests with specific user identifiers or unique data identifiers, they are guided to the corresponding retrieval cluster according to pre-defined traffic migration rules. This allows for precise control over the traffic received by each retrieval cluster, ensuring the accuracy and controllability of traffic migration, avoiding traffic chaos, and further improving the stability and reliability of traffic migration during multi-cluster index updates.
[0050] This application embodiment dynamically allocates business traffic based on the actual performance of the retrieval cluster. In a multi-cluster environment, it can effectively avoid the overall performance degradation caused by the overload of a single cluster, while preventing the waste of resources in low-load clusters. Then, through a phased migration strategy, it achieves gradual traffic redirection, which not only ensures the controllability of the migration process, but also maintains the continuity of system services and meets personalized business needs.
[0051] In one embodiment, the step of migrating the service traffic of the target retrieval cluster to the remaining retrieval clusters in stages according to the traffic flow ratio based on a multi-stage traffic migration strategy includes: At the current stage, based on a multi-stage traffic migration strategy, the business traffic of the target retrieval cluster is migrated to the other retrieval clusters according to the traffic flow ratio; Determine the interface response efficiency of the remaining retrieval clusters; If the interface response efficiency is less than a preset response efficiency threshold, the multi-stage traffic migration strategy is adjusted to obtain an adjusted multi-stage traffic migration strategy. The adjusted multi-stage traffic migration strategy is determined as the multi-stage traffic migration strategy, and the next stage is determined as the current stage. The steps of migrating the business traffic of the target retrieval cluster to the other retrieval clusters according to the traffic flow ratio are iteratively executed in the current stage based on the multi-stage traffic migration strategy until the last stage is executed. If the interface response efficiency is greater than or equal to the preset response efficiency threshold, the multi-stage traffic migration strategy is iteratively executed. In the current stage, the business traffic of the target retrieval cluster is migrated to the other retrieval clusters according to the traffic flow ratio, until the last stage is executed.
[0052] Specifically, during the execution of a multi-stage traffic migration strategy for business traffic migration, key metrics such as interface response time and throughput of the remaining retrieval clusters need to be monitored upon completion of each stage. The interface response efficiency is then calculated based on these metrics. The interface response efficiency of the remaining retrieval clusters is compared with a preset response efficiency threshold to determine whether dynamic adjustments to the multi-stage traffic migration strategy are necessary. The preset response efficiency threshold can be flexibly set according to actual business needs and system performance to ensure system stability and efficiency during traffic migration.
[0053] When the interface response efficiency is less than the preset response efficiency threshold, the multi-stage traffic migration strategy is adjusted. Adjustments may include, but are not limited to, reducing the proportion of traffic migration in subsequent stages or increasing the number of migration stages, to alleviate the load on other retrieval clusters and improve their interface response efficiency. Subsequently, based on the adjusted multi-stage traffic migration strategy, traffic migration for the target retrieval cluster is performed in subsequent stages.
[0054] During the iterative execution of traffic migration steps, the interface response efficiency of the remaining retrieval clusters is continuously monitored, and the multi-stage traffic migration strategy is dynamically adjusted according to the actual situation until all stages of traffic migration are successfully completed.
[0055] This application embodiment employs a multi-stage traffic migration strategy and dynamically adjusts it based on the interface response efficiency of other retrieval clusters. This ensures that the interface response of each retrieval cluster remains within a reasonable range during the smooth traffic migration process, avoiding system performance degradation due to sudden traffic surges. This greatly reduces the impact of traffic migration on business operations and ensures the normal operation of the system under high load conditions.
[0056] In one embodiment, updating the data index of the target retrieval cluster includes: Import the dataset to be imported into the target retrieval cluster; Obtain the total current imported data volume and the total historical imported data volume of the target retrieval cluster; Based on the data volume fluctuation algorithm and the total amount of historical imported data, the normal range of the total amount of imported data is determined; If the current total amount of imported data does not meet the normal range of the total amount of imported data, then return to the step of importing the dataset to be imported into the target retrieval cluster; If the total amount of imported data meets the normal range of the total amount of imported data, then the first index alias will be switched from the index of the original dataset to the index of the newly imported dataset.
[0057] Specifically, an index for the dataset to be imported is built in the target retrieval cluster. After completing the full migration of business traffic in the target retrieval cluster, the dataset to be imported is imported into the storage space of its corresponding index in the target retrieval cluster, and the total amount of data imported is recorded after the import is complete.
[0058] At the same time, obtain the total amount of historical imported data. The total amount of historical imported data can be the total amount of data imported last time, or the average of the total amounts of data imported multiple times in the past. There is no limitation here.
[0059] Based on the data volume fluctuation algorithm and the total amount of historical imported data, a normal range for the total amount of imported data is determined. The data volume fluctuation algorithm can be set according to actual business needs and data characteristics. A percentage fluctuation range based on the total amount of historical imported data can be set as the normal range, for example, setting ±10% of the total amount of historical imported data as the normal fluctuation range. This setting takes into account both the natural growth or decrease trend of data volume and can tolerate a certain range of temporary fluctuations.
[0060] If the total amount of data being imported does not meet the normal range, it indicates that the dataset being imported is invalid, triggering an alarm mechanism and prompting a re-import operation. Before re-importing, the import process can be checked and optimized, such as checking whether the network connection is normal and whether the data format meets the requirements, to improve the success rate of subsequent imports.
[0061] If the total amount of data being imported meets the normal range, it indicates that the imported dataset is normal. The first index alias will then be switched from the original dataset's index to the index of the newly imported dataset (i.e., the index of the dataset to be imported). Switching the first index alias allows for rapid index updates without directly impacting ongoing queries. After switching the first index alias, the retrieval system will automatically redirect subsequent query requests to the newly imported dataset's index, ensuring data real-time performance and accuracy.
[0062] This application embodiment calculates the normal range of the total imported data volume using a data volume fluctuation algorithm and the total historical imported data volume. Based on the normal range of the total imported data volume, it identifies abnormal data import situations. When abnormal data import situations are detected, it automatically alarms and re-imports the data without manual intervention, ensuring the quality of the imported data and significantly improving the response speed of data import anomaly repair. Under normal imported data conditions, it achieves seamless index updates through index alias switching, ensuring the stability of the retrieval service.
[0063] In one embodiment, before updating the data index of each of the retrieval clusters, the method further includes: The dataset to be imported is divided according to its data type, resulting in multiple grouped datasets; For each grouped dataset, a group index is constructed in each of the retrieval clusters based on the group identifier and the current date information.
[0064] Specifically, the datasets to be imported are grouped according to data type, such as user identifier datasets, data identifier datasets, and business datasets. Based on data type grouping, further subdivision can be performed based on similar data attributes. For example, the user identifier dataset can be processed using a hash algorithm, and then subdivided according to the similarity of hash values.
[0065] The dataset to be imported is divided into multiple grouped datasets using the method described above. The number of groups can be set according to the actual situation. For each grouped dataset, before importing into each retrieval cluster, a group index corresponding to the grouped dataset is built in each retrieval cluster based on the group identifier and the current date information, ensuring that each grouped dataset has its own independent index structure. For example, if the current date information is XXXX year YY month ZZ day and the group identifier is 2, then the group index of the grouped dataset can be named "XXXXYYZZ_group2".
[0066] This application embodiment groups the dataset to be imported and constructs independent group indexes for each group based on group identifiers and current date information. This enables refined management and maintenance of large datasets, dynamically adapts to changes in data scale, and solves the problem of poor scalability of a single index. When dealing with massive amounts of data, this grouped indexing method can significantly improve the efficiency of index import and query operations.
[0067] In one embodiment, updating the data index of the target retrieval cluster includes: Based on each of the grouping indexes, the grouping datasets are imported into the target retrieval cluster in parallel; For each grouped dataset, switch the second index alias from the group index of the original grouped dataset to the group index of the newly imported grouped dataset.
[0068] Specifically, multi-threaded or distributed tasks are initiated to import each grouped dataset into the storage space of its corresponding index in the target retrieval cluster in parallel, thereby maximizing the utilization of cluster resources and shortening the overall import time. After the import is complete, the total amount of the imported grouped datasets is recorded.
[0069] At the same time, obtain the total amount of historically imported grouped datasets. The total amount of historically imported grouped datasets can be the total amount of the last imported grouped dataset or the total amount of multiple historically imported grouped datasets, and there is no limitation here.
[0070] Based on the data volume fluctuation algorithm and the total amount of historical imported group datasets, the normal range of the total amount of imported group data is determined. The data volume fluctuation algorithm can be set according to actual business needs and data characteristics.
[0071] If the total amount of the currently imported group dataset does not meet the normal range for the total amount of imported group data, it indicates that the currently imported group dataset is invalid, triggering an alarm mechanism and requiring the group dataset to be re-imported. You can also check and optimize the import process before performing the re-import operation to improve the success rate of subsequent imports.
[0072] If the total volume of the currently imported grouped dataset meets the normal range for total imported grouped data, it indicates that the currently imported grouped dataset is normal. The second index alias will then be switched from the grouped index of the original grouped dataset to the grouped index of the newly imported grouped dataset. Switching the second index alias allows for rapid index updates without directly affecting ongoing queries. After switching the second index alias, the retrieval system will automatically redirect subsequent query requests to the newly imported grouped dataset index, ensuring data real-time performance and accuracy.
[0073] It should be noted that if this multi-cluster index update process is the first parallel import based on the grouped index, then the total amount of all grouped datasets is counted to obtain the current total amount of imported data. At the same time, the total amount of historical imported data is obtained. The normal range of the total amount of imported data is calculated by using the data volume fluctuation algorithm and the total amount of historical imported data. Abnormal data import situations are identified based on the normal range of the total amount of imported data.
[0074] This application embodiment imports grouped datasets in parallel through each grouped index. Within the limits of cluster resources, efficiency is positively correlated with the grouped data. It fully utilizes distributed capabilities, significantly shortens the import cycle, and also supports mechanisms for automatic detection and recovery of import anomalies to ensure data quality. Under normal import conditions, it achieves seamless index updates through index alias switching, ensuring the stability of the retrieval service.
[0075] Based on all the above embodiments, the following is combined with Figure 5 This paper describes a specific implementation process of a multi-cluster index update method. Figure 5 This is the second flowchart of the multi-cluster index update method provided in this application.
[0076] The datasets to be imported are grouped according to different data types, resulting in multiple grouped datasets. Business traffic from the Elasticsearch cluster currently preparing to update the index is gradually switched to the remaining Elasticsearch clusters. Multiple grouped datasets are processed simultaneously, and a corresponding grouped index is created for each grouped dataset based on its group identifier and current date.
[0077] After creating the group index for each group dataset and completing the full business traffic switchover for the current Elasticsearch cluster, import each group dataset in parallel into the storage space of its corresponding group index within the current Elasticsearch cluster. After the import is complete, verify the current imported data volume against the historical imported data volume to determine if there are any anomalies. If an anomaly is found, the import failed and needs to be re-performed. If no anomalies are found, the import was successful, and proceed to the next step: switching the index alias from the group index of the old group dataset to the group index of the newly imported group dataset, until all group indexes have been replaced.
[0078] Repeat the index update process described above until all Elasticsearch clusters have completed the index update. Further, redistribute the service traffic across the Elasticsearch clusters, ensuring balanced configuration across all Elasticsearch clusters.
[0079] The above method is particularly suitable for high-concurrency scenarios, enabling efficient data import and seamless updates to the data index. It solves the problem of rapidly importing large amounts of data into the retrieval system under high concurrency conditions while avoiding business query delays caused by data import and index switching updates, thus meeting the needs for timely data updates and immediate, efficient queries from the user side.
[0080] Figure 6 This is a schematic diagram of the structure of the multi-cluster index update device provided in this application.
[0081] like Figure 6 As shown, the multi-cluster index update device includes: The multi-cluster index update module 610 is used to determine each retrieval cluster in the retrieval cluster set as the target retrieval cluster for the current index update, and smoothly migrate the business traffic of the target retrieval cluster to the remaining retrieval clusters, wherein the remaining retrieval clusters are the retrieval clusters in the retrieval cluster set other than the target retrieval cluster; and update the data index of the target retrieval cluster after the business traffic diversion of the target retrieval cluster is completed. The traffic configuration module 620 is used to allocate corresponding business traffic to each of the search clusters after each search cluster has completed the data index update.
[0082] The multi-cluster index update device provided in this application smoothly migrates business traffic to other clusters before updating a single cluster, and performs independent updates for each cluster after the full traffic migration is completed. This avoids the instability of interface response caused by updating the index of the entire cluster at the same time. Traffic allocation is restored uniformly after all clusters have been updated. This keeps the business impact generated during the index update iteration process at a very low level, greatly improves the overall performance and stability of the retrieval system, and thus enhances the user experience.
[0083] In one embodiment, the multi-cluster index update module 610 is further configured to: Based on the performance metrics of the remaining retrieval clusters, the traffic flow ratio of the remaining retrieval clusters is determined; the traffic flow ratio is used to indicate the proportion of business traffic migrated to the remaining retrieval clusters. Based on a multi-stage traffic migration strategy, the business traffic of the target retrieval cluster is migrated to the remaining retrieval clusters in stages according to the traffic flow ratio.
[0084] In one embodiment, the multi-cluster index update module 610 is further configured to: At the current stage, based on a multi-stage traffic migration strategy, the business traffic of the target retrieval cluster is migrated to the other retrieval clusters according to the traffic flow ratio; Determine the interface response efficiency of the remaining retrieval clusters; If the interface response efficiency is less than a preset response efficiency threshold, the multi-stage traffic migration strategy is adjusted to obtain an adjusted multi-stage traffic migration strategy. The adjusted multi-stage traffic migration strategy is determined as the multi-stage traffic migration strategy, and the next stage is determined as the current stage. The steps of migrating the business traffic of the target retrieval cluster to the other retrieval clusters according to the traffic flow ratio are iteratively executed in the current stage based on the multi-stage traffic migration strategy until the last stage is executed. If the interface response efficiency is greater than or equal to the preset response efficiency threshold, the multi-stage traffic migration strategy is iteratively executed. In the current stage, the business traffic of the target retrieval cluster is migrated to the other retrieval clusters according to the traffic flow ratio, until the last stage is executed.
[0085] In one embodiment, the multi-cluster index update module 610 is further configured to: Import the dataset to be imported into the target retrieval cluster; Obtain the total current imported data volume and the total historical imported data volume of the target retrieval cluster; Based on the data volume fluctuation algorithm and the total amount of historical imported data, the normal range of the total amount of imported data is determined; If the current total amount of imported data does not meet the normal range of the total amount of imported data, then return to the step of importing the dataset to be imported into the target retrieval cluster; If the total amount of imported data meets the normal range of the total amount of imported data, then the first index alias will be switched from the index of the original dataset to the index of the newly imported dataset.
[0086] In one embodiment, the multi-cluster index update module 610 is further configured to: The dataset to be imported is divided according to its data type, resulting in multiple grouped datasets; For each grouped dataset, a group index is constructed in each of the retrieval clusters based on the group identifier and the current date information.
[0087] In one embodiment, the multi-cluster index update module 610 is further configured to: Based on each of the grouping indexes, the grouping datasets are imported into the target retrieval cluster in parallel; For each grouped dataset, switch the second index alias from the group index of the original grouped dataset to the group index of the newly imported grouped dataset.
[0088] It should be noted that the multi-cluster index update device provided in this application can execute the multi-cluster index update method described in any of the above embodiments during actual operation, which will not be elaborated in this embodiment.
[0089] Figure 7 This is a schematic diagram of the structure of the electronic device provided in this application, such as... Figure 7 As shown, the electronic device may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740, wherein the processor 710, communications interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute a multi-cluster index update method. This method includes: for each retrieval cluster in the retrieval cluster set, determining the retrieval cluster as the target retrieval cluster for the current index update; smoothly migrating the service traffic of the target retrieval cluster to the remaining retrieval clusters, wherein the remaining retrieval clusters are the retrieval clusters remaining in the retrieval cluster set besides the target retrieval cluster; updating the data index of the target retrieval cluster after the service traffic diversion of the target retrieval cluster is completed; and allocating corresponding service traffic to each retrieval cluster after each retrieval cluster has completed its data index update.
[0090] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0091] On the other hand, this application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer can execute the multi-cluster index update method provided in the above embodiments. The method includes: for each retrieval cluster in the retrieval cluster set, determining the retrieval cluster as the target retrieval cluster for the current index update; smoothly migrating the business traffic of the target retrieval cluster to the remaining retrieval clusters, wherein the remaining retrieval clusters are the retrieval clusters remaining in the retrieval cluster set besides the target retrieval cluster; updating the data index of the target retrieval cluster when the business traffic diversion of the target retrieval cluster is completed; and allocating corresponding business traffic to each retrieval cluster when the data index update of each retrieval cluster is completed.
[0092] In another aspect, this application also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program is implemented to perform the multi-cluster index update method provided in the above embodiments. The method includes: for each retrieval cluster in a set of retrieval clusters, determining the retrieval cluster as the target retrieval cluster for the current index update; smoothly migrating the service traffic of the target retrieval cluster to the remaining retrieval clusters, wherein the remaining retrieval clusters are the retrieval clusters remaining in the set of retrieval clusters excluding the target retrieval cluster; updating the data index of the target retrieval cluster when the service traffic diversion of the target retrieval cluster is completed; and allocating corresponding service traffic to each retrieval cluster when the data index update of each retrieval cluster is completed.
[0093] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0094] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A multi-cluster index update method, characterized in that, The multi-cluster index update method includes: For each retrieval cluster in the retrieval cluster set, the retrieval cluster is determined as the target retrieval cluster for the current index update, and the business traffic of the target retrieval cluster is smoothly migrated to the remaining retrieval clusters, wherein the remaining retrieval clusters are the retrieval clusters in the retrieval cluster set other than the target retrieval cluster; after the business traffic diversion of the target retrieval cluster is completed, the data index of the target retrieval cluster is updated. Once each retrieval cluster has completed its data index update, corresponding business traffic is allocated to each retrieval cluster.
2. The multi-cluster index update method according to claim 1, characterized in that, The step of smoothly migrating the service traffic of the target retrieval cluster to other retrieval clusters includes: Based on the performance metrics of the remaining retrieval clusters, the traffic flow ratio of the remaining retrieval clusters is determined; the traffic flow ratio is used to indicate the proportion of business traffic migrated to the remaining retrieval clusters. Based on a multi-stage traffic migration strategy, the business traffic of the target retrieval cluster is migrated to the remaining retrieval clusters in stages according to the traffic flow ratio.
3. The multi-cluster index update method according to claim 2, characterized in that, The multi-stage traffic migration strategy, which migrates the service traffic of the target retrieval cluster to the remaining retrieval clusters in stages according to the traffic flow ratio, includes: At the current stage, based on a multi-stage traffic migration strategy, the business traffic of the target retrieval cluster is migrated to the other retrieval clusters according to the traffic flow ratio; Determine the interface response efficiency of the remaining retrieval clusters; If the interface response efficiency is less than a preset response efficiency threshold, the multi-stage traffic migration strategy is adjusted to obtain an adjusted multi-stage traffic migration strategy. The adjusted multi-stage traffic migration strategy is determined as the multi-stage traffic migration strategy, and the next stage is determined as the current stage. The steps of migrating the business traffic of the target retrieval cluster to the other retrieval clusters according to the traffic flow ratio are iteratively executed in the current stage based on the multi-stage traffic migration strategy until the last stage is executed. If the interface response efficiency is greater than or equal to the preset response efficiency threshold, the multi-stage traffic migration strategy is iteratively executed. In the current stage, the business traffic of the target retrieval cluster is migrated to the other retrieval clusters according to the traffic flow ratio, until the last stage is executed.
4. The multi-cluster index update method according to claim 1, characterized in that, The step of updating the data index of the target retrieval cluster includes: Import the dataset to be imported into the target retrieval cluster; Obtain the total current imported data volume and the total historical imported data volume of the target retrieval cluster; Based on the data volume fluctuation algorithm and the total amount of historical imported data, the normal range of the total amount of imported data is determined; If the current total amount of imported data does not meet the normal range of the total amount of imported data, then return to the step of importing the dataset to be imported into the target retrieval cluster; If the total amount of imported data meets the normal range of the total amount of imported data, then the first index alias will be switched from the index of the original dataset to the index of the newly imported dataset.
5. The multi-cluster index update method according to claim 1, characterized in that, Before updating the data index for each of the aforementioned retrieval clusters, the following steps are also included: The dataset to be imported is divided according to its data type, resulting in multiple grouped datasets; For each grouped dataset, a group index is constructed in each of the retrieval clusters based on the group identifier and the current date information.
6. The multi-cluster index update method according to claim 5, characterized in that, The step of updating the data index of the target retrieval cluster includes: Based on each of the grouping indexes, the grouping datasets are imported into the target retrieval cluster in parallel; For each grouped dataset, switch the second index alias from the group index of the original grouped dataset to the group index of the newly imported grouped dataset.
7. A multi-cluster index update device, characterized in that, The multi-cluster index update device includes: The multi-cluster index update module is used to determine each retrieval cluster in the retrieval cluster set as the target retrieval cluster for the current index update, and smoothly migrate the business traffic of the target retrieval cluster to the remaining retrieval clusters, wherein the remaining retrieval clusters are the retrieval clusters in the retrieval cluster set other than the target retrieval cluster; after the business traffic diversion of the target retrieval cluster is completed, the data index of the target retrieval cluster is updated. The traffic configuration module is used to allocate corresponding business traffic to each of the search clusters after each search cluster has completed the data index update.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the multi-cluster index update method as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium, wherein a computer program is stored on the non-transitory computer-readable storage medium, characterized in that, When the computer program is executed by a processor, it implements the steps of the multi-cluster index update method as described in any one of claims 1 to 6.
10. A computer program product, the computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the multi-cluster index update method as described in any one of claims 1 to 6.