Distributed information storage and retrieval optimization method

By employing a hierarchical node cluster organization and intelligent load balancing strategy, the bottleneck of central nodes and the problem of static indexes in distributed information storage and retrieval are solved, resulting in an efficient, reliable, and easily scalable information storage and retrieval solution.

CN120892618APending Publication Date: 2025-11-04GUANGZHOU ZHIBO TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511263648.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-11-04

Smart Images

  • Figure CN120892618A_ABST
    Figure CN120892618A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of distributed information storage and retrieval, in particular to a distributed information storage and retrieval optimization method, which comprises a layered node cluster architecture, a dynamic index updating strategy, a decentration collaboration mechanism and an intelligent load distribution method. The system pressure is dispersed through a multi-layer architecture, the maintenance efficiency is improved through dynamic indexing, the fault-tolerant capability is enhanced through local consensus decision, and the delay and the network overhead are reduced through combination of a prediction model and proximity sensing routing optimization. According to the method, the problems of bottleneck of a central node, low efficiency of a static index, lengthy retrieval path and the like can be effectively solved, and a distributed information storage and retrieval solution which is efficient, reliable and high in expansibility is provided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0002] The application belongs to the technical field of information technology and data processing, and specifically relates to a distributed information storage and retrieval optimization method. BACKGROUND

[0003] With the rapid development of distributed information system technology, distributed information storage and retrieval technology has been widely applied in cloud computing, big data processing, Internet of Things and industrial automation fields. However, in the face of the demand for efficient storage and rapid and accurate retrieval of massive data, the existing technology still faces problems such as high system response delay, complex data consistency guarantee, insufficient retrieval efficiency and limited architecture scalability. Especially in the high-concurrency access scenario, the traditional distributed architecture is difficult to balance the storage performance and retrieval real-time performance at the same time, which restricts its further promotion in intelligent and real-time applications.

[0004] According to the search, the patent document with the publication number CN106528876B discloses an information processing method and a distributed information processing system of a distributed system, and the publication date is August 23, 2019. The patent proposes to separate the meta-information operation and the data operation, to manage the meta-information through the setting of the center node, and to take multiple data nodes to be responsible for data storage and processing, so as to realize the function division to improve the overall processing efficiency of the system. However, this scheme still has obvious deficiencies in actual application: it relies on the center node to uniformly process all meta-information requests, and in the high-concurrency scenario, it is easy to form a performance bottleneck, resulting in an increase in meta-information access delay; in addition, the architecture does not optimize the data index mechanism, and the retrieval operation still needs to traverse or broadcast the query, resulting in large network overhead and slow response speed, which is difficult to meet the information retrieval demand of low delay and high throughput.

[0005] According to the search, the patent document with the publication number CN108376177B discloses a method and a distributed system for processing information, and the publication date is October 25, 2019. The patent proposes that the main control node distributes retrieval requests to multiple data processing nodes, each node constructs a picture feature index based on the cluster center and performs local retrieval, and finally aggregates the results to improve the retrieval efficiency. This method realizes parallel retrieval to a certain extent and improves the query performance of picture data. However, it still has the following limitations: first, the index construction relies on the static division of the cluster center, and when facing a dynamically changing data set, the model needs to be frequently retrained, which has a high maintenance cost; second, the main control node bears all the request distribution and result aggregation tasks, which has a single point failure risk, and as the node scale expands, the calculation and communication burden of the main control node significantly increases; in addition, this scheme mainly faces the image feature matching scene, and lacks efficient index support for general structured or semi-structured data, which limits the universality and scalability.

[0006] The above problems show that the current existing distributed information storage and retrieval technology still generally relies on centralized control or static index mechanism in architecture design, and it is difficult to realize efficient parallel retrieval and low delay response while ensuring system scalability. Especially in the context of continuous growth of data size and increasing complexity of query mode, the existing scheme has obvious short board in load balancing, dynamic index updating, and decentralized cooperation. Therefore, a new type of distributed information storage and retrieval optimization method is needed, which can realize intelligent index distribution, dynamic load scheduling and efficient parallel retrieval on the basis of decentralized architecture, so as to comprehensively improve the storage efficiency and retrieval performance of the system.

[0007] The present application aims to provide a distributed information storage and retrieval optimization method to overcome the performance limitations caused by the central node bottleneck, static index structure and inefficient retrieval path in the prior art, and meet the comprehensive needs of modern information systems for high concurrency, low delay and strong scalability. SUMMARY

[0008] The present application provides a distributed information storage and retrieval optimization method, which aims to solve the performance limitations caused by the central node bottleneck, static index structure and inefficient retrieval path in the prior art by introducing a multi-level dynamic index mechanism, a decentralized cooperative architecture and an intelligent load distribution strategy, and meet the comprehensive needs of modern information systems for high concurrency, low delay and strong scalability.

[0009] The technical scheme of the present application is as follows:

[0010] In the overall architecture design, the method adopts a hierarchical node cluster organization form. The bottom layer is a data storage unit group, each unit including a plurality of storage modules, and the modules are connected through a consistent hash algorithm to realize data distribution and redundancy backup. The middle layer is an index management unit group, each unit consisting of a plurality of index generators and index updaters, responsible for dynamically building and maintaining partitioned indexes. The top layer is a task scheduling unit group, including a plurality of cooperative schedulers, used for coordinating the distribution and result aggregation of retrieval requests. The above three layers interact through a high-speed communication protocol to ensure the efficiency of information transmission.

[0011] In the index mechanism, the present application proposes a dynamic index updating strategy based on time window and access frequency. Specifically, each index generator statistically analyzes the data access mode according to the preset time window, marks the high-frequency accessed data items as hot data, and preferentially loads the corresponding index entries into the memory cache area. At the same time, through the sliding window mechanism, the change of data access frequency is monitored in real time, and the storage location of the index entry is dynamically adjusted, so as to reduce the invalid index operation on cold data. In addition, the index updater regularly merges and compresses the partitioned index to avoid index fragmentation problem.

[0012] To realize the decentralized collaborative architecture, the application designs a distributed decision mechanism based on local consensus. The schedulers in each task scheduling unit group share the current task load state through local broadcast, and select the optimal retrieval path according to the load balancing algorithm. When a certain scheduler detects that its load exceeds the preset threshold, it will automatically trigger the task migration process and distribute part of the retrieval request to other idle schedulers. This process is completed through a lightweight message queue, ensuring the low delay characteristics of task migration.

[0013] In the load distribution strategy, the application adopts a dynamic weight distribution method based on a prediction model. Each storage module trains a resource consumption prediction model based on its historical access records to estimate the computing resources required to process different retrieval requests. The scheduler assigns appropriate weights to each storage module when distributing tasks, giving priority to modules with sufficient resources for complex requests. At the same time, the weight parameters are adjusted through a periodic feedback mechanism to ensure the fairness and stability of load distribution.

[0014] To address data consistency issues, the application introduces a multi-version concurrency control mechanism. Each storage module generates a unique timestamp for each write request and temporarily stores the updated data item in the buffer area. Only when all related modules confirm the update success, the data item will be formally submitted to the main storage area. In this process, read requests can access the latest stable version or temporary version according to the timestamp, thereby improving read efficiency while ensuring consistency.

[0015] To reduce network overhead and improve response speed, the application designs a routing optimization algorithm based on proximity awareness. Each storage module records the information of several modules closest to it in initialization and registers these information to the index management unit group. When a retrieval request involves cross-module queries, the index generator will preferentially select adjacent modules as query targets, reducing the number of data transmission hops. In addition, by establishing direct channels between modules, the data transmission path is further shortened.

[0016] The application achieves the following technical effects through the above technical means: First, the hierarchical node cluster organization effectively disperses system pressure and avoids single nodes becoming performance bottlenecks; second, the dynamic index update strategy significantly improves index maintenance efficiency and reduces resource waste caused by invalid operations; third, the decentralized collaborative architecture improves the fault tolerance of the system and reduces the risk of single point failure; finally, the load distribution method based on the prediction model and the routing optimization algorithm based on proximity awareness work together to significantly reduce system latency and improve throughput.

[0017] In summary, the present application solves the problems of central node bottleneck, static index structure and inefficient search path in the prior art by means of multi-level architecture design, dynamic index mechanism, decentralized cooperation strategy and intelligent load distribution, and provides an efficient, reliable and easily expandable solution for the field of distributed information storage and retrieval. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 The figure is a schematic diagram of the overall architecture of the present application, showing a hierarchical node cluster organization form, including a bottom layer of data storage unit group, a middle layer of index management unit group and a top layer of task scheduling unit group, and a high-speed communication protocol connection therebetween.

[0019] Figure 2 The figure is a flowchart of the dynamic index update strategy, which describes in detail the index entry priority adjustment process based on time window and access frequency, including hot data identification, memory cache area loading and index merging and compression operation.

[0020] Figure 3 The figure is a schematic diagram of the task migration mechanism of the decentralized cooperation architecture, showing the specific process of the task scheduler in the task scheduling unit group sharing the load state through local broadcast and triggering task migration in the case of high load.

[0021] The reference signs are as follows: 1, data storage unit group; 2, index management unit group; 3, task scheduling unit group; 4, hot data; 5, memory cache area; 6, index generator; 7, index updater; 8, scheduler; 9, task migration process; 10, load state broadcast. DETAILED DESCRIPTION

[0022] The present application provides a distributed information storage and retrieval optimization method, and the specific embodiments are combined with the structures and processes shown in the drawings for detailed description. Figure 1 to the drawings Figure 3 The following description will focus on the core parts such as overall architecture design, dynamic index update strategy, decentralized cooperation architecture and load distribution strategy, and the connection relationship, position relationship and mutual cooperation relationship between each part will be described one by one.

[0023] As Figure 1As shown, the overall architecture adopts a hierarchical node cluster organization form, including the bottom layer data storage unit group 1, the middle layer index management unit group 2 and the top layer task scheduling unit group 3. The data storage unit group 1 is composed of several storage modules, which realize data distribution and redundant backup through a consistent hash algorithm. Specifically, each storage module internally contains a data storage area and a buffer area, the data storage area is used to store primary data items, and the buffer area is used to temporarily store temporary data in update operations. The storage modules are directly connected through a high-speed communication protocol, ensuring the efficiency of data synchronization and backup operations. The index management unit group 2 is located in the middle layer and is composed of multiple index generators 6 and index updaters 7. The index generator 6 is responsible for dynamically constructing index entries according to the data access mode, while the index updater 7 periodically merges and compresses the partitioned index to reduce fragmentation problems. The task scheduling unit group 3 is located at the top layer and contains multiple cooperative schedulers 8, which share the current task load state through a local broadcast mechanism and select the optimal path to complete the distribution and result aggregation of retrieval requests according to the load balancing algorithm. The three layers are connected through a high-speed communication protocol, which is based on low-latency message queue technology and can support fast transmission of large-scale concurrent requests.

[0024] In terms of dynamic index update strategy, as shown in Figure 2 The index generator 6 first analyzes the data access mode according to the preset time window, identifies the hot data 4 with high frequency access, and loads the index entries of the hot data 4 into the memory cache area 5 first, thereby improving the retrieval efficiency. The sliding window mechanism plays a key role in this process, which monitors the changes in data access frequency in real time and dynamically adjusts the storage location of the index entries. For example, when the access frequency of a certain data item decreases, its corresponding index entry will be migrated from the memory cache area 5 to the disk storage area to release memory resources. At the same time, the index updater 7 periodically performs merging and compression operations on the partitioned index to avoid index fragmentation problems caused by frequent updates. The specific implementation of the above process relies on the timestamp marking mechanism, that is, each index entry is attached with a timestamp to record the time point of its last access or update, thereby providing a basis for subsequent priority adjustment.

[0025] The design of the decentralized collaborative architecture is as shown in Figure 3As shown, the schedulers 8 within the task scheduling unit group 3 share the load status information 10 through local broadcast. Specifically, each scheduler 8 periodically sends messages containing the current load status to other schedulers within the same group, and the messages are delivered through lightweight message queues. When a scheduler 8 detects that its load exceeds a preset threshold, it triggers the task migration process 9. The core of the task migration process 9 is to select a suitable idle scheduler 8 as the target node and distribute part of the retrieval request to the node. In order to ensure the low latency characteristics of the migration process, the task migration process 9 adopts a proximity-aware routing optimization algorithm. The algorithm records the information of several modules closest to each storage module in the initialization stage and registers these information to the index management unit group 2. When the retrieval request involves cross-module queries, the index generator 6 will preferentially select adjacent modules as query targets, thereby reducing the number of hops of data transmission and improving response speed.

[0026] On the load distribution strategy, the application adopts a dynamic weight distribution method based on a prediction model. Each storage module trains a resource consumption prediction model based on its historical access records. The model estimates the computing resources required to process different retrieval requests by analyzing the complexity characteristics of these requests. The scheduler 8 assigns appropriate weights to each storage module when distributing tasks, taking into account the prediction results of each storage module. For example, for retrieval requests with high complexity, the scheduler 8 will preferentially assign them to storage modules with sufficient resources to avoid performance bottlenecks caused by insufficient resources. In addition, the system also introduces a periodic feedback mechanism to dynamically adjust the weight parameters by collecting the actual resource consumption of each storage module, thereby ensuring the fairness and stability of load distribution.

[0027] For data consistency guarantee, the application introduces a multi-version concurrency control mechanism. When each storage module receives a write request, it first generates a unique timestamp for it and temporarily stores the updated data item in the buffer area. Only when all related modules confirm that the update is successful will the data item be formally submitted to the main storage area. In this process, read requests can choose to access the latest stable version or temporary version according to the timestamp, thereby improving read efficiency while ensuring consistency. Specifically, when a read request arrives, the storage module will determine whether it needs to access the temporary version in the buffer area according to the timestamp of the request. If the timestamp is newer, it will return the temporary version directly; otherwise, it will return the stable version in the main storage area.

[0028] To reduce network overhead and improve response speed, the application designs a routing optimization algorithm based on proximity awareness. In the initialization phase, the algorithm records the information of several modules closest to each module in the storage module, and registers these information to the index management unit group 2. When the retrieval request involves cross-module query, the index generator 6 will preferentially select the adjacent module as the query target, reducing the number of data transmission hops. In addition, by establishing direct channels between modules, the data transmission path is further shortened. For example, in a certain cross-module retrieval scenario, the index generator 6 first determines the module where the target data is located according to the request content, and then selects the closest module as the proxy node through the proximity awareness algorithm, and finally returns the query result to the requester through the proxy node.

[0029] The specific implementation process of each part is as follows: when the system receives a retrieval request, the scheduler 8 in the task scheduling unit group 3 first obtains the load state information 10 of each scheduler through local broadcast mechanism, and selects the optimal retrieval path according to the load balancing algorithm. Subsequently, the scheduler 8 distributes the retrieval request to the index generator 6 in the index management unit group 2. The index generator 6 queries the index entry in the memory cache area 5 according to the request content, and if it hits the hot data 4, it directly returns the result; otherwise, it queries the partition index in the disk storage area through the index updater 7. In the cross-module query scenario, the index generator 6 will preferentially select the adjacent module as the query target to reduce the number of data transmission hops. Finally, the query result is aggregated by the task scheduling unit group 3 and returned to the requester.

[0030] In the whole process, the connection relationship and position relationship between each component cooperate closely to achieve the goal of the application. For example, the high-speed communication protocol between the data storage unit group 1 and the index management unit group 2 ensures the real-time performance of index entry update, and the local broadcast mechanism between the index management unit group 2 and the task scheduling unit group 3 improves the sharing efficiency of load state information. In addition, the routing optimization algorithm based on proximity awareness effectively reduces the network overhead of cross-module query, thereby improving the overall performance of the system. In order to better enable relevant persons in the art to fully understand and implement the application, the following supplementary explanation of the specific implementation principle of the application is made in conjunction with a specific application scenario.

[0031] In a certain distributed information storage and retrieval system, assume that the system is applied to the commodity data management scenario of a certain large e-commerce platform. The platform needs to process millions of commodity query requests every day, and the data size continues to grow, and the query mode is complex and diverse. The following are the specific running steps and implementation principles of the technical solution based on the application in this scenario.

[0032] Firstly, when the system receives a user's commodity retrieval request, the scheduler 8 in the task scheduling unit group 3 obtains the load state information 10 of the current schedulers through the local broadcast mechanism. In this process, each scheduler 8 periodically sends a message containing its load state to other schedulers in the same group, and these messages are delivered through a lightweight message queue. The scheduler 8 selects the optimal path according to the load balancing algorithm and distributes the retrieval request to the index generator 6 in the index management unit group 2. On this basis, if a certain scheduler 8 detects that its load exceeds the preset threshold, it will trigger the task migration process 9. The core of the task migration process 9 is to select a suitable idle scheduler as the target node and distribute part of the retrieval request to the node. In order to ensure the low delay characteristics of the migration process, the task migration process 9 uses a routing optimization algorithm based on proximity awareness. This algorithm records the information of several modules closest to it in the initialization stage for each storage module and registers these information to the index management unit group 2. In this way, the system can quickly locate the adjacent module as the query target, thereby reducing the number of hops of data transmission and improving the response speed.

[0033] Subsequently, the index generator 6 queries the index entries in the memory cache area 5 according to the content of the retrieval request. As shown in Figure 2 , the index generator 6 first analyzes the data access pattern according to the preset time window and identifies the hot data 4 with high frequency access. The index entries of the hot data 4 are preferentially loaded into the memory cache area 5, thereby improving the retrieval efficiency. If the retrieval request hits the hot data 4, the result is returned directly; otherwise, the index generator 6 queries the partition index in the disk storage area through the index updater 7. The sliding window mechanism plays a key role in this process, which monitors the changes of data access frequency in real time and dynamically adjusts the storage location of the index entries. For example, when the access frequency of a certain commodity data item decreases, its corresponding index entry will be migrated from the memory cache area 5 to the disk storage area to release the memory resources. At the same time, the index updater 7 regularly performs merging and compression operations on the partition index to avoid the problem of index fragmentation caused by frequent updates. The specific implementation of the above process relies on the timestamp marking mechanism, that is, each index entry is attached with a timestamp recording the time point of its last access or update, thereby providing a basis for subsequent priority adjustment.

[0034] In the cross-module query scenario, the index generator 6 will prefer to select the adjacent module as the query target. For example, when the user queries a certain type of specific goods, the index generator 6 first determines the target data module according to the request content, and then selects the nearest module as the proxy node through the proximity awareness algorithm. In this way, the system can significantly reduce the hop count of data transmission, thereby reducing network overhead and improving response speed. In addition, by establishing a direct channel between modules, the data transmission path is further shortened. For example, in a certain cross-module retrieval scenario, the index generator 6 first determines the target data module according to the request content, and then selects the nearest module as the proxy node through the proximity awareness algorithm, and finally returns the query result to the requester via the proxy node.

[0035] In the load distribution strategy, the system adopts a dynamic weight distribution method based on a prediction model. Each storage module trains a resource consumption prediction model based on its historical access records. The model estimates the computing resources required to process different retrieval requests by analyzing the complexity characteristics of these requests. The scheduler 8 assigns appropriate weights to each storage module when distributing tasks, based on the prediction results. For example, for retrieval requests with high complexity, the scheduler 8 will preferentially assign them to storage modules with sufficient resources to avoid performance bottlenecks caused by insufficient resources. In addition, the system also introduces a periodic feedback mechanism to dynamically adjust the weight parameters based on the actual resource consumption of each storage module, thereby ensuring the fairness and stability of load distribution.

[0036] For data consistency guarantee, the system introduces a multi-version concurrency control mechanism. When receiving a write request, each storage module first generates a unique timestamp for it and temporarily stores the updated data item in the buffer area. Only when all related modules confirm the update success, the data item will be formally submitted to the main storage area. In this process, read requests can access the latest stable version or temporary version according to the timestamp, thereby improving read efficiency while ensuring consistency. Specifically, when a read request arrives, the storage module will determine whether to access the temporary version in the buffer area according to the timestamp of the request. If the timestamp is new, the temporary version is returned directly; otherwise, the stable version in the main storage area is returned.

[0037] In the whole process, the connection relationship and position relationship between each component cooperate closely to achieve the goal of the invention. For example, the high-speed communication protocol between the data storage unit group 1 and the index management unit group 2 ensures the real-time updating of index entries, while the local broadcast mechanism between the index management unit group 2 and the task scheduling unit group 3 improves the sharing efficiency of load state information. In addition, the routing optimization algorithm based on proximity awareness effectively reduces the network overhead of cross-module queries, thereby improving the overall performance of the system.

[0038] Through the combination of the above steps and principles, the application realizes efficient information storage and retrieval in the application scenario of the e-commerce platform. The system not only can cope with the performance requirements in the high-concurrency access scenario, but also can maintain low latency and high throughput in the case of continuous growth of data size. This design not only solves the problems of center node bottleneck, static index structure and inefficient retrieval path in the prior art, but also provides an efficient, reliable and easily scalable solution for the field of distributed information storage and retrieval.

Claims

1. A distributed information storage and retrieval optimization method, characterized in that, Includes the following steps: Data distributed by the storage modules in the data storage unit group (1) through the consistent hashing algorithm is obtained; the index generator (6) in the index management unit group (2) analyzes the data access pattern according to the preset time window and identifies hot data (4), and loads the index entries of hot data (4) into the memory cache area (5) first; the scheduler (8) in the task scheduling unit group (3) selects the optimal retrieval path according to the load balancing algorithm and distributes the retrieval request to the index management unit group (2); when the scheduler (8) detects that its own load exceeds the preset threshold, it triggers the task migration process (9) and distributes some retrieval requests to other idle schedulers (8).

2. The distributed information storage and retrieval optimization method according to claim 1, characterized in that, The index updater (7) in the index management unit group (2) periodically performs merging and compression operations on the partition indexes to avoid index fragmentation problems.

3. The distributed information storage and retrieval optimization method according to claim 1, characterized in that, Each storage module trains a resource consumption prediction model based on its historical access records. The scheduler (8) assigns dynamic weights to each storage module based on its prediction results, and prioritizes assigning more complex retrieval requests to storage modules with sufficient resources.

Citation Information

Patent Citations

  • Information processing methods and distributed information processing systems

    CN106528876B

  • Methods and distributed systems for processing information

    CN108376177B