Method and system for AI-driven real-time database analysis based on distributed storage
By using AI technology in real-time analysis databases for data classification and hot data storage node management in distributed storage architecture, the problem of unbalanced storage nodes is solved, data processing performance and system stability are improved, and operation and maintenance costs are reduced.
Patent Information
- Application Number
- CN202510669061.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-23
AI Technical Summary
When existing real-time analytical databases process large amounts of data, storage node pressure may occur, and some nodes have excessive storage load, resulting in overall performance degradation.
Using a distributed storage architecture based on AI technology, through logical classification of data sets, similar or related data are divided into the same data block, and a hot data storage node is set up in the distributed storage system to monitor the real-time access of the data block, and frequently called data blocks are migrated to the hot data storage node, marked as hot data blocks, and the data migration buffer is managed when necessary.
It improves the response speed of data query and data processing performance, meets the requirements of high concurrency and low latency of real-time analysis systems, reduces frequent data migration caused by periodic fluctuations, and reduces the operation and maintenance costs and resource waste of the overall system.
Smart Images

Figure CN120196683A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of analysis library data processing, and specifically to a method and system for an AI-driven real-time analysis database based on distributed storage. Background Art
[0002] A real-time analysis database is a type of database system specifically designed to support real-time data processing and analysis. Its main goal is to be able to efficiently process, query, and analyze data simultaneously or shortly after the data is generated, so as to provide instant feedback for business decisions.
[0003] The real-time analysis database quickly collects and processes a continuous stream of data (such as logs, sensor data, user behavior, etc.). It usually takes only a few seconds or even less time for the data to enter the system and generate analysis results, ensuring that the business system can respond to various changes in real time. However, in the case of a large amount of data, the system may experience delays during real-time data collection, processing, and calculation. Especially in complex analysis and aggregation operations, the delay may exceed the expected range. This is because a large amount of data requires the distributed storage system to have sufficient scalability and load balancing capabilities. In practical applications, it is easy to have an uneven pressure on storage nodes, with some nodes having an overly heavy storage load, resulting in a decline in overall performance. Summary of the Invention
[0004] Aiming at the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide a method and system for an AI-driven real-time analysis database based on distributed storage, which can shard the data set of the database and set up a hot data storage node with high load performance, and preferentially migrate the hot storage data to the hot data storage node to reduce the occurrence of overloaded storage node pressure.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A method for an AI-driven real-time analysis database based on distributed storage, the method comprising the following steps:
[0006] Obtain the data set of the analysis database, logically classify the data in the data set based on AI technology, and divide the data set into multiple data blocks according to the classification result, and each data block is stored in a different storage node respectively;
[0007] Set a high access volume threshold, obtain the total access volume of the stored data in each data block within a preset time, and when the total access volume of the corresponding data block is greater than or equal to the high access volume threshold, mark the stored data in this data block as hot storage data;
[0008] Establish a hot data storage node, migrate the data blocks storing the hot storage data to the hot data storage node, mark the data blocks migrated to the hot data storage node as hot data blocks, and mark the remaining data blocks not in the hot data storage node as cold data blocks;
[0009] Set up a data migration buffer in the hot data storage node. When the total access volume of the data stored in the hot data block within a preset time is less than the high access volume threshold, determine whether to migrate its stored data to the data migration buffer in the hot data storage node. Implement an access volume monitoring strategy for the data blocks in the data migration buffer to identify the warm data blocks as hot data blocks or cold data blocks after the preset time.
[0010] In some embodiments, when the total access volume of the data stored in the hot data block within a preset time is less than the high access volume threshold, obtain the access change volume of the data stored in the hot data block within the preset time, set a change volume threshold, and compare the access change volume with the change volume threshold. When the access change volume of the data stored in the hot data block within the preset time is less than the change volume threshold, migrate the hot data block back to the original storage node and mark it as a cold data block; when the access change volume of the data stored in the hot data block within the preset time is greater than or equal to the change volume threshold, mark this data block as a warm data block and migrate its stored data to the data migration buffer in the hot data storage node.
[0011] In some embodiments, the specific method of obtaining the access change volume of the data stored in the hot data block within a preset time is: obtain the access volume of this stored data within the preset time, mark it as the current access volume, and divide the current access volume by the total access volume when this data block was migrated to the hot data storage node to obtain the access change volume.
[0012] In some embodiments, the access volume monitoring strategy includes starting timing when the data block is migrated to the data migration buffer and ending the timing until the time reaches the preset time. Count the number of times the data stored in the warm data block is accessed and called during the timing period and mark it as the observed access volume. Compare the observed access volume with the high access volume threshold and make corresponding responses according to the comparison result.
[0013] In some embodiments, if the observed access volume is greater than or equal to the high access volume threshold, migrate the data block out of the data migration buffer and re-mark it as a hot data block; if the observed access volume is less than the high access volume threshold, migrate the data block from the data migration buffer of the hot data storage node to the original storage node and mark it as a cold data block.
[0014] In some embodiments, when the observed access volume is less than the high access volume threshold and the stored data access volume of the data blocks in the data migration buffer is not zero, it is determined whether the stored data called in the data block is isolated data. When the called stored data is isolated data, the stored data in this data block is split. The called stored data is split into hot stored data and migrated to the hot data storage node, the unused stored data is split into cold stored data and migrated to the original storage node, and this data block is marked as a split data block.
[0015] In some embodiments, the specific method for determining whether the stored data called in the data block is isolated data is as follows: a minimum access volume threshold is set, the observed access volume of the warm data block is compared with the minimum access volume threshold. If the observed access volume is less than the minimum access volume threshold, the operation of migrating this data block to the original storage node and marking it as a cold data block is maintained; if the observed access volume is greater than or equal to the minimum access volume threshold, the proportion of the called stored data in the entire data block in this data block is further obtained, and this proportion is marked as the call proportion. A proportion threshold is set. When the call proportion is less than the proportion threshold, it is determined that the called stored data in this data block is isolated data.
[0016] In some embodiments, after the data block is marked as a split data block, when the access volume of the hot stored data stored in the hot data storage node within a preset time is lower than the minimum access volume threshold, the split data block is merged, so that all the stored data of this data block is migrated to the original storage node, and this data block is marked as a cold data block; when the sum of the access volumes of the stored data stored in the original storage node and the hot data storage node within a preset time reaches the high access volume threshold, the split data block is merged, so that all the stored data of this data block is migrated to the hot data storage node, and this data block is marked as a hot data block.
[0017] The present invention also provides the following technical solution: A system of an AI-driven real-time analysis database based on distributed storage, including:
[0018] A data sharding module, which includes obtaining a data set of the analysis database, logically classifying the data in the data set based on AI technology, and splitting the data set into multiple data blocks according to the classification results, and each data block is stored in different storage nodes respectively;
[0019] A data recognition module, which includes setting a high access volume threshold, obtaining the total access volume of the stored data in each data block within a preset time, and when the total access volume of the corresponding data block is greater than or equal to the high access volume threshold, marking the stored data in this data block as hot stored data;
[0020] A data migration module, which includes establishing a hot data storage node, migrating data blocks storing hot storage data to the hot data storage node, marking the data blocks migrated to the hot data storage node as hot data blocks, and marking the remaining data blocks not in the hot data storage node as cold data blocks;
[0021] A data buffer module, which includes setting up a data migration buffer area in the hot data storage node. When the total access volume of the data stored in the hot data block within a preset time is less than the high access volume threshold, it is judged whether to migrate its stored data to the data migration buffer area in the hot data storage node. An access volume monitoring strategy is executed for the data blocks in the data migration buffer area to be used to mark warm data blocks as hot data blocks or cold data blocks after a preset time.
[0022] The present invention further provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the above method of an AI-driven real-time analysis database based on distributed storage.
[0023] The technical solution provided by the present invention has the following beneficial effects compared with the prior art:
[0024] First, the present invention logically classifies a data set through AI technology, cuts similar or related data into the same data block, and uses a specially configured hot data storage node in the distributed storage architecture to monitor the real-time access volume of the data block, and can timely identify the data blocks that will be frequently called, and migrate the frequently called data blocks to the hot data storage node to improve the query response speed and data processing performance of the actually used data, and meet the requirements of high concurrency and low latency of the real-time analysis system.
[0025] Second, by setting the change amount threshold, the present invention can more accurately distinguish between persistent hot data and short-term fluctuating data, and only trigger the migration of data blocks from the high-cost hot data storage node when the data access significantly decreases, avoiding the premature migration of originally normally accessed data back to the cold node due to data fluctuations, thereby reducing the frequent data migration caused by periodic fluctuations, and reducing the overall system operation and maintenance costs and resource waste.
[0026] Third, through the judgment and splitting of isolated data, the present invention can separately manage hot data and cold data, so that only the truly frequently accessed data (hot storage data) is migrated to the high-performance hot data storage node, while the data that has not been called or has a low call frequency remains on the original storage node. The splitting avoids unnecessary frequent migrations of the entire data block due to a small number of hot data, making the system more stable in the face of instantaneous access fluctuations, and not causing frequent switching of data between hot and cold nodes due to misjudgment, reducing the risk of system scheduling jitter. Description of the Drawings
[0027] Figure 1 This is a schematic flowchart of a method for an AI-driven real-time analysis database based on distributed storage according to the present invention;
[0028] Figure 2 This is a schematic diagram of the modules of a system for an AI-driven real-time analysis database based on distributed storage according to the present invention. Detailed implementation manners
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0030] It can be understood that the term "one" should be understood as "at least one" or "one or more". That is, in one embodiment, the number of one element can be one, and in other embodiments, the number of this element can be multiple. The term "one" cannot be understood as a limitation on the quantity.
[0031] The present invention provides a method for an AI-driven real-time analysis database based on distributed storage, as Figure 1 shown. The method includes the following steps:
[0032] Step 1: Obtain the data set of the analysis database, logically classify the data in the data set based on AI technologies (such as clustering algorithms, neural network models) and assign corresponding labels. For example, according to different data types, application scenarios, query patterns, etc., and according to the AI classification results, divide the data set into multiple data blocks, and each data block is stored in different storage nodes respectively. Each data block should contain data with the same or similar classification labels;
[0033] Step 2: Set a high access volume threshold, obtain the total access volume of the stored data in each data block within a preset time, and compare the total access volume of the stored data in the data block with the high access volume threshold. When the total access volume is less than the high access volume threshold, it means that the stored data in this data block has not been called multiple times during the real-time analysis of the database, so keep the storage state of this stored data in the corresponding data block without additional processing; when the total access volume is greater than or equal to the high access volume threshold, it means that the stored data in this data block has been called multiple times during the real-time analysis of the database, which means that the current real-time analysis of the database requires a large amount of the stored data in this data block, so mark the stored data in this data block as hot storage data;
[0034] Step 3: In the overall distributed storage architecture, set up a part of the nodes as hot data storage nodes. The hot data storage nodes adopt high-cost hardware configurations with high access concurrency and high I / O performance, as well as in-memory computing solutions, and pre-configure dedicated buffers and optimized storage algorithms on these nodes to support fast read and write operations for frequently accessed data. Migrate the data blocks storing hot storage data to the hot data storage nodes, mark the data blocks migrated to the hot data storage nodes as hot data blocks, and mark the remaining data blocks not in the hot data storage nodes as cold data blocks;
[0035] Step 4: Set up a data migration buffer in the hot data storage nodes. When the total access volume of the data stored in the hot data block within the preset time is less than the high access volume threshold, obtain the access change volume of the data stored in the hot data block within the preset time, and set a change volume threshold. Compare the access change volume with the change volume threshold. When the access change volume of the data stored in the hot data block within the preset time is less than the change volume threshold, it indicates that the access volume of the data stored in this data block has significantly decayed after being migrated to the hot data storage node. Then, migrate the hot data block back to the original storage node and mark it as a cold data block; when the access change volume of the data stored in the hot data block within the preset time is greater than or equal to the change volume threshold, it indicates that although the access volume of the data stored in this data block has decreased after being migrated to the hot data storage node, the decay amplitude is small. Then, mark this data block as a warm data block, migrate its stored data to the data migration buffer in the hot data storage node, and execute an access volume monitoring strategy for the data blocks in the data migration buffer to be used to mark the warm data block as a hot data block or a cold data block after the preset time.
[0036] The specific method for obtaining the access change amount of the stored data in the hot data block within the preset time is as follows: Obtain the access amount of this stored data within the preset time, mark it as the current access amount, and divide the current access amount by the total access amount when this data block is migrated to the hot data storage node to obtain the access change amount. For example, set the preset time to 3 minutes and the high access amount threshold to 100 times. The stored data in a data block sliced from the data set of the analysis database is accessed 150 times within 3 minutes, that is, the total access amount is 150 times. Then, this data block is migrated to the hot data storage node and marked as a hot data block. In the next 3 minutes (reaching the preset time again), the total access amount of this hot data block is 90 times. Since the total access amount at this time is less than the high access amount threshold, the obtained access change amount is 90÷150 = 0.6. Set the change amount threshold to 0.5. Since the access change amount is greater than the change amount threshold, this data block is marked as a warm data block, and its stored data is migrated to the data migration buffer area in the hot data storage node. This is because data access often has volatility, and certain business scenario requirements have inherent periodic fluctuations. Even if there is a slight decrease in a short period, it does not fully represent that the data has cooled down. This may be because the access requirements of the business within a certain time window have natural fluctuations, and there is only a slight decrease in the short term. If the hot data block is directly marked as a cold data block and migrated out of the hot data storage node at this time, it is easy to cause frequent data block migration operations, resulting in additional I / O and network loads, consuming additional computing resources, wasting system resources, and increasing the system burden.
[0037] The access amount monitoring strategy includes starting timing when the data block is migrated to the data migration buffer and ending timing until the time reaches the preset time. Count the number of times the stored data in the warm data block is accessed and called during the timing period and mark it as the observed access amount. Compare the size of the observed access amount with the high access amount threshold. If the observed access amount is greater than or equal to the high access amount threshold, it means that after the data block is migrated to the data migration buffer, the access amount of its stored data has returned to the qualified state within the preset time, indicating that the previous decrease in the access amount belongs to a normal fluctuation phenomenon. Then, this data block is migrated out of the data migration buffer and re-marked as a hot data block. If the observed access amount is less than the high access amount threshold, it means that after the data block is migrated to the data migration buffer, the access amount of its stored data still fails to meet the standard of the hot data block, indicating that the stored data of this data block has cooled down, that is, the call frequency has decreased. Then, this data block is migrated from the data migration buffer of the hot data storage node to the original storage node and marked as a cold data block.
[0038] After executing the above method, when the observed access volume is less than the high access volume threshold and the stored data access volume of the data block in the data migration buffer is not zero, it should also be determined whether the stored data called in this data block is isolated data. When the called stored data is isolated data, the stored data in this data block is split. The called stored data is split into hot stored data and migrated to the hot data storage node, the uncalled stored data is split into cold stored data and migrated to the original storage node, and this data block is marked as a split data block.
[0039] The specific method for determining whether the stored data called in the data block is isolated data is as follows: Set a minimum access volume threshold, which can be 20%-50% of the high access volume threshold. Compare the observed access volume of the warm data block with the minimum access volume threshold. If the observed access volume is less than the minimum access volume threshold, it indicates that the call volume of the stored data in this data block is extremely small. Then maintain the operation of migrating this data block to the original storage node and marking it as a cold data block. If the observed access volume is greater than or equal to the minimum access volume threshold, further obtain the proportion of the stored data called in this data block in the entire data block, and mark this proportion as the call proportion. Set a proportion threshold, such as 0.1-0.2. If the call proportion is greater than or equal to the proportion threshold, it means that the analysis database calls a relatively large proportion of the stored data in this module. Then maintain the operation of migrating this data block to the original storage node and marking it as a cold data block. If the call proportion is less than the proportion threshold, it means that the analysis database only calls a small amount of stored data in this module. Since the observed access volume is greater than or equal to the minimum access volume threshold, it indicates that a small part of the stored data called in this data block is called repeatedly in large quantities. Then determine that the stored data called in this data block is isolated data. The data within a data block usually should have similar access characteristics and usage patterns due to the same or similar categories and labels. Therefore, when the real-time analysis database is used, the entire data block needs to be migrated to the hot data storage node for use. However, due to differences in business requirements and the refinement of call query methods, it may occur that some data items are called repeatedly frequently, while most of the other data in the same data block is rarely accessed. Then define this part of the data that is called in large quantities but accounts for a small proportion as isolated data. For a data block with isolated data, if the entire data block is migrated to the hot data storage node, but in fact only a small part of the data is called frequently, it will lead to a waste of high-performance resources. After splitting, migrating the hot part (hot stored data) to the high-cost hot data storage node can, while ensuring high access performance, keep the cold data on the original storage node, thus saving more resources. Taking the above embodiments as an example, when the high access volume threshold is set to 100 times, the minimum access volume threshold is 30% of the high access volume threshold, that is, the minimum access volume threshold is 30 times. Set the observed access volume of the warm data block to 40 times. Since the observed access volume is greater than the minimum access volume threshold, obtain the proportion of the stored data called in this data block in the entire data block. Suppose there are 1000 stored data in this data block. Among the observed access volume, 20 data are called and accessed 2 times, that is, there are 20 stored data that have been called. The call proportion is 2%. Set the proportion threshold to 5%. Then it can be determined that there is isolated data in this data block.
[0040] After the data block is marked as a split data block, a part of its stored data is stored in the hot data storage node, and the other part of the stored data is stored in the original storage node. When the access volume of the hot storage data stored in the hot data storage node is lower than the minimum access volume threshold within the preset time, the split data block is merged, so that all the stored data of this data block is migrated to the original storage node, and this data block is marked as a cold data block; when the sum of the access volumes of the stored data stored in the original storage node and the hot data storage node reaches the high access volume threshold within the preset time, the split data block is merged, so that all the stored data of this data block is migrated to the hot data storage node, and this data block is marked as a hot data block.
[0041] Generally speaking, the present invention aims to design a method for an AI-driven real-time analysis database based on distributed storage. Aiming at the problem that some node storage loads are overweight in the case of the real-time analysis database processing a large amount of data, the present invention logically classifies the data set through AI technology, and slices similar or related data into the same data block. By using the specially configured hot data storage node in the distributed storage architecture to monitor the real-time access volume of the data block, data blocks that will be frequently called can be identified in time, and the frequently called data blocks are migrated to the hot data storage node to improve the query response speed and data processing performance of the actually used data, meet the requirements of high concurrency and low latency of the real-time analysis system, and according to the data access conditions in different stages, the system can accurately judge and manage the heat, warm, and cold of the data block, forming a set of adaptive and intelligent data scheduling mechanisms. By setting the change amount threshold, continuous hot data and short-term fluctuating data can be more accurately distinguished. Only when the data access significantly decreases is the data block triggered to move out of the high-cost hot data storage node, avoiding the premature migration of originally normally accessed data back to the cold node due to data fluctuations, thereby reducing the frequent data migration caused by periodic fluctuations, and reducing the overall system operation and maintenance costs and resource waste; by judging and splitting the isolated data, the hot data and the cold data can be managed separately, so that only the truly frequently accessed data (hot storage data) is migrated to the high-performance hot data storage node, while the data that has not been called or has a low call frequency remains on the original storage node. The splitting avoids unnecessary frequent migrations of the entire data block due to a small number of hot data, making the system more stable in the face of instantaneous access fluctuations and not switching frequently between hot and cold nodes due to misjudgment, reducing the risk of scheduling jitter of the system.
[0042] The present invention provides a system for an AI-driven real-time analysis database based on distributed storage, as Figure 2 shown, including:
[0043] A data sharding module, which includes obtaining a data set from an analysis database, logically classifying the data in the data set based on AI technology, and splitting the data set into multiple data blocks according to the classification results, with each data block stored in different storage nodes respectively;
[0044] A data identification module, which includes setting a high access volume threshold, obtaining the total access volume of the stored data in each data block within a preset time, and marking the stored data in this data block as hot stored data when the total access volume of the corresponding data block is greater than or equal to the high access volume threshold;
[0045] A data migration module, which includes establishing a hot data storage node, migrating the data blocks storing hot stored data to the hot data storage node, marking the data blocks migrated to the hot data storage node as hot data blocks, and marking the remaining data blocks not in the hot data storage node as cold data blocks;
[0046] A data buffer module, which includes setting up a data migration buffer area in the hot data storage node. When the total access volume of the stored data in the hot data block within a preset time is less than the high access volume threshold, it is judged whether to migrate its stored data to the data migration buffer area in the hot data storage node, and an access volume monitoring strategy is executed for the data blocks in the data migration buffer area to be used to mark the warm data blocks as hot data blocks or cold data blocks after a preset time.
[0047] Embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. Embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or installed from a removable medium. When the computer program is executed by the central processing unit, the above functions defined in the methods of the present application are executed. It should be noted that the computer-readable medium described above in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wire segments, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. And in the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program codes. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program codes contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless segments, wire segments, optical cables, RF, etc., or any suitable combination of the above.
[0048] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0049] Those skilled in the art should understand that the above description is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application.
Claims
1. A method for an AI-driven real-time analysis database based on distributed storage, characterized in that, The method includes the following steps: Obtain the data set of the analysis database, perform logical classification on the data in the data set based on AI technology, and split the data set into multiple data blocks according to the classification results. Each data block is stored in different storage nodes respectively; Set a high access volume threshold, obtain the total access volume of the stored data in each data block within a preset time, and when the total access volume of the corresponding data block is greater than or equal to the high access volume threshold, mark the stored data in this data block as hot stored data; Establish a hot data storage node, migrate the data block storing the hot stored data to the hot data storage node, mark the data block migrated to the hot data storage node as a hot data block, and mark the remaining data blocks not in the hot data storage node as cold data blocks; Set up a data migration buffer area in the hot data storage node. When the total access volume of the stored data in the hot data block within a preset time is less than the high access volume threshold, determine whether to migrate its stored data to the data migration buffer area in the hot data storage node. Implement an access volume monitoring strategy for the data blocks in the data migration buffer area to be used to mark the warm data block as a hot data block or a cold data block after a preset time.
2. The method of an AI-driven real-time analysis database based on distributed storage according to claim 1, wherein When the total access volume of the stored data in the hot data block within a preset time is less than the high access volume threshold, obtain the access change volume of the stored data in the hot data block within a preset time, set a change volume threshold, compare the access change volume with the change volume threshold. When the access change volume of the stored data in the hot data block within a preset time is less than the change volume threshold, migrate the hot data block back to the original storage node and mark it as a cold data block; When the access change volume of the stored data in the hot data block within a preset time is greater than or equal to the change volume threshold, mark this data block as a warm data block and migrate its stored data to the data migration buffer area in the hot data storage node.
3. A method for an AI-driven real-time analysis database based on distributed storage according to claim 2, characterized in that, The specific way to obtain the access change volume of the stored data in the hot data block within a preset time is: obtain the access volume of this stored data within a preset time, mark it as the current access volume, and divide the current access volume by the total access volume when this data block is migrated to the hot data storage node to obtain the access change volume.
4. A method for an AI-driven real-time analysis database based on distributed storage according to claim 3, wherein, The access volume monitoring strategy includes starting timing when the data block is migrated to the data migration buffer area and ending the timing until the time reaches the preset time. Count the number of times the stored data in the warm data block is accessed and called during the timing period and mark it as the observed access volume, compare the observed access volume with the high access volume threshold, and make corresponding responses according to the comparison results.
5. A method for an AI-driven real-time analysis database based on distributed storage according to claim 4, characterized in that, If the observed access volume is greater than or equal to the high access volume threshold, move the data block out of the data migration buffer area and re-mark it as a hot data block; if the observed access volume is less than the high access volume threshold, move the data block from the data migration buffer area of the hot data storage node to the original storage node and mark it as a cold data block.
6. A method for an AI-driven real-time analysis database based on distributed storage according to claim 5, characterized in that, When the observed access volume is less than the high access volume threshold and the stored data access volume of the data block in the data migration buffer is not 0, determine whether the stored data called in this data block is orphan data. When the called stored data is orphan data, split the stored data in this data block, split the called stored data into hot stored data and migrate it to the hot data storage node, split the uncalled stored data into cold stored data and migrate it to the original storage node, and mark this data block as a split data block.
7. A method for an AI-driven real-time analysis database based on distributed storage according to claim 6, characterized in that, The specific method for determining whether the stored data called in this data block is orphan data is as follows: Set the minimum access volume threshold, compare the observed access volume of the warm data block with the minimum access volume threshold. If the observed access volume is less than the minimum access volume threshold, maintain the operation of migrating this data block to the original storage node and mark it as a cold data block; if the observed access volume is greater than or equal to the minimum access volume threshold, further obtain the proportion of the called stored data in this data block occupying the entire data block, mark the proportion as the call ratio, set the ratio threshold, and when the call ratio is less than the ratio threshold, determine that the called stored data in this data block is orphan data.
8. A method for an AI-driven real-time analysis database based on distributed storage according to claim 7, characterized in that, After the data block is marked as a split data block, when the access volume of the hot stored data stored in the hot data storage node within the preset time is lower than the minimum access volume threshold, merge the split data block so that all the stored data of this data block is migrated to the original storage node, and mark this data block as a cold data block; When the sum of the access volumes of the stored data stored in the original storage node and the hot data storage node within the preset time reaches the high access volume threshold, merge the split data block so that all the stored data of this data block is migrated to the hot data storage node, and mark this data block as a hot data block.
9. A system of an AI-driven real-time analysis database based on distributed storage, characterized in that, The method for an AI-driven real-time analysis database based on distributed storage according to any one of claims 1-8, includes: A data sharding module, which includes obtaining a data set of the analysis database, logically classifying the data in the data set based on AI technology, and splitting the data set into multiple data blocks according to the classification results, and each data block is stored in different storage nodes respectively; A data identification module, which includes setting a high access volume threshold, obtaining the total access volume of the stored data in each data block within the preset time, and when the total access volume of the corresponding data block is greater than or equal to the high access volume threshold, marking the stored data in this data block as hot stored data; A data migration module, which includes establishing a hot data storage node, migrating the data block storing the hot stored data to the hot data storage node, marking the data block migrated to the hot data storage node as a hot data block, and marking the remaining data blocks not in the hot data storage node as cold data blocks; A data buffer module, which includes establishing a data migration buffer in a hot data storage node. When the total access volume of the data stored in a hot data block within a preset time is less than a high access volume threshold, it is determined whether to migrate its stored data to the data migration buffer in the hot data storage node. An access volume monitoring policy is executed for the data blocks in the data migration buffer to be used to identify warm data blocks as hot data blocks or cold data blocks after a preset time.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method of an AI-driven real-time analysis database based on distributed storage according to any one of claims 1-8 above.
Citation Information
Patent Citations
Optomagnetoelectric fusion media asset data distributed storage and management method and device
CN115883590A
Full-link tracking isolated tree identification and alarm method and device, equipment and medium
CN118551216A
Data storage method and device and computing equipment
CN118708561A
Cited By
A method and apparatus for sharding a distributed database
CN122692133A