An archive management system based on a distributed data archive

CN122777481APending Publication Date: 2026-09-18SHANDONG RUITU INTELLIGENT TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611243432.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-17
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0005]为此,本发明提供一种基于分布式数据档案库的档案管理系统,用以克服现有技术中分布式存储难以动态调整档案的存储位置与存储方式,进而造成档案提取效率低、占用内存较大的问题

Benefits of technology

[0033] Compared with existing technologies, the beneficial effects of this invention are as follows: The archive management system based on a distributed data archive provided by this invention solves the core pain points of current distributed storage, such as the difficulty in dynamically adjusting the location and method of archive storage, resulting in low archive retrieval efficiency and large memory consumption, through multi-module collaborative design. This provides effective support for the optimization and upgrading of archive management: This system, with the help of the analysis module, determines the search representation status by combining the number of searches, retrieval, and read/write attributes of the archive within the analysis period. At the same time, it associates storage nodes and search nodes to determine the storage validity and virtual/real positions, realizing accurate judgment of archive storage needs. The management module dynamically formulates and executes storage adjustment measures based on the judgment results, breaking the limitations of the static allocation of traditional distributed storage, allowing the archive storage location and method to be flexibly optimized to adapt to the actual usage frequency. With the real-time monitoring of search nodes, storage nodes, and search attributes by the search module, as well as the global synchronization of the lookup table, the system can quickly locate the target archive and remotely retrieve data by selecting the calling node with the lowest CPU utilization, greatly improving retrieval efficiency. At the same time, dynamic adjustment avoids archives occupying redundant node resources for a long time, effectively reducing memory consumption, and taking into account the massive expansion advantages of distributed storage and the flexible and efficient needs of archive management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122777481A_ABST
    Figure CN122777481A_ABST
Patent Text Reader

Abstract

This invention relates to the field of archival management technology, and more particularly to an archival management system based on a distributed data archive. The system comprises distributed storage nodes and a data management center. The data management center includes an analysis module and a management module. The analysis module determines the search representation status, storage node validity, and search real and virtual positions of each archive based on the number of searches and search attributes, thereby determining whether to adjust the storage measures for the corresponding archive. The management module determines the storage adjustment measures for the corresponding archive based on the reasons for adjustment, and generates a table showing the correspondence between each archive and its storage node, storing this table in the entire data archive. This invention enables dynamic storage adaptation through collaboration among the modules of the data management center, solving the problems of insufficient adjustment in traditional distributed storage, improving archive retrieval efficiency, optimizing memory usage, ensuring data security, and strongly supporting the digital upgrade of archival management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of archives management technology, and in particular to an archives management system based on a distributed data archive. Background Technology

[0002] Distributed file systems can support the decentralized storage and collaborative access of massive amounts of paper scans, electronic documents, and other archives, solving the expansion bottleneck of traditional centralized storage. Data sharding technology splits large-capacity archives and distributes them across different nodes, improving storage parallelism. Redundant backup technology uses multiple data copies on multiple nodes to avoid the risk of archive loss due to single-point failures, meeting the needs of long-term archive preservation. Metadata management technology records information such as archive indexes and storage nodes to assist in quickly locating archives. Meanwhile, distributed database technology is adapted to structured archive data management, ensuring efficient retrieval of related information. These technologies collectively solve the problems of limited storage capacity, low security, and restricted access in traditional archives, becoming the core support solution for digital archive management. However, in current applications of these technologies, distributed storage struggles to dynamically adjust the storage nodes and storage methods of archives, resulting in low archive retrieval efficiency and high memory consumption, hindering the optimization and upgrading of archive management.

[0003] Chinese Patent Publication No. CN119669150B discloses an archive management system based on distributed storage. This invention matches the archive types of pre-stored video archives in a smart campus with corresponding storage nodes, obtains archive features of the pre-stored video archives based on image feature recognition algorithms, and uses archive continuity similarity and video archive feature data length as the criteria for determining distributed storage segmentation points to facilitate subsequent distributed storage. By traversing the similarity of the segmented pre-stored video archives, the features of the first and second keyframe video archives are obtained as the basis for data backup and restoration, and distributed storage is performed based on archive type. This method improves the security of smart campus video archives, and the feature selection based on image recognition algorithms and distributed storage ensures the integrity of video backup and restoration. Therefore, the existing technology has the following problems:

[0004] Distributed storage makes it difficult to dynamically adjust the storage location and method of archives, resulting in low archive retrieval efficiency and large memory consumption, which restricts the optimization and upgrading of archive management. Summary of the Invention

[0005] To address this, the present invention provides an archive management system based on a distributed data archive, which overcomes the problem in the prior art that distributed storage is difficult to dynamically adjust the storage location and storage method of archives, resulting in low archive retrieval efficiency and large memory consumption.

[0006] To achieve the above objectives, the present invention provides an archive management system based on a distributed data archive, comprising several distributed storage nodes and a data management center, wherein each of the storage nodes is communicatively connected to the data management center, and the data management center further comprises:

[0007] The sorting module is used to match several files in the database based on search keywords and sort the files in real time according to the actual click rate of each file.

[0008] The search module obtains the search node and storage node of each file, and monitors the search attributes of each file.

[0009] The search attributes include retrieval attributes and read / write attributes;

[0010] An analysis module, connected to the search module, is used to determine the search representation status of each file based on the number of searches and search attributes within the analysis period, determine the validity of each storage node of the corresponding file and its search real position and search virtual position based on the storage node and search node, and determine whether to adjust the storage measures of the corresponding file based on the determination result of the search representation status combined with the search virtual position and the validity of the storage node.

[0011] The management module, which is connected to the analysis module, determines the storage adjustment measures for the corresponding archives based on the judgment result of adjusting the corresponding archives and the reasons for the adjustment, and generates a table of correspondence between each archive and its storage node and stores the table of correspondence in the entire data archive database.

[0012] As a preferred technical solution for an archive management system based on a distributed data archive, the sorting module sorts the matched archives in real time according to the actual click rate of each archive, from largest to smallest.

[0013] As a preferred technical solution for an archive management system based on a distributed database, the analysis module acquires all archives within the analysis period and determines the search count and search attributes of each archive. Based on the search count, it determines whether to combine the search attributes to determine the search representation status, including:

[0014] Based on the determination result that the number of searches is greater than or equal to the preset number, the search representation status of the corresponding file is determined to be an explicit frequent state;

[0015] Based on the determination that the number of searches is less than the preset number, the search representation status of the corresponding file is determined by combining the search attributes of each search.

[0016] As a preferred technical solution for an archive management system based on a distributed database, the analysis module determines the proportion of read / write attributes based on the ratio of the number of times the search attribute is read / write attribute to the total number of searches, and determines the search representation status of the corresponding archive based on the read / write attribute proportion, including:

[0017] Based on the determination result that the proportion of read and write attributes is greater than or equal to the preset proportion, the search representation state is determined to be a latent frequent state.

[0018] Based on the determination result that the proportion of read and write attributes is less than the preset proportion, the search representation state is determined to be an infrequent state.

[0019] As a preferred technical solution for an archive management system based on a distributed data archive, the analysis module determines the nodes in the search nodes that are the same as the storage nodes of the corresponding archive as the real search nodes of the archive, and determines the nodes in the search nodes that are different from the storage nodes of the corresponding archive as the virtual search nodes of the archive.

[0020] As a preferred technical solution for an archive management system based on a distributed database, the analysis module determines the validity of a storage node based on whether the storage node of the archive is also a search node, wherein:

[0021] In response to the fact that it is also a search node, the storage node is determined to be a valid storage node;

[0022] If a storage node is not simultaneously being searched, it is determined to be an invalid storage node.

[0023] As a preferred technical solution for an archive management system based on a distributed data archive, the analysis module determines whether to adjust the corresponding archive based on the judgment result of the search representation status, combined with the search vacancy and the validity of the storage node, including:

[0024] Based on the determination of infrequent status, the storage measures for the corresponding files will not be adjusted.

[0025] Based on the results of the frequent status determinations, the storage measures for the corresponding files are adjusted, among which,

[0026] In response to the explicit frequent state, the storage measures of the corresponding file are adjusted based on the determination results of the existence of invalid storage nodes or search vacancy.

[0027] In response to the latent frequent state, the storage measures for the corresponding files are adjusted based on the determination results of the simultaneous existence of invalid storage nodes and search vacant positions.

[0028] As a preferred technical solution for an archive management system based on a distributed data repository, the management module determines the storage adjustment measures for the corresponding archives based on the judgment result of adjusting the corresponding archives and the reason for the adjustment, including:

[0029] In response to the reason for the adjustment being the existence of invalid storage nodes, the storage adjustment measure is determined to be to compress the files of the invalid storage nodes and store the compressed files as archives.

[0030] In response to the reason for the adjustment being the existence of search voids, storage adjustment measures were formulated based on the void search ratio of each search void.

[0031] As a preferred technical solution for an archive management system based on a distributed data archive, the management module stores the archive in search vacancy spaces where the proportion of vacancy searches is greater than or equal to the proportion threshold, and stores the archive's compressed file in search vacancy spaces where the proportion of vacancy searches is less than the proportion threshold.

[0032] As a preferred technical solution for an archive management system based on a distributed data archive, the management module decompresses the compressed package when the corresponding file is retrieved from the storage node where the compressed package is located, or deletes the compressed package when the duration of existence of a single compressed package exceeds a preset duration.

[0033] Compared with existing technologies, the beneficial effects of this invention are as follows: The archive management system based on a distributed data archive provided by this invention solves the core pain points of current distributed storage, such as the difficulty in dynamically adjusting the location and method of archive storage, resulting in low archive retrieval efficiency and large memory consumption, through multi-module collaborative design. This provides effective support for the optimization and upgrading of archive management: This system, with the help of the analysis module, determines the search representation status by combining the number of searches, retrieval, and read / write attributes of the archive within the analysis period. At the same time, it associates storage nodes and search nodes to determine the storage validity and virtual / real positions, realizing accurate judgment of archive storage needs. The management module dynamically formulates and executes storage adjustment measures based on the judgment results, breaking the limitations of the static allocation of traditional distributed storage, allowing the archive storage location and method to be flexibly optimized to adapt to the actual usage frequency. With the real-time monitoring of search nodes, storage nodes, and search attributes by the search module, as well as the global synchronization of the lookup table, the system can quickly locate the target archive and remotely retrieve data by selecting the calling node with the lowest CPU utilization, greatly improving retrieval efficiency. At the same time, dynamic adjustment avoids archives occupying redundant node resources for a long time, effectively reducing memory consumption, and taking into account the massive expansion advantages of distributed storage and the flexible and efficient needs of archive management.

[0034] In particular, the analysis module provides precise decision support for the dynamic storage adjustment of the file management system through scientific judgment logic and parameter settings, effectively making up for the shortcomings of the crude judgment of traditional distributed storage. The analysis module takes the number of searches and the ratio of read and write attributes as the core dimensions, combined with the preset number of searches and reasonable preset ratios that are positively correlated with the analysis cycle, to accurately classify three characteristic states: explicit frequent, implicit frequent, and infrequent. It not only identifies files with high access frequency, but also captures the file needs with low access frequency but high modification frequency, avoiding the one-sidedness of judging solely by the number of searches. Moreover, the hierarchical judgment logic makes storage adjustment more targeted, providing a reliable basis for the management module to dynamically optimize storage location and method, helping the system solve the core problems of low retrieval efficiency and high memory consumption, while adapting to diverse scenarios of file access and modification, ensuring the efficient utilization of storage resources in the distributed data archive.

[0035] In particular, the analysis module provides precise node-level decision-making basis for the dynamic storage optimization of the archive management system by clarifying the definition criteria of real and virtual search positions and the logic for determining the validity of storage nodes. This further strengthens the system's core capability to solve problems such as insufficient dynamic adjustment of distributed storage, low retrieval efficiency, and high memory consumption. The analysis module uses the consistency between storage nodes and search nodes as the core judgment dimension, clearly distinguishing between real search positions (node ​​overlap) and virtual search positions (node ​​non-overlap), and accurately identifying effective storage nodes (which are also search nodes) and invalid storage nodes (which are only storage nodes). It intuitively presents the matching situation between archive storage nodes and actual search needs. This precise division allows the management module to optimize in a targeted manner, that is, retaining effective nodes to ensure access efficiency, cleaning up invalid nodes to release redundant memory, and supplementing storage in virtual nodes to meet cross-node needs. This makes the storage layout more in line with the actual use scenario, providing solid node analysis support for the system to dynamically adjust storage locations, improve retrieval efficiency, and reduce memory consumption.

[0036] In particular, the management module, through differentiated and refined storage adjustment strategies, accurately implements the decisions of the analysis module, effectively addressing the core pain points of insufficient dynamic adjustment of distributed storage, low retrieval efficiency, and high memory consumption, providing key support for the optimization and upgrading of the document management system. The management module formulates adaptation measures for different adjustment reasons, including: using compressed storage for invalid storage nodes to reduce redundant memory consumption; deploying original data blocks or compressed packages differently for search virtual slots based on the virtual slot search ratio, balancing access efficiency and resource conservation; and achieving a balance between redundancy cleanup and data security through intelligent deletion rules for compressed packages and the constraint of retaining at least two compressed packages. These measures make storage adjustments more targeted, ensuring rapid retrieval of high-frequency access and high-demand nodes while releasing redundant resources through compression and intelligent deletion, perfectly meeting the core objectives of dynamic system adaptation and optimized resource allocation. Attached Figure Description

[0037] Figure 1 This is a connection diagram of an archive management system based on a distributed data archive according to an embodiment of the present invention;

[0038] Figure 2 This is a flowchart of the analysis module in an embodiment of the present invention. Detailed Implementation

[0039] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0040] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0041] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., indicating directions or node relationships are based on the directions or node relationships shown in the drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.

[0042] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0043] It should be understood that the core forms of storage nodes in distributed storage include physical servers, virtual nodes, and storage array nodes. Since the contents stored inside physical servers vary depending on their location, they should also be related to their physical location. This application considers the analysis of each storage node's extraction and storage of each stored content to dynamically adjust the storage location and storage method of the files, thereby improving the file extraction efficiency and reducing memory usage to achieve the purpose of optimizing and upgrading file management.

[0044] Please see Figure 1 The diagram shown is a connection diagram of an archive management system based on a distributed data archive according to an embodiment of the present invention. This embodiment of the present invention provides an archive management system based on a distributed data archive, comprising several distributed storage nodes and a data management center. Each storage node is communicatively connected to the data management center, which further includes:

[0045] The sorting module is used to match several files in the database based on search keywords and sort the files in real time according to the actual click rate of each file.

[0046] The search module is connected to the sorting module, and obtains the search node and storage node of each file, as well as monitors the search attributes of each file.

[0047] The search attributes include retrieval attributes and read / write attributes;

[0048] An analysis module, connected to the search module, is used to determine the search representation status of each file based on the number of searches and search attributes within the analysis period, determine the validity of each storage node of the corresponding file and its search real position and search virtual position based on the storage node and search node, and determine whether to adjust the storage measures of the corresponding file based on the determination result of the search representation status combined with the search virtual position and the validity of the storage node.

[0049] The management module, which is connected to the analysis module, determines the storage adjustment measures for the corresponding archives based on the judgment result of adjusting the corresponding archives and the reasons for the adjustment, and generates a table of correspondence between each archive and its storage node and stores the table of correspondence in the entire data archive database.

[0050] In practice, when a file 'a' is stored on a certain storage node, if the storage node (i.e., the search node) does not store the original data block or compressed package of 'a', then according to the lookup table, several storage nodes that store the original data block of 'a' are determined, and the storage node with the lowest CPU utilization is selected as the calling node. After communicating with the calling node, the original data block is opened remotely.

[0051] Understandably, this system, through its analysis module combining search frequency, access / read / write attributes, and other data, accurately determines the usage needs and storage status of files. The management module then dynamically adjusts the file storage location and method accordingly, changing the traditional distributed storage model of fixed allocation and fundamentally solving the core problem of storage's inability to dynamically adapt, thus prioritizing storage resources for frequently used files. The search module monitors search and storage node information in real time, using a relationship table to globally associate files with storage nodes, avoiding the inefficiency of location caused by distributed storage. Simultaneously, by selecting the node with the lowest CPU utilization for remote data retrieval, latency caused by node congestion is reduced, significantly improving performance. This system significantly improves the smoothness and speed of file retrieval; a dynamic adjustment mechanism removes low-frequency files from redundant storage nodes, retaining only the necessary storage format (i.e., compressed files), while high-frequency files are deployed on high-efficiency access nodes, avoiding the waste of memory usage by all files indiscriminately, achieving reasonable allocation of storage resources, and effectively reducing overall memory usage costs; the cluster architecture of distributed storage nodes, coupled with a globally synchronized lookup table, ensures that file data will not be lost due to single-point failures; at the same time, the division of labor and cooperation among modules takes into account both storage scalability and management flexibility, adapting to the long-term storage needs of massive files, and meeting the dynamic management needs of file retrieval and updates, thus contributing to the digital upgrade of file management.

[0052] Specifically, the sorting module sorts the matched files in real time according to their actual click-through rates from highest to lowest.

[0053] It is understandable that entering a search keyword may result in at least one corresponding file. Based on the position of the displayed file in adjacent sentences, the searcher will usually select the specific file they want to search for. Therefore, the sorting module will count the actual click rate of each keyword and its corresponding files to facilitate the sorting of files when entering the same keyword in the future.

[0054] Please see Figure 2 The diagram shown illustrates the workflow of the analysis module in this embodiment of the invention. Specifically, the analysis module acquires all files within the analysis period and determines the search count and search attributes of each file. Based on the search count, it determines whether to combine the search attributes to determine the search representation status, including:

[0055] Based on the determination that the number of searches is greater than or equal to the preset number, the search characteristic of the corresponding file is determined to be an explicit and frequent state. It should be understood that an explicit and frequent state indicates that the search volume for the file is relatively large (high search frequency) during this analysis period. Therefore, it is necessary to consider other data to determine whether the storage location of the file is appropriate and whether it needs to be adjusted to meet the current search situation.

[0056] Based on the determination that the number of searches is less than the preset number, the search representation status of the corresponding file is determined by combining the search attributes of each search.

[0057] Understandably, the preset number of searches is 30. When the number of searches exceeds 30, the search representation state is determined to be an explicit frequent state. The search representation state includes explicit frequent state, implicit frequent state, and infrequent state.

[0058] In practice, the analysis cycle is from one week to one month, preferably once every two weeks. The length of the analysis cycle is positively correlated with the number of preset times. That is, when the analysis cycle is two weeks, the preset number of times is 30. If the analysis cycle is extended, the preset number of times should be increased proportionally. If the analysis cycle is shortened, the preset number of times should be reduced proportionally.

[0059] Specifically, the analysis module determines the proportion of read / write attributes based on the ratio of the number of times the search attribute is read / write to the total number of searches, and determines the search representation status of the corresponding file based on the read / write attribute proportion, including:

[0060] Based on the determination result that the proportion of read and write attributes is greater than or equal to the preset proportion, the search representation state is determined to be a latent frequent state. It should be understood that the latent frequent state means that the number of searches for the file in this analysis period is relatively small, but it is usually accompanied by modifications to the file content. At this time, it is necessary to consider each file specifically to determine whether its storage needs to be adjusted and how to adjust it.

[0061] Based on the determination result that the proportion of read and write attributes is less than the preset proportion, the search representation state is determined to be an infrequent state. It should be understood that the infrequent state means that there are not many searches and modifications of files during this analysis period, and there is no need to adjust the storage of files at this time.

[0062] It should be understood that the total number of searches is the sum of the number of searches with the read / write attribute and the number of searches with the retrieval attribute.

[0063] In practice, the preset percentage is usually ≥50%. The higher the preset percentage, the more modifications are made to the files. Preferably, the preset percentage is 70%.

[0064] Understandably, this approach breaks away from the traditional limitation of judging file importance solely by search frequency. By combining search frequency with read / write attribute ratios, it identifies both frequently accessed (explicitly frequent) files and less frequently accessed (implicitly frequent) files with high modification needs. This comprehensively covers diverse file usage scenarios, ensuring that storage adjustments do not overlook critical requirements. Through clear characterization of file statuses, it provides the management module with a clear basis for adjustments: explicitly frequent files require optimized storage locations to improve access efficiency, implicitly frequent files require targeted adjustments to adapt to modification needs, and infrequent files maintain their current state to reduce redundant operations, ensuring the accuracy of storage adjustments from the source.

[0065] Specifically, the analysis module identifies the nodes in the search nodes that are the same as the storage nodes of the corresponding files as the real search nodes of the files, and identifies the nodes in the search nodes that are different from the storage nodes of the corresponding files as the virtual search nodes of the files.

[0066] In one implementation, the distributed data archive has nine storage nodes: x1, x2, ..., x9. File b has four storage nodes: x3, x4, x5, and x8. During the analysis period, there are five storage nodes (i.e., search nodes) that retrieve file b: x2, x4, x5, x7, and x8. Therefore, the real search nodes for file b are x4, x5, and x8, and the virtual search nodes for file b are x2 and x7.

[0067] Specifically, the analysis module determines the validity of a storage node based on whether the storage node of the file is also a search node, where:

[0068] In response to the fact that it is also a search node, the storage node is determined to be a valid storage node;

[0069] If a storage node is not simultaneously being searched, it is determined to be an invalid storage node.

[0070] In the above implementation, the valid storage nodes among x3, x4, x5 and x8 include x4, x5 and x8, and the invalid storage node is x3.

[0071] Understandably, by distinguishing between real and virtual search nodes, the compatibility between actual search nodes and storage nodes can be clearly understood. This clarifies which nodes are core areas for high-frequency searches (real nodes) and which nodes have search needs but no storage (virtual nodes), avoiding the resource waste caused by the blind deployment of traditional distributed storage and allowing storage adjustments to better meet actual usage needs. Accurately identifying invalid nodes that only perform storage functions but are not called by searches (i.e., the aforementioned x3 nodes) provides a clear basis for the management module to clean up such nodes, preventing invalid storage from occupying memory resources for extended periods and effectively reducing the memory consumption cost of distributed storage. Based on search... Identifying vacant storage locations provides the management module with decision-making direction for cross-node storage supplementation, reducing remote call latency caused by inconsistencies between search nodes and storage nodes during subsequent searches. Simultaneously, it enhances resource allocation for effective storage nodes, further improving the smoothness of file retrieval. By combining node-level supply and demand matching data (search real locations, search vacant locations, effective storage nodes, and invalid storage nodes) with previous search representations, the management module's storage adjustments not only adapt to file usage frequency but also align with node access needs. This forms a complete closed loop from demand determination to node optimization, helping the system overcome the static limitations of traditional distributed storage.

[0072] Specifically, the analysis module determines whether to adjust the corresponding file based on the search representation status judgment result, combined with its search vacancy and storage node validity, including:

[0073] Based on the determination of infrequent status, the storage measures for the corresponding files will not be adjusted.

[0074] Based on the results of the frequent status determinations, the storage measures for the corresponding files are adjusted, among which,

[0075] In response to the explicit frequent state, the storage measures for the corresponding files are adjusted based on the determination results of the existence of invalid storage nodes or search vacancy. It can be understood that the explicit frequent state means that the files are frequently accessed objects, and their storage layout will affect the overall retrieval efficiency and memory usage of the system. The existence of invalid storage nodes means that such nodes only occupy memory but are not called by search, which is a redundant resource consumption. In high-frequency scenarios, long-term memory occupation will exacerbate resource tension. The existence of search vacancy means that the search node and the storage node are inconsistent. High-frequency cross-node remote calls will accumulate latency and seriously affect retrieval efficiency. Therefore, in high-frequency usage scenarios, any resource mismatch (redundancy or efficiency risks) will be amplified and needs to be adjusted in a timely manner to ensure core requirements.

[0076] In response to the implicitly frequent state, the storage measures for the corresponding files are adjusted based on the judgment results of the simultaneous existence of invalid storage nodes and search vacancy. It should be understood that files in the implicitly frequent state are objects with low access frequency but high modification demand. The adjustment needs to balance both modification convenience and resource consumption: when only invalid storage nodes exist, the memory occupation of invalid nodes has a limited impact due to the low number of searches, and occasional access can be met by remote retrieval by referring to the relationship table, without the need for additional adjustments; when only search vacancy exists, the impact of low-frequency cross-node calls on retrieval efficiency is negligible. If storage is blindly deployed on vacancy nodes, it will increase the memory burden due to low-frequency use and multiple node storage; only when memory redundancy and access efficiency risks exist simultaneously will the benefits of adjustment (i.e., releasing memory and improving modification access convenience) increase to avoid invalid operations.

[0077] Understandably, cleaning up invalid storage nodes for frequently accessed files can avoid memory waste caused by unused storage and improve the overall utilization of storage resources; supplementing storage in search dummy nodes reduces the latency of frequent cross-node calls, making the retrieval of frequent files smoother; focusing on the core of frequent access, prioritizing the storage optimization of core files aligns with the actual need for prioritizing key resources in file management.

[0078] Understandably, rejecting ineffective adjustments in a single mismatch scenario reduces system computational overhead and storage layout redundancy, aligning with the principle of lightweight optimization. A high read / write ratio means that files need to be updated frequently, while cleaning up invalid nodes (saving memory) and supplementing virtual storage (reducing latency) makes modification and access more convenient. In addition, it does not ignore the potential demand for hidden and frequently accessed files, nor does it consume resources due to excessive adjustments, achieving on-demand optimization and taking into account the management needs of niche core files.

[0079] Specifically, the management module determines the storage adjustment measures for the corresponding files based on the judgment result of the adjustment and the reason for the adjustment, including:

[0080] In response to the reason for the adjustment being the existence of invalid storage nodes, the storage adjustment measure is determined to be to compress the files of the invalid storage nodes and store the compressed files as archives.

[0081] In response to the reason for the adjustment being the existence of search voids, storage adjustment measures were formulated based on the void search ratio of each search void.

[0082] Specifically, the management module stores the original data block of the file (i.e., data that can be opened directly) in the search virtual space where the virtual space search ratio is greater than or equal to the ratio threshold, and stores the compressed package of the file (data that can only be opened after decompression) in the search virtual space where the virtual space search ratio is less than the ratio threshold.

[0083] In implementation, the percentage of virtual space searches = the number of times virtual spaces are searched ÷ the total number of searches for the file; the percentage threshold is usually <50%, preferably set to 20%; it should be understood that the larger the percentage threshold, the more frequently the virtual spaces are searched, that is, the more the original data blocks should be stored on the storage node for subsequent calls.

[0084] Specifically, the management module decompresses the compressed package when the corresponding file is retrieved on the storage node where the compressed package is located (and deletes the compressed package synchronously during decompression), or deletes the compressed package when the duration of a single compressed package exceeds a preset duration (this indicates that the storage node has not searched for the content for a long time, so there is no need for the node to store it and it can be deleted directly).

[0085] It should be understood that the compressed package corresponding to the original data block of each file should exist in at least 2 storage nodes. Therefore, when deleting a compressed package, the number of compressed packages should be determined according to the lookup table before deletion. If there are only 2 compressed packages of the same file, they should not be deleted.

[0086] In practice, the preset duration is usually two to three times the analysis cycle.

[0087] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

[0088] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An archive management system based on a distributed data archive, comprising several distributed storage nodes and a data management center, characterized in that, Each of the aforementioned storage nodes is communicatively connected to the data management center, which further includes: The sorting module is used to match several files in the database based on search keywords and sort the files in real time according to the actual click rate of each file. The search module obtains the search node and storage node of each file, and monitors the search attributes of each file. The search attributes include retrieval attributes and read / write attributes; An analysis module, connected to the search module, is used to determine the search representation status of each file based on the number of searches and search attributes within the analysis period, determine the validity of each storage node of the corresponding file and its search real position and search virtual position based on the storage node and search node, and determine whether to adjust the storage measures of the corresponding file based on the determination result of the search representation status combined with the search virtual position and the validity of the storage node. The management module, which is connected to the analysis module, determines the storage adjustment measures for the corresponding archives based on the judgment result of adjusting the corresponding archives and the reasons for the adjustment, and generates a table of correspondence between each archive and its storage node and stores the table of correspondence in the entire data archive database.

2. The archive management system based on a distributed data archive according to claim 1, characterized in that, The sorting module sorts the matched files in real time according to their actual click-through rates, from highest to lowest.

3. The archive management system based on a distributed data archive according to claim 1, characterized in that, The analysis module acquires all files within the analysis period and determines the search count and search attributes of each file. Based on the search count, it determines whether to combine the search attributes to determine the search representation status, including: Based on the determination result that the number of searches is greater than or equal to the preset number, the search representation status of the corresponding file is determined to be an explicit frequent state; Based on the determination that the number of searches is less than the preset number, the search representation status of the corresponding file is determined by combining the search attributes of each search.

4. The archive management system based on a distributed data archive according to claim 3, characterized in that, The analysis module determines the read / write attribute ratio based on the ratio of the number of searches with read / write attributes to the total number of searches, and determines the search representation status of the corresponding file based on the read / write attribute ratio, including: Based on the determination result that the proportion of read and write attributes is greater than or equal to the preset proportion, the search representation state is determined to be a latent frequent state. Based on the determination result that the proportion of read and write attributes is less than the preset proportion, the search representation state is determined to be an infrequent state.

5. The archive management system based on a distributed data archive according to claim 1, characterized in that, The analysis module identifies nodes in the search nodes that are the same as the storage nodes of the corresponding file as the real search nodes of that file, and identifies nodes in the search nodes that are different from the storage nodes of the corresponding file as the virtual search nodes of that file.

6. The archive management system based on a distributed data archive according to claim 1, characterized in that, The analysis module determines the validity of a storage node based on whether the storage node of the file is also a search node, wherein: In response to the fact that it is also a search node, the storage node is determined to be a valid storage node; If a storage node is not simultaneously being searched, it is determined to be an invalid storage node.

7. The archive management system based on a distributed data archive according to claim 1, characterized in that, The analysis module determines whether to adjust the corresponding file based on the search representation status judgment result, combined with its search vacancy and storage node validity, including: Based on the determination of infrequent status, the storage measures for the corresponding files will not be adjusted. Based on the results of the frequent status determinations, the storage measures for the corresponding files are adjusted, among which, In response to the explicit frequent state, the storage measures of the corresponding file are adjusted based on the determination results of the existence of invalid storage nodes or search vacancy. In response to the latent frequent state, the storage measures of the corresponding file are adjusted based on the determination result that there are both invalid storage nodes and search vacant positions.

8. The archive management system based on a distributed data archive according to claim 1, characterized in that, The management module determines the storage adjustment measures for the corresponding files based on the judgment result of the adjustment and the reason for the adjustment, including: In response to the reason for the adjustment being the existence of invalid storage nodes, the storage adjustment measure is determined to be to compress the files of the invalid storage nodes and store the compressed files as archives. In response to the reason for the adjustment being the existence of search voids, storage adjustment measures were formulated based on the void search ratio of each search void.

9. The archive management system based on a distributed data archive according to claim 8, characterized in that, The management module stores the file in search vacancy slots where the vacancy search ratio is greater than or equal to the ratio threshold, and stores the compressed file in search vacancy slots where the vacancy search ratio is less than the ratio threshold.

10. The archive management system based on a distributed data archive according to claim 8, characterized in that, The management module decompresses the compressed package when it retrieves the corresponding file on the storage node where the compressed package is located, or deletes the compressed package when the duration of a single compressed package exceeds a preset duration.

Citation Information

Patent Citations

  • An archive management system based on distributed storage

    CN119669150B