Archive information storage method and system

By classifying the access frequency and storage time of electronic archive files and adjusting the file storage status according to the classification label, the long-term readability of electronic archives is solved, and efficient access and long-term storage of archive files are achieved.

CN120045516APending Publication Date: 2025-05-27SICHUAN JIAQI LIXIN INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510086947.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to ensure the long-term readability of electronic files, and physical storage devices degenerate over time and rely on specific hardware devices and software applications, resulting in data being unreadable or processed.

Method used

By obtaining the access frequency and storage time of archive files, classifying them, adding classification tags, dividing archive files into high-frequency usage files and low-frequency usage files, and adjusting the file storage status according to the classification tag to ensure that it matches the access requirements.

Benefits of technology

It realizes that the long-term readability of archive files is ensured while ensuring the efficiency of archive files access, and adapts to the access needs of different archive files by flexibly adjusting the file storage status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045516A_ABST
    Figure CN120045516A_ABST
Patent Text Reader

Abstract

The invention provides an archive information storage method and system, and relates to the technical field of data processing, and the method comprises the steps: obtaining the access frequency and storage time of each archive file in a preset time interval; classifying the archive files according to the access frequency and the storage time, adding classification labels, and dividing the archive files into high-frequency use files and low-frequency use files; obtaining a file storage state of the archive file, wherein the file storage state comprises a file storage format, a file storage medium and a file index mode; judging whether the file storage state of the archive file is matched with the access requirement corresponding to the classification tag or not; and if not, adjusting the file storage state of the archive file according to the access requirement corresponding to the classification tag. According to the scheme, the file storage state of the archive file can be flexibly adjusted according to the access requirements and storage time of different archive files, and the long-term readability of the archive file is ensured on the premise that the access efficiency of the archive file is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a method and system for storing archival information. Background Art

[0002] Archives refer to various forms of original records with preservation value directly formed in various social activities by people. With the development of information technology and the advancement of digital transformation, traditional paper archives are gradually being replaced by electronic archives, and many organizations and institutions store archival information in an electronic manner. Compared with the preservation of traditional paper archives, although electronic archives provide convenience and efficiency improvement in many aspects, there are also some specific disadvantages. Electronic storage media still degrade over time, especially physical storage devices such as magnetic disks and optical discs. At the same time, the access and management of electronic archives rely on specific hardware devices and software applications. Over time, the original technology may be phased out or no longer supported, resulting in the inability to read or process data.

[0003] Therefore, how to provide a method for storing archival information that ensures the long-term readability of archival data is an urgent problem to be solved at present. Summary of the Invention

[0004] To improve the above problems, the present invention provides a method and system for storing archival information.

[0005] In the first aspect of the embodiments of the present invention, a method for storing archival information is provided. The method includes: Obtain the access frequency and storage time of each archival file within a preset time interval; Classify the archival files according to the access frequency and storage time and add classification labels, and divide the archival files into frequently used files and infrequently used files; Obtain the file storage status of the archival file, where the file storage status includes file storage format, file storage medium, and file indexing method; Determine whether the file storage status of the archival file matches the access requirements corresponding to the classification label; If not, adjust the file storage status of the archival file according to the access requirements corresponding to the classification label.

[0006] Optionally, the step of classifying the archival files according to the access frequency and storage time specifically includes: Obtain an access frequency adjustment coefficient according to the storage time; Determine whether the adjusted access frequency based on the access frequency adjustment coefficient reaches a preset classification threshold; If it reaches, it is determined as a frequently used file; if it does not reach, it is determined as an infrequently used file.

[0007] Optionally, the step of determining whether the access frequency adjusted by the access frequency adjustment coefficient reaches a preset classification threshold specifically includes: Based on the access rights of the visitor, calculate the access frequencies of visitors with different access rights respectively; For each level of access rights, determine whether the access frequency adjusted by the access frequency adjustment coefficient reaches the preset classification threshold for this access right respectively; If any level reaches, it is determined as a frequently used file; if none of the levels reach, it is determined as an infrequently used file.

[0008] Optionally, the step of determining whether the file storage status of the archive file matches the access requirements corresponding to the classification label specifically includes: When the classification label is a frequently used file, calculate the average feedback time of the file information when the archive file is accessed; Determine whether the average feedback time is greater than the preset high-frequency feedback time. If it is greater, it is considered not to match the access requirements; When the classification label is an infrequently used file, calculate the theoretical readable storage time of the archive file based on the file storage status; Determine whether the theoretical readable storage time is less than the preset low-frequency readable time. If it is less, it is considered not to match the access requirements.

[0009] Optionally, the step of adjusting the file storage status of the archive file according to the access requirements corresponding to the classification label specifically includes: When the classification label is a frequently used file, determine at least one influencing factor that causes the average feedback time to be greater than the high-frequency feedback time according to the influence ratio of the file storage format, file storage medium, and file indexing method on the average feedback time; Adjust the file storage status corresponding to the influencing factor.

[0010] Optionally, the step of adjusting the file storage status corresponding to the influencing factor specifically includes: When the influencing factor is the file storage format, adjust the file storage format to the file storage format with the highest matching degree to the file format used by the visitor; When the influencing factor is the file storage medium, adjust the file storage medium to the storage medium with a faster response speed in the current archive information storage system; When the influencing factor is the file indexing method, optimize the index information and metadata of the archive file.

[0011] Optionally, the step of adjusting the file storage status of the archive file according to the access requirements corresponding to the classification label specifically includes: When the classification label is a low-frequency used file, obtain the available storage medium type according to the low-frequency readable time, and adjust the file storage medium to the available storage medium type supported in the current archive information storage system; Adjust the file storage format to the widely supported file format supported in the current archive information storage system; Standardize the index information and metadata of the archive file.

[0012] Optionally, the step of adjusting the file storage status of the archive file according to the access requirements corresponding to the classification label specifically further includes: When the classification label is a low-frequency used file, generate a copy of the archive file for off-site backup.

[0013] Optionally, the method further includes: After classifying the archive file, determine whether the archive file already has a classification label; If not, add a classification label; If it exists, determine whether the current classification label is the same as the classification result; If they are different, delete the current classification label, and add a new classification label to the archive file according to the classification result.

[0014] In the second aspect of the embodiments of the present invention, an archive information storage system is provided, including: A data acquisition unit, configured to acquire the access frequency and storage time of each archive file within a preset time interval; A file classification unit, configured to classify the archive file according to the access frequency and storage time and add a classification label, and classify the archive file into high-frequency used files and low-frequency used files; A status acquisition unit, configured to acquire the file storage status of the archive file, where the file storage status includes file storage format, file storage medium, and file indexing method; A requirement matching unit, configured to determine whether the file storage status of the archive file matches the access requirements corresponding to the classification label; A status adjustment unit, configured to, if not matching, adjust the file storage status of the archive file according to the access requirements corresponding to the classification label.

[0015] In summary, the present invention provides an archive information storage method and system, which can flexibly adjust the file storage status of archive files according to the access requirements and storage time of different archive files, and ensure the long-term readability of archive files while ensuring the access efficiency of archive files. Brief Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0017] Figure 1 It is a flowchart of the method for storing file information according to an embodiment of the present invention; Figure 2 It is a block diagram of the functional modules of the file information storage system according to an embodiment of the present invention.

[0018] Reference Signs: Data acquisition unit 110; file classification unit 120; status acquisition unit 130; demand matching unit 140; status adjustment unit 150. Detailed Embodiments

[0019] A file refers to various forms of original records with preservation value directly formed in various social activities by people. With the development of information technology and the advancement of digital transformation, traditional paper files are gradually being replaced by electronic files, and many organizations and institutions adopt electronic methods to store file information. Compared with the preservation of traditional paper files, although electronic files provide convenience and efficiency improvement in many aspects, there are also some unique disadvantages. Electronic storage media still degrades over time, especially physical storage devices such as magnetic disks and optical discs. At the same time, the access and management of electronic files rely on specific hardware devices and software applications. Over time, the original technology may be phased out or no longer supported, resulting in data being unreadable or unprocessable.

[0020] Therefore, how to provide a method for storing file information that ensures the long-term readability of file data is an urgent problem to be solved currently.

[0021] In view of this, the designers of the present invention designed a method and system for storing file information.

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated in the drawings here can be arranged and designed in various different configurations.

[0023] Accordingly, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0024] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0025] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by terms such as "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the inventive product is customarily placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. In addition, terms such as "first" and "second" are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.

[0026] In the description of the present invention, it should also be noted that unless otherwise clearly specified and defined, the terms "set", "install", "connect", and "couple" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0027] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0028] A method for storing file information provided in this embodiment will be specifically described below.

[0029] Please refer to Figure 1 , a method for storing file information provided in this embodiment, the method includes: Step S101, obtaining the access frequency and storage time of each file within a preset time interval.

[0030] The access frequency of archive files can be obtained through the logging system of the archive information storage system. First, ensure that the archive information storage system enables the detailed access logging function to record information such as the time of each user access, the identity of the visitor, and the specific archive accessed. By reading the content of the access log, the access frequency of each archive file within a preset time interval can be calculated.

[0031] The storage time refers to the time when the archive file is stored in the archive information storage system, which can be calculated based on the archiving time of the archive file.

[0032] For a preset time interval, one or more preset time intervals can be defined according to the actual archive information management requirements, such as 3 months, 6 months, 1 year, etc., for evaluating the access frequency of archives. At the same time, as the management requirements change, regularly review and adjust these time intervals to adapt to new usage scenarios.

[0033] Step S102, classify the archive files according to the access frequency and storage time and add classification labels, and divide the archive files into frequently used files and infrequently used files.

[0034] Based on the access frequency of archive files within a preset time interval, the archive files can be divided into frequently used files and infrequently used files. Generally, as the storage time increases, the access frequency of users to archive files also decreases.

[0035] In the embodiments of the present invention, a flexible classification system can be established to allow re-evaluation and adjustment of the access frequency for distinguishing archive categories as time goes by and the actual usage situation changes. When specifically classifying, a relatively simple method can be adopted, that is, set corresponding access frequency classification thresholds in advance for different storage times, and classify by comparing the actual access frequency with the classification thresholds. It is also possible to use the method of coefficient addition to preset an influence mapping relationship between the storage time and the classification threshold, and then calculate the corresponding classification threshold according to the specific value of the storage time. Specifically, as a preferred method in the embodiments of the present invention, step S102 specifically includes: Obtain the access frequency adjustment coefficient according to the storage time; Judge whether the access frequency adjusted based on the access frequency adjustment coefficient reaches the preset classification threshold; If it reaches, it is judged as a frequently used file; if it does not reach, it is judged as an infrequently used file.

[0036] It should be noted that the setting method of the classification threshold is not fixed and can dynamically adjust the calculation strategy of the classification threshold according to the change of the demand scenario, so as to adjust the classification result.

[0037] Meanwhile, as a preferred implementation, for the classification of archival documents, in addition to being divided into frequently used documents and infrequently used documents, a more hierarchical classification method can be adopted, such as "frequently used", "medium-frequency used", "infrequently used", "historical archives", etc. Or after the first classification of frequently used documents and infrequently used documents, secondary classification can be carried out separately. A more refined classification method facilitates more refined adjustment during subsequent processing.

[0038] When classifying, in addition to considering the access frequency of archival documents, the situation of the visitors also needs to be considered. When accessing archival information, different access permissions are usually set for different visitors to restrict the content they can access, and different access permissions correspond to users with different identities. Therefore, the classification method can be further refined based on the permissions of the visitors. Specifically, as a preferred implementation, step S102 further specifically includes: Based on the access permissions of the visitors, calculate the access frequencies of visitors with different access permissions respectively; For each level of access permission, determine whether the access frequency adjusted based on the access frequency adjustment coefficient reaches the classification threshold preset for this access permission; If any level reaches, it is determined as a frequently used document; if none of the levels reach, it is determined as an infrequently used document.

[0039] In some systems, the number of visitors with high-level permissions may be relatively small. Therefore, different classification thresholds can be set for different access permission levels to reflect the impact of the access times of visitors with different levels of permissions on the classification results. For example, a visitor with high-level permissions who accesses more than 2 times per month can be classified as a frequently used document, while a visitor with low-level permissions who accesses more than 10 times per month can be classified as a frequently used document. The actual access frequency obtained is 3 times per month for high-level permissions and 5 times per month for low-level permissions. At this time, it is still classified as a frequently used document.

[0040] It should be noted that after classifying the archival documents, determine whether there is already a classification label for the archival document; If not, add a classification label; If it exists, determine whether the current classification label is the same as the classification result; If they are different, delete the current classification label and add a new classification label for the archival document according to the classification result.

[0041] Since this method is executed for stored archive files, if these archive files have already executed this method during storage, it means that they have already been assigned classification labels. If they have not executed this method, they do not have classification labels. For archive files that have already been assigned classification labels, due to a certain time interval since the last execution, their access frequency may have changed within the preset time interval, or the classification strategy may have changed, which may all lead to the classification result being inconsistent with the previous classification label. At this time, subsequent steps need to be executed based on the new classification result. Therefore, for the case where the current classification label is different from the classification result, the current classification label needs to be deleted, and new classification labels need to be added to the archive files according to the classification result.

[0042] Step S103: Obtain the file storage status of the archive file, where the file storage status includes the file storage format, the file storage medium, and the file indexing method.

[0043] The file storage format refers to the file format in which the archive file is saved. The access and management of electronic archives rely on specific hardware devices and software applications. With the rapid development of information technology, file formats are constantly updated. Archive files archived at different time points may be stored in different formats. The file storage medium refers to the type of storage medium used in the storage device for storing the archive file. When storing archive files, different storage media may be used according to their storage requirements. The file indexing method refers to the form of file indexing method used to search for and obtain the corresponding archive file when accessing stored archive files, which may specifically include the indexing strategy and the management method of metadata.

[0044] By obtaining the above information, it is possible to know the specific storage status of the archive file.

[0045] It should be noted that as time goes by, the original technology may be phased out or no longer supported, resulting in data being unreadable or unprocessable. With the rapid development of information technology and the continuous update of file formats, it may be difficult to open or parse old-version files, which all pose various requirements for the file storage status of archive files.

[0046] Step S104: Determine whether the file storage status of the archive file matches the access requirements corresponding to the classification label.

[0047] By matching the file storage status with the access requirements corresponding to the classification label, it is determined whether the current file storage status of the archive file is reasonable, and then corresponding adjustment measures are taken. For archive files of different classifications, their corresponding access requirements are different, so they need to be matched separately.

[0048] As a preferred embodiment of the present invention, step S104 specifically includes: When the classification label is a frequently used file, calculate the average feedback time of the file information when the archive file is accessed; Determine whether the average feedback time is greater than a preset high-frequency feedback time. If it is greater, it is considered that it does not match the access requirement; When the classification label is a rarely used file, calculate the theoretical readable storage time of the archive file based on the file storage status; Determine whether the theoretical readable storage time is less than a preset low-frequency readable time. If it is less, it is considered that it does not match the access requirement.

[0049] For archive files with high and low access frequencies, different strategies need to be adopted to optimize resource utilization, ensure data security and long-term availability.

[0050] For frequently used files, the focus of access requirements is on high-performance storage and fast retrieval. Therefore, the average feedback time of the file information when the archive file is accessed is selected as the measurement standard. The average feedback time refers to the average value of the feedback times during multiple accesses of the archive file within a preset time interval. Selecting the average value can more accurately reflect the overall state within a period of time.

[0051] For rarely used files, the focus of access requirements is on long-term storage to ensure future readability. Therefore, the theoretical readable storage time is adopted as the measurement standard. By comparing the current state of the file with the preset high-frequency feedback time or low-frequency readable time, it is determined whether the file storage state matches the access requirements corresponding to the classification label. The theoretical readable storage time is calculated by subtracting the storage time obtained in step S101 from the theoretical storage time calculated based on the storage medium and technical environment currently used by the file storage state. It represents the storage time during which the archive file can be normally read under theoretical conditions, and the data will not be damaged due to the aging of the storage medium.

[0052] For archive files that match the access requirements, there is no need to adjust their file storage status for the time being. When it is judged again next time and a mismatch occurs, the subsequent steps need to be executed. For archive files that do not match the access requirements, step S105 needs to be executed to adjust the file storage state.

[0053] Step S105, if there is a mismatch, adjust the file storage state of the archive file according to the access requirements corresponding to the classification label.

[0054] It should be noted that when adjusting the file storage status, since the access requirements corresponding to frequently used files and infrequently used files are different, it is necessary to adopt an appropriate adjustment strategy based on the current file storage status of the file.

[0055] Specifically, when the classification label is a frequently used file, at least one influencing factor that causes the average feedback time to be greater than the high-frequency feedback time is determined according to the influence ratio of the file storage format, file storage medium, and file indexing method on the average feedback time; The file storage status corresponding to the influencing factor is adjusted.

[0056] The file storage format, file storage medium, and file indexing method may all affect the feedback time of archival files. Specifically, if the file storage format does not match the file format used by the visitor, format conversion is required, increasing the feedback time. Different storage media have different feedback speeds due to their structures. For example, solid-state drives (SSDs) or NVMe SSDs provide extremely fast data read and write speeds and are very suitable for frequently accessed data. Whether the indexing structure adopted for storing archival files is efficient and supports fast location and retrieval of information will also affect the feedback speed. Therefore, the influence ratio of these three factors on the average feedback time can be calculated separately to determine which factor specifically causes the average feedback time to be greater than the high-frequency feedback time. Depending on the actual situation, the influencing factor may be one or more of the three. Based on the judgment result, the specific adjustment method is determined.

[0057] Specifically, when the influencing factor is the file storage format, the file storage format is adjusted to the file storage format with the highest matching degree to the file format used by the visitor; When the influencing factor is the file storage medium, the file storage medium is adjusted to a storage medium with a faster response speed in the current archival information storage system; When the influencing factor is the file indexing method, the indexing information and metadata of the archival files are optimized.

[0058] When adjusting the file storage format, the file format with the highest usage ratio by the visitor can be counted, and the file storage format of the archival file can be adjusted to the corresponding file format. Or based on the permission level of the visitor, the file storage format of the archival file can be adjusted to the file format used by high-level visitors. In this way, the feedback time increased due to format conversion can be eliminated or significantly reduced.

[0059] When making adjustments for file storage media, it is possible to first determine whether there is a storage media with a faster reaction speed among the storage media used by the current file information storage system. For multiple archive files that need to adjust the storage media simultaneously, they can be stored together on the same high-performance storage media.

[0060] When making adjustments to the file indexing method, targeted optimization can be carried out, including providing detailed indexes and real-time updated metadata. For multiple archive files that need to adjust the file indexing method simultaneously, new index structures and metadata can be created and stored together for subsequent quick search and access.

[0061] Among the above three adjustment actions, depending on the number of determined influencing factors, only one of them may be executed, or multiple of them may be executed simultaneously.

[0062] When the classification label is a low-frequency used file, obtain the available storage media type according to the low-frequency readable time, and adjust the file storage media to the available storage media type supported by the current file information storage system; Adjust the file storage format to the widely supported file format supported by the current file information storage system; Standardize the index information and metadata of the archive file.

[0063] For low-frequency used files, the purpose of the adjustment is to increase the theoretical readable storage time on the premise of ensuring readability. From the calculation method of the theoretical readable storage time, it can be seen that adjustments can be made from two perspectives. One is to increase the theoretical storage time, and the other is to reduce the storage time. These two perspectives can be achieved through one operation, that is, replacing the file storage media. Select a low-cost but reliable cold storage solution, such as a Blu-ray disc library or deep cold cloud storage. These devices are suitable for storing data that is not active for a long time. Determine which types of storage media can meet the requirements of the low-frequency readable time through calculation, and then determine which available storage media types are supported by the current file information storage system. Adjust the low-frequency used files to the storage media of the available storage media type for storage. After adjusting the storage media, its storage time is naturally updated and needs to be recalculated. On this basis, for archive data stored for a long time, periodic migration can be carried out. Every certain period of time (for example, every 5 years), all electronic archives are migrated from the current storage media to a new generation of media, and updated to the latest file format version synchronously. This process should include comprehensive data verification steps to ensure the consistency and integrity of the data before and after migration.

[0064] On the other hand, in order to ensure readability for long-term preservation, the file storage format and file indexing method also need to be properly handled. Efficiently compress infrequently used files to reduce space occupied; use internationally recognized and widely supported file formats, such as PDF / A (for documents), TIFF (for images), MPEG-4 (for videos), etc. These formats are designed with long-term preservation needs in mind and have good backward compatibility. Minimize dependence on specific manufacturers or software versions, as they may no longer be supported over time. For proprietary formats that must be used, the corresponding readers or conversion tools should also be saved. For file indexing methods, the use of standardized metadata schemas (such as Dublin Core) can help improve interoperability and long-term availability. In addition, a "activation on demand" mechanism can be designed, that is, the relevant archives are transferred from cold storage to hot storage for access only when they are really needed.

[0065] On this basis, as a preferred implementation, when the classification label is a low-frequency-used file, a copy is generated for the archive file as an off-site backup.

[0066] Ensure that there are at least two complete copies in different geographical locations to prevent data loss due to a disaster in a single location. Multi-copy offsite backup increases data redundancy and creates multiple copies stored in different geographical locations to ensure that data can be recovered even if a major accident occurs.

[0067] The above adjustment plan, through reasonable classification and storage strategies, ensures that high-frequency and low-frequency used files each obtain the most suitable resource allocation, thereby improving overall efficiency.

[0068] In summary, the archival information storage method provided by the implementation of the present invention can flexibly adjust the file storage status of the archival files according to the access requirements and storage time of different archival files, thereby ensuring the long-term readability of the archival files while ensuring the access efficiency of the archival files.

[0069] like Figure 2 As shown, the archival information storage system provided by the present invention comprises: The data acquisition unit 110 is used to acquire the access frequency and storage time of each archive file within a preset time interval; A file classification unit 120, configured to classify the archive files according to the access frequency and storage time and add classification labels to classify the archive files into high-frequency use files and low-frequency use files; A status acquisition unit 130 is used to acquire the file storage status of the archive file, wherein the file storage status includes the file storage format, the file storage medium and the file indexing method; A requirement matching unit 140 is configured to determine whether the file storage status of the file archive matches the access requirement corresponding to the classification label. A status adjustment unit 150 is configured to, if they do not match, adjust the file storage status of the file archive according to the access requirement corresponding to the classification label.

[0070] The file information storage system provided by the embodiment of the present invention is used to implement the above file information storage method. Therefore, the specific implementation manner is the same as the above method and will not be described herein again.

[0071] In summary, the present invention provides a file information storage method and system, which can flexibly adjust the file storage status of file archives according to the access requirements and storage time of different file archives, and ensure the long-term readability of file archives while guaranteeing the access efficiency of file archives.

[0072] In several embodiments disclosed in the present application, it should be understood that the disclosed apparatus and method can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of apparatuses, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0073] In addition, each functional module in each embodiment of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0074] If the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

Claims

1. A method for storing archive information, characterized in that: The method comprises: Obtain the access frequency and storage time of each archive file within a preset time interval; Classify the archive files according to the access frequency and storage time and add classification labels to divide the archive files into high-frequency use files and low-frequency use files; Acquire the file storage status of the archive file, wherein the file storage status includes the file storage format, the file storage medium and the file indexing method; Determine whether the file storage status of the archive file matches the access requirement corresponding to the classification label; If there is no match, the file storage status of the archive file is adjusted according to the access requirements corresponding to the classification label.

2. The archival information storage method according to claim 1, characterized in that: The step of classifying the archive files according to the access frequency and storage time specifically includes: Obtaining an access frequency adjustment coefficient according to the storage time; Determining whether the access frequency adjusted based on the access frequency adjustment coefficient reaches a preset classification threshold; If it is reached, it is judged as a high-frequency use file, if not reached, it is judged as a low-frequency use file.

3. The archival information storage method according to claim 2, characterized in that: The step of judging whether the access frequency adjusted based on the access frequency adjustment coefficient reaches a preset classification threshold specifically includes: Based on the visitor access rights, the access frequencies of visitors with different access rights are calculated respectively; For each level of access rights, determining whether the access frequency adjusted based on the access frequency adjustment coefficient reaches a preset classification threshold of the access rights; If any of the levels are reached, it is judged as a high-frequency use file. If all the levels are not reached, it is judged as a low-frequency use file.

4. The archival information storage method according to any one of claims 1 to 3, characterized in that: The step of determining whether the file storage status of the archive file matches the access requirement corresponding to the classification label specifically includes: When the classification label is a high-frequency file, the average feedback time of the file information when the archive file is accessed is calculated; Determine whether the average feedback time is greater than a preset high-frequency feedback time, and if so, consider that it does not match the access requirement; When the classification label is a low-frequency-used file, the theoretical readable storage time of the archive file is calculated based on the file storage status; It is determined whether the theoretical readable storage time is less than a preset low-frequency readable time. If so, it is considered that the theoretical readable storage time does not match the access requirement.

5. The archival information storage method according to claim 4, characterized in that: The step of adjusting the file storage status of the archive file according to the access requirements corresponding to the classification tags specifically includes: When the classification label is a frequently used file, at least one influencing factor causing the average feedback time to be greater than the high-frequency feedback time is determined according to the influence ratio of the file storage format, the file storage medium, and the file indexing method on the average feedback time; Adjust the file storage status corresponding to the influencing factors.

6. The archival information storage method according to claim 5, characterized in that: The step of adjusting the file storage status corresponding to the influencing factor specifically includes: When the influencing factor is the file storage format, the file storage format is adjusted to the file storage format that best matches the file format used by the visitor; When the influencing factor is the file storage medium, the file storage medium is adjusted to a storage medium with a faster response speed in the current archive information storage system; When the influencing factor is the file indexing method, the index information and metadata of the archive file are optimized.

7. The archival information storage method according to claim 4, characterized in that: The step of adjusting the file storage status of the archive file according to the access requirements corresponding to the classification tags specifically includes: When the classification label is a low-frequency use file, the available storage medium type is obtained according to the low-frequency readable time, and the file storage medium is adjusted to the available storage medium type supported by the current archive information storage system; Adjust the file storage format to a widely supported file format supported in current archival information storage systems; Standardize index information and metadata for archival documents.

8. The archival information storage method according to claim 7, characterized in that: The step of adjusting the file storage status of the archive file according to the access requirements corresponding to the classification tags specifically includes: When the classification label is a low-frequency use file, a copy of the archive file is generated for off-site backup.

9. The archival information storage method according to claim 1, characterized in that: The method further comprises: After classifying the archive files, determine whether the archive files already have classification tags; Add the category label if it does not exist; If it exists, determine whether the current classification label is the same as the classification result; If they are different, the current classification label is deleted and a new classification label is added to the archive file according to the classification result.

10. An archive information storage system, characterized in that: include: A data acquisition unit, used to acquire the access frequency and storage time of each archive file within a preset time interval; A file classification unit, used to classify the archive files according to the access frequency and storage time and add classification labels to classify the archive files into high-frequency use files and low-frequency use files; A status acquisition unit, used to acquire the file storage status of the archive file, wherein the file storage status includes the file storage format, the file storage medium and the file indexing method; A demand matching unit, used to determine whether the file storage status of the archive file matches the access demand corresponding to the classification label; The state adjustment unit is used to adjust the file storage state of the archive file according to the access requirements corresponding to the classification tags if there is a mismatch.

Citation Information

Cited By

  • Real estate registration archive arrangement service method and system

    CN120386901A

  • Real estate registration file sorting service method and system

    CN120386901B