A file management method and system
By obtaining the basic data of archives and using the classification scoring formula to calculate the classification and importance score of archives, the problem of the inability to accurately evaluate the significance of archives in existing archive management is solved, and scientific management and efficient storage of archives are achieved.
Patent Information
- Application Number
- CN202411625274.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-11-14
AI Technical Summary
The existing archive management methods fail to adopt differentiated management and preservation strategies according to different types of archives, resulting in the inability to accurately assess the actual significance of archives, affecting the setting of management periods and low retrieval efficiency.
By obtaining the basic data of the archives, including the time of creation, content and confidentiality period, the classification score and importance score of the archives are calculated using the classification scoring formula, and the archiving category and retention period are determined in combination with preset standards. Scientific methods are used to classify, score and archive the archives.
It has achieved accurate assessment and scientific management of archives, improved the efficiency and security of archive management, optimized storage resource allocation, and enhanced information security and retrieval efficiency.
Smart Images

Figure CN119557419B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of archive management, and in particular to an archive management method and system. Background Art
[0002] In today's information explosion, driven by rapid socioeconomic development and the widespread adoption of information technology, the amount of documents generated by various organizations and individuals has become richer and more diverse than ever before. These documents encompass not only traditional paper documents but also a vast array of electronic documents, audio, video, and other media formats. This diversified development has led to a rapid increase in the number of archives and a significant surge in the demand for storage space. Traditional physical storage methods, such as filing cabinets and boxes, are no longer able to cope with this massive volume of data and cannot meet the requirements for efficient management and long-term preservation.
[0003] One existing technology is shifting towards digital storage solutions, using electronic archive management systems to manage and preserve vast amounts of archival materials. These institutions store large amounts of paper archival materials on hard drives, cloud servers, or other network platforms through digital processing methods such as scanning and photographing. This technology not only effectively conserves physical space and reduces reliance on physical storage facilities, but also significantly improves archival retrieval efficiency. Users can quickly locate the archival information they need through various methods, such as keyword searches and category browsing, significantly reducing search time and improving work efficiency. Digital storage also facilitates the long-term preservation and security of archives, reducing the risk of information loss due to physical damage or loss, and ensuring the integrity and availability of archival materials.
[0004] Despite this, current archival management methods still have some shortcomings. Existing management methods often focus primarily on the storage of archives, failing to adopt differentiated management and preservation strategies based on different types of archives. For example, most archives are simply classified and stored based on their generation time and category, ignoring factors such as the archive's confidentiality level and integrity. This makes it impossible to accurately assess the actual significance of different archives, which in turn affects the specific setting of archive management periods. In addition, due to the relatively simple management methods, archive management as a whole appears to be relatively chaotic, and information retrieval efficiency is low. Summary of the Invention
[0005] The present invention provides a file management method and system to improve file management efficiency.
[0006] In a first aspect, in order to solve the above technical problems, the present invention provides a file management method, comprising:
[0007] Obtaining basic data of the archive, including: the time when the archive was created, the content of the archive, and the confidentiality period of the archive;
[0008] Perform classification scoring calculation based on the basic data of the archive to obtain an archive classification score;
[0009] Based on the archive classification score and the preset classification score standard, the archive filing category is obtained;
[0010] Calculating an importance rating factor based on the basic data of the archive and the archiving category of the archive;
[0011] Performing an importance score calculation based on the importance rating factor to obtain a file importance score;
[0012] According to the importance score of the archive and the preset retention period standard, a retention period is selected for archiving.
[0013] In an optional embodiment, performing classification score calculation based on the basic data of the archive to obtain the archive classification score includes:
[0014] The file creation time score is calculated using the following formula:
[0015] A=100-Y now +Y Start
[0016] Among them, A is the score of the time when the file was formed, Y now is the current year, Y Start The year in which the archive was created;
[0017] The file classification score is calculated using the following formula:
[0018] W=A+B+C
[0019] Among them, W is the file classification score, B is the file content score, and C is the file confidentiality score;
[0020] The archive confidentiality score is calculated based on the confidentiality level of the archive.
[0021] In an optional embodiment, the calculation of the archive content score includes:
[0022] Extract keywords from the archive content to obtain a preset number of archive keywords;
[0023] Finding the keyword score corresponding to the archive keyword based on the keyword score preset in the keyword database;
[0024] The keyword scores are summed to obtain the archive content score.
[0025] In an optional embodiment, obtaining the archive filing category based on the archive classification score and a preset classification score standard includes:
[0026] The archival categories include: general archives, confidential archives, and top secret archives;
[0027] When the file classification score exceeds the preset top secret file score threshold, the file classification category is top secret file.
[0028] In an optional implementation, calculating the importance rating factor based on the basic archive data and the archive filing category includes:
[0029] The importance rating factors include category factor, completeness factor, and confidentiality factor;
[0030] The categorical factor is calculated using the following formula:
[0031]
[0032] Among them, α is the classification factor, m is the number of categories of archived data, A i is the corresponding importance weight of the i-th category, C i is the classification factor of the i-th category;
[0033] The confidentiality factor is calculated using the following formula:
[0034]
[0035] Among them, S is the confidentiality factor, L is the current file confidentiality level, and M is the maximum limit of the preset confidentiality level;
[0036] The integrity factor is calculated using the following formula:
[0037]
[0038] Where η is the completeness factor, T is the completeness of the archive data, and μ is the amount of archived data.
[0039] In an optional embodiment, performing importance score calculation based on the importance rating factor to obtain the archive importance score includes:
[0040] The importance score of the archive is calculated using the following formula:
[0041] I=α×(S×β-η) / t
[0042] Among them, I is the archive importance score, α is the category factor, η is the completeness factor, S is the confidentiality factor, t is the normalization factor, and β is the importance parameter.
[0043] In an optional embodiment, selecting a retention period for archiving based on the archive importance score and a preset retention period standard includes:
[0044] The retention period standard classifies the retention period as: permanent, long-term or short-term;
[0045] When the importance score of the archive is greater than a first score threshold, the retention period is permanent.
[0046] In a second aspect, the present invention provides a file management system, comprising:
[0047] The file acquisition module is used to obtain basic file data, including: the time when the file was created, the content of the file and the confidentiality period of the file;
[0048] A classification scoring module is used to calculate the classification score based on the basic data of the archive to obtain the archive classification score;
[0049] An archive classification module, configured to obtain an archive filing category based on the archive classification score and a preset classification scoring standard;
[0050] A factor calculation module, used to calculate an importance rating factor based on the basic data of the archive and the archive category of the archive;
[0051] A scoring calculation module, configured to calculate an importance score based on the importance rating factor to obtain an archive importance score;
[0052] The archive archiving module is used to select a retention period for archiving based on the archive importance score and the preset retention period standard.
[0053] In a third aspect, the present invention further provides an electronic device comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for managing an archive as described in any one of the above is implemented.
[0054] In a fourth aspect, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute an archive management method described in any one of the above.
[0055] Compared to the prior art, the present invention has the following beneficial effects: The present invention discloses a method and system for managing archives, the method comprising obtaining basic archive data, the basic archive data including: the archive creation time, the archive content, and the archive confidentiality period; calculating an archive classification score based on the basic archive data using a classification scoring formula; obtaining an archive filing category based on the archive classification score and a preset classification scoring standard; calculating an importance rating factor based on the basic archive data and the archive filing category; calculating an archive importance score based on the importance rating factor combined with an importance scoring formula; and selecting a filing period for filing based on the archive importance score and a preset filing period standard. The present invention proposes a management system for classifying, scoring, and filing archives using a scientific method, significantly improving the efficiency and accuracy of archive management. By processing basic archive data such as the creation time, content, and confidentiality period through an automated process, and calculating the archive classification score using a classification scoring formula, the filing category is determined based on the scoring standard, thereby achieving an objective assessment of the significance and importance of the archive. Specifically, the archive importance score is calculated using a formula. This method comprehensively considers the archive's category, confidentiality level, completeness, normalization factor, and importance parameters, achieving an accurate assessment of the archive's importance. At the same time, this method enhances information security, especially by implementing stricter security measures for archives with high confidentiality levels to prevent the leakage of sensitive information. Furthermore, through scientific classification and importance assessment, this method can better meet the needs of different users and provide personalized information services. Users can set their own weights to adapt to different usage scenarios. In summary, this calculation method greatly improves the efficiency and security of archive management.
[0056] Based on the importance score of the archive and the preset retention period standard, a retention period is selected for archiving. By classifying archives into permanent, long-term, and short-term storage, resource allocation can be optimized to ensure that archives of high importance and sensitivity (such as national laws and records of major historical events) are preserved long-term for future reference and research; archives of certain importance are properly stored for a certain period of time, while archives of lower importance are stored for a short period of time and then destroyed in a timely manner to free up storage space.
[0057] In summary, this method proposes a specific and scientific calculation method that can optimize the allocation of storage resources based on the importance and retention period of archives, ensuring the security of information. At the same time, it establishes a comprehensive information index system, greatly facilitating the retrieval and use of archives, and improving the storage and retrieval efficiency of archives. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 This is a flowchart of a file management method provided by the first embodiment of the present invention;
[0059] Figure 2 This is a schematic diagram of the structure of an archive management system provided by the second embodiment of the present invention. DETAILED DESCRIPTION
[0060] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0061] In today's information explosion, driven by rapid socioeconomic development and the widespread adoption of information technology, the amount of documents generated by various organizations and individuals has become richer and more diverse than ever before. These documents encompass not only traditional paper documents but also a vast array of electronic documents, audio, video, and other media formats. This diversified development has led to a rapid increase in the number of archives and a significant surge in the demand for storage space. Traditional physical storage methods, such as filing cabinets and boxes, are no longer able to cope with this massive volume of data and cannot meet the requirements for efficient management and long-term preservation.
[0062] One existing technology is shifting towards digital storage solutions, using electronic archive management systems to manage and preserve vast amounts of archival materials. These institutions store large amounts of paper archival materials on hard drives, cloud servers, or other network platforms through digital processing methods such as scanning and photographing. This technology not only effectively conserves physical space and reduces reliance on physical storage facilities, but also significantly improves archival retrieval efficiency. Users can quickly locate the archival information they need through various methods, such as keyword searches and category browsing, significantly reducing search time and improving work efficiency. Digital storage also facilitates the long-term preservation and security of archives, reducing the risk of information loss due to physical damage or loss, and ensuring the integrity and availability of archival materials.
[0063] Despite this, current archival management methods still have some shortcomings. Existing management methods often focus primarily on the storage of archives, failing to adopt differentiated management and preservation strategies based on different types of archives. For example, most archives are simply classified and stored based on their generation time and category, ignoring factors such as the archive's confidentiality level and integrity. This makes it impossible to accurately assess the actual significance of different archives, which in turn affects the specific setting of archive management periods. In addition, due to the relatively simple management methods, archive management as a whole appears to be relatively chaotic, and information retrieval efficiency is low.
[0064] To solve the above problems, refer to Figure 1The first embodiment of the present invention provides a file management method, comprising the following steps:
[0065] S11, obtaining basic data of the archive, wherein the basic data of the archive includes: the time when the archive was created, the content of the archive, and the confidentiality period of the archive;
[0066] S12, performing classification scoring calculation based on the basic data of the archive to obtain an archive classification score;
[0067] S13, obtaining the archive filing category based on the archive classification score and the preset classification score standard;
[0068] S14, calculating an importance rating factor based on the basic archive data and the archive filing category;
[0069] S15, performing importance scoring calculation based on the importance rating factor to obtain a file importance score;
[0070] S16, selecting a retention period for archiving based on the file importance score and a preset retention period standard.
[0071] In step S11, basic data of the archive is obtained, and the basic data of the archive includes: the time when the archive is created, the content of the archive and the confidentiality period of the archive.
[0072] In one embodiment, basic data of electronic archives, including the time of creation, content, and confidentiality period, can be obtained by directly reading information such as the creation date and modification date from common electronic document formats through metadata extraction technology. For example, for documents in PDF format, libraries such as PyPDF2 or pdfminer in Python can be used to extract the document's metadata. These libraries provide a rich API that can easily read attributes such as the document's title, author, creation date, and modification date. Similarly, for Microsoft Office documents (such as DOCX, XLSX, etc.), Java libraries such as Apache POI or Python's python-docx library can be used to access and read the file's metadata information. These tools not only support reading basic document information, but also help identify the document's version history, revision history, etc., thereby improving the accuracy and completeness of the archive. In addition, for electronic archives stored in a database, the required information can be directly obtained through means such as SQL queries. Archive management systems provide API interfaces that allow third-party applications to call and automatically obtain archive data. In cases where the technology cannot be fully automated, manual input or verification methods can be used.
[0073] In another embodiment, for scanned documents or other documents in non-text formats, optical character recognition (OCR) technology is an effective means of converting them into editable and searchable text. OCR technology can automatically identify the text content in a document through image processing and pattern recognition technology, and convert it into a computer-readable text format. For example, using the Tesseract OCR engine, not only can text in multiple languages be recognized, but it also supports multiple operating systems and can improve the recognition effect of specific fonts or scenes through training. In addition, commercial software such as Adobe Acrobat Pro DC also integrates advanced OCR functions, which can efficiently process complex documents, such as documents containing charts, tables and different fonts, to ensure that the converted text is both accurate and complete. Through OCR technology, this method can extract key content from documents in non-text formats as a data source for subsequent document classification, archiving and retrieval.
[0074] It is worth noting that the basic data of the archives include: the time when the archives were created, the content of the archives and the confidentiality period of the archives. For example, scientific research papers and policy documents are two very typical types of archives. Scientific research papers usually contain content such as research objectives, research methods, experimental data, analysis results, conclusions and references. For example, a research paper on the impact of climate change will provide a detailed introduction to the selection of the research area, the climate model used, the type of data collected, the data analysis method, the main trends found and the significance for future climate predictions. Policy documents record the policy regulations, implementation guidelines, relevant legal provisions, etc. issued by the government or organization. For example, a new environmental protection policy document includes information such as the background of the policy, the scope of application, the main measures, the implementing agency, the supervision mechanism, and the penalties for violations.
[0075] In step S12, a classification score is calculated based on the basic data of the archive to obtain the archive classification score.
[0076] In one embodiment, the profile creation time score is calculated using the following formula:
[0077] A=100-Y now +Y Start
[0078] Among them, A is the score of the time when the file was formed, Y now is the current year, Y Start The year in which the archive was created;
[0079] The file classification score is calculated using the following formula:
[0080] W=A+B+C
[0081] Among them, W is the file classification score, B is the file content score, and C is the file confidentiality score;
[0082] The archive confidentiality score is calculated based on the confidentiality level of the archive, and the archive confidentiality score for confidential documents is 30 points.
[0083] The calculation of the archive content score includes: extracting keywords from the archive content to obtain a preset number of archive keywords; finding keyword scores of the archive keywords based on preset keyword scores in a keyword database; and summing the keyword scores to obtain the archive content score.
[0084] It's worth noting that this formula reflects the historical significance of an archive by calculating the difference between the archive's creation year and the current year. The earlier the archive's creation date, the higher the score. Secondly, the archive content score is calculated by performing keyword extraction on the archive content to obtain a preset number of archive keywords. Then, based on the preset keyword scores in the keyword database, the keyword scores of these archive keywords are found. Finally, these keyword scores are summed to obtain the archive content score. The archive confidentiality score is calculated based on the archive's confidentiality level. For example, a confidential document has a confidentiality score of 30, while a general document has a confidentiality score of 10. This score is determined by actual usage scenarios and is not limited by this invention. For example, consider a scientific research paper archive, created in 2005, and the current year is 2024. First, the archive creation time score is calculated as 81. Next, keyword extraction is performed on the archive content. The extracted keywords are "climate change," "model prediction," and "data analysis." These keywords have scores of 20, 15, and 10, respectively, in the keyword database. Therefore, the archive content score is 45. Furthermore, since this archive is a confidential document, its confidentiality score is 30. The final total score is 156 points.
[0085] In one embodiment, a keyword extraction method based on a topic model is selected for keyword extraction. The specific steps for keyword extraction from archive content are as follows: First, the archive content is preprocessed, including word segmentation and stop word removal, to ensure the cleanliness and standardization of the text. The preprocessed document collection is then trained using the Latent Dirichlet Allocation (LDA) model to infer the document's topic distribution and topic-word distribution. Each topic is represented by a probability distribution of a set of words. By analyzing these probability distributions, the core words within each topic can be determined. Finally, the words with the highest probability are selected from each topic as keywords. Based on the scores in a preset keyword database, the scores for each keyword are calculated and summed to obtain the archive content score. This method can extract keywords from the archive content. The preset number is three in this example, but this method is not limited to this. For archive keyword databases, the following method can be used: Keywords are extracted from the preprocessed text using a topic model (such as LDA) or other keyword extraction techniques (such as TF-IDF or TextRank). The extracted keywords need to be deduplicated and normalized to ensure keyword consistency and accuracy. The extracted keywords and their associated attributes (such as frequency of occurrence and importance score) are then stored in a database, either a relational database (such as MySQL, PostgreSQL) or a NoSQL database (such as MongoDB). The database table structure should include fields such as keywords, document IDs, number of occurrences, and TF-IDF values to facilitate subsequent query and analysis. This approach allows for efficient management and utilization of keywords in archives, improving retrieval efficiency during archival content analysis.
[0086] In another embodiment, multiple linear regression is used to calculate the file classification score. The multiple linear regression is a statistical technique for analyzing the relationship between multiple independent variables and dependent variables. In the file management system, a prediction model is mainly established by collecting historical data and real-time data of the file, such as file classification, file content and file confidentiality level. The core purpose of multiple linear regression is to use these known input variables (independent variables) to predict the target output variable (dependent variable). Specifically, it determines the degree of influence of these variables on the dependent variable by weighted combination of multiple independent variables, thereby deriving a linear equation. The regression coefficients in this equation, such as a1, b1, and c1, are obtained by fitting historical data, and they reflect the contribution of each independent variable to the dependent variable. In the present invention, multiple linear regression is used to predict the classification category of the file. First, by collecting the current basic information of the file, the system inputs this data as independent variables into the regression model. Then, the model uses these independent variables to perform calculations, and through the previously fitted regression coefficients, the independent variables are combined into a linear equation, and the classification category of the file is finally predicted.
[0087] In step S13, the archive filing category is obtained based on the archive classification score and the preset classification score standard.
[0088] In one embodiment, the file archiving categories include: general files, confidential files, and top secret files;
[0089] When the file classification score exceeds a preset top secret score threshold, the file is classified as a top secret file. In one embodiment, the classification scoring criteria are established by establishing a top secret threshold. The maximum score for the file classification score is 100. For example, the top secret threshold is set at 80. Files with a score exceeding 80 are classified as top secret files. The present invention does not limit the setting of the top secret threshold. Similarly, the classification of confidential and general files is determined based on the set threshold. The rationale behind this approach lies in that, by comprehensively considering the file creation time score, file content score, and file confidentiality score, it can more comprehensively and accurately reflect the importance and sensitivity of the file. Specifically, the file creation time score reflects the historical significance and timeliness of the file; the file content score evaluates the core content and information significance of the file through keyword extraction technology; and the file confidentiality score measures the sensitivity and confidentiality requirements of the file. These three scores jointly determine the final classification of the file, ensuring the scientific and accurate classification. In this way, files of varying importance and sensitivity can be rationally archived, facilitating subsequent retrieval and management, and improving the efficiency and security of file retrieval. For example, confidential and top-secret files can adopt stricter access control measures, such as multi-factor authentication, access logging, and regular audits, to ensure that only authorized personnel can access this sensitive information, thereby preventing information leakage and misuse. General files, on the other hand, can adopt more relaxed access control measures, such as simple username and password verification, allowing more people to easily access and share information, thereby improving work efficiency. This invention does not limit access control measures.
[0090] In another embodiment, in addition to classifying archive categories into general archives, confidential archives, and top secret archives, archives can also be classified according to their frequency of use, for example, into frequently used archives, occasionally used archives, and rarely used archives. The specific classification threshold can be set as follows: archives that have been accessed more than 50 times in the past year are classified as frequently used archives; archives that have been accessed 5 to 50 times are classified as occasionally used archives; and archives that have been accessed less than 5 times are classified as rarely used archives. For example, an annual financial report has been consulted and cited many times in the past year and is therefore classified as a frequently used archive and stored in an easily accessible location; a historical project summary report is only consulted under specific circumstances and is therefore classified as an occasionally used archive and stored in a secondary storage location; a meeting minutes from ten years ago is almost never consulted and is therefore classified as a rarely used archive and stored in cold storage.
[0091] In step S14, an importance rating factor is calculated based on the basic archive data and the archive filing category.
[0092] In one embodiment, the importance rating factors include a category factor, a completeness factor, and a confidentiality factor;
[0093] The categorical factor is calculated using the following formula:
[0094]
[0095] Among them, α is the classification factor, m is the number of categories of archived data, A i is the corresponding importance weight of the i-th category, C i is the classification factor of the i-th category;
[0096] The confidentiality factor is calculated using the following formula:
[0097]
[0098] Where S is the confidentiality factor, L is the current confidentiality level of the archive, and M is the maximum limit of the preset confidentiality level. The integrity factor is calculated using the following formula:
[0099]
[0100] Where η is the completeness factor, T is the completeness of the archive data, and μ is the amount of archived data.
[0101] It is worth noting that the importance rating factor is calculated based on the basic data and archival category of the archive. This factor includes a category factor, a completeness factor, and a confidentiality factor. The category factor is calculated using a formula and reflects the relative importance of the archive's category. Different archive categories (such as general archives, confidential archives, and top secret archives) have different importance weights, and the category factor indicates the specific importance of the archive within that category. The confidentiality factor is calculated using a formula and reflects the confidentiality requirements and sensitivity of the archive. A higher confidentiality level indicates a more sensitive archive, requiring stricter protection measures. The completeness factor is calculated using a formula and reflects the integrity and reliability of the archive data. A higher degree of completeness indicates more complete and reliable information within the archive. The combined calculation of these factors provides a comprehensive assessment of the archive's importance. For example, if a archive belongs to a high-importance category (such as top secret archives), has a high confidentiality level (such as top secret), and has good data integrity, its category factor, confidentiality factor, and completeness factor will all be high, resulting in a high final importance rating factor, indicating that the archive is of great significance for management and protection.
[0102] In one embodiment, data integrity can be obtained through the following steps: First, a comprehensive inspection of the archive content is carried out to confirm whether all necessary information contained in the archive is complete, such as key information such as file title, author, date, and content. Secondly, the missing or incomplete information items in the archive are counted, such as missing signatures, unclear dates, missing content, etc. Then, the data integrity is calculated based on the total number of information items and the number of missing information items in the archive. The specific formula is: data integrity = (total number of information items - number of missing information items) / total number of information items. For example, if a certain archive has a total of 10 information items, of which 2 are missing, the data integrity is (10-2) / 10 = 0.8, or 80%. Through this method, the integrity of the archive can be quantified, providing data for the subsequent calculation of the integrity factor.
[0103] In step S15, an importance score is calculated based on the importance rating factor to obtain a file importance score.
[0104] In one embodiment, the file importance score is calculated using the following formula:
[0105] I=α×(S×β-η) / t
[0106] Among them, I is the archive importance score, α is the category factor, η is the completeness factor, S is the confidentiality factor, t is the normalization factor, and β is the importance parameter.
[0107] It's worth noting that the archive importance score represents the overall importance of a file. The classification factor, calculated using a formula, reflects the relative importance of the file's category. The completeness factor, calculated using a formula, reflects the integrity and reliability of the archive data. The higher the integrity of the archive data, the more complete and reliable the information contained in the archive. The confidentiality factor, calculated using a formula, reflects the confidentiality requirements and sensitivity of the archive. A higher confidentiality level indicates a more sensitive archive, requiring stricter protection measures. The normalization factor standardizes the importance score, ensuring that the same scoring criteria are applied to different archives. The importance parameter adjusts the weight of the confidentiality factor in the importance score. For example, consider a file classified as confidential, with an importance weight of 0.8 and a classification factor of 0.9, resulting in a classification factor of 0.72. The file's confidentiality level is 3, and the preset maximum confidentiality level is 5, so the confidentiality factor is 0.6. The archive's completeness factor is 0.8. The normalization factor is 0.001, and the importance parameter is 1.5. By plugging these values into the formula, we can calculate an importance score of 72, which will be used to determine the subsequent storage period.
[0108] In step S16, a retention period is selected for archiving based on the archive importance score and a preset retention period standard.
[0109] In one embodiment, the storage period standard divides the storage period into: permanent, long-term or short-term; wherein, when the archive importance score is greater than 80 points, the storage period is permanent. The specific content of the storage period standard is as follows: When the archive importance score is greater than 80 points, the storage period is permanent. Such archives usually have extremely high historical, legal or strategic significance and need to be preserved for a long time for future reference and research. When the archive importance score is between 50 and 80 points, the storage period is long-term, specifically 20 to 50 years. Such archives still have high use significance within a certain period of time, such as long-term cooperation agreements. When the archive importance score is less than 50 points, the storage period is short-term, usually 1 to 10 years. Such archives have a short-term use demand, but there is little significance in long-term preservation, such as weekly work plans.
[0110] It's worth noting that archives that have exceeded their retention period must be handled according to the following steps: First, the archive management system evaluates and reviews the archives, including their actual usage and historical significance. Second, based on the evaluation results, the archives are identified and screened. Archives deemed no longer valuable can be destroyed. Archives that still have value can apply for an extension of their retention period. Destruction methods include physical cold destruction (such as storing the storage media in a warehouse) and electronic destruction (such as deleting electronic files) to ensure the confidentiality of archive contents. Following destruction, a destruction certificate must be issued and a new archive created. Applications for an extension of retention period must include the reason and duration of the extension, and the retention period records must be updated to ensure that the archive's management and use comply with the new retention requirements. After processing is complete, relevant records, including evaluation reports, approval documents, and destruction certificates, must be archived. These records are not only crucial for archive management but also serve as valuable information for future audits and traceability. These steps ensure that archives that have exceeded their retention period are properly handled, avoiding unnecessary waste of storage resources.
[0111] In summary, the present invention proposes a file management method and system designed to improve the efficiency and accuracy of file management by scientifically classifying, scoring, and archiving files. This method first obtains basic file data, including the file's creation date, content, and confidentiality period. Using this basic information, the system calculates a classification score for the file using a classification scoring formula. This score comprehensively considers the file's creation date, content, and confidentiality level. Based on the classification score and pre-set classification scoring criteria, the system determines the file's archiving category, including general, confidential, and top secret. Subsequently, based on the file's basic data and archiving category, the system calculates an importance rating factor, which includes a category factor, a completeness factor, and a confidentiality factor. The category factor reflects the relative importance of the file's category; the confidentiality factor reflects the confidentiality requirements and sensitivity of the file; and the completeness factor reflects the integrity and reliability of the file data. Based on these importance rating factors, the system further calculates the file's importance score using the importance scoring formula. Finally, the system selects an appropriate retention period for archiving based on the file's importance score and pre-set retention period criteria. The retention period standard divides the retention period into permanent, long-term, or short-term. When the archive importance score is greater than 80 points, the retention period is permanent. Such archives have extremely high historical, legal, or strategic significance and need to be preserved for a long time for future reference and research. When the archive importance score is between 50 and 80 points, the retention period is long-term, usually 20 to 50 years. Such archives still have high use value within a certain period of time, such as long-term cooperation agreements. When the archive importance score is less than 50 points, the retention period is short-term, usually one to ten years. Such archives have short-term use needs but are not very meaningful for long-term preservation, such as weekly work plans. Through this scientific classification and scoring method, the present invention not only optimizes the allocation of archive resources and improves the efficiency and security of archive management. In addition, this method also enhances the security of information, especially for archives with high confidentiality levels, taking more stringent security measures to prevent the leakage of sensitive information. Furthermore, users can set weights to adapt to different usage scenarios, thereby improving the storage efficiency and retrieval efficiency of archive management.
[0112] Reference Figure 2 A second embodiment of the present invention provides a file management method, comprising:
[0113] The file acquisition module is used to obtain basic file data, including: the time when the file was created, the content of the file and the confidentiality period of the file;
[0114] A classification scoring module is used to calculate the classification score of the archive based on the basic data of the archive;
[0115] An archive classification module, configured to obtain an archive filing category based on the archive classification score and a preset classification scoring standard;
[0116] A factor calculation module, used to calculate an importance rating factor based on the basic data of the archive and the archive category of the archive;
[0117] A scoring calculation module, configured to calculate an importance score based on the importance rating factor to obtain a file importance score;
[0118] The archive archiving module is used to select a retention period for archiving based on the archive importance score and the preset retention period standard.
[0119] Preferably, the archive acquisition module is used to:
[0120] Obtain basic data of the archive, including: the time when the archive was created, the content of the archive and the confidentiality period of the archive.
[0121] Preferably, the classification scoring module is used to:
[0122] The file classification score is calculated based on the basic data of the file using the classification scoring formula, including:
[0123] The file creation time score is calculated using the following formula:
[0124] A=100-Y now +Y Start
[0125] Among them, A is the score of the time when the file was formed, Y now is the current year, Y Sta rt is the year in which the archive was created;
[0126] The file classification score is calculated using the following formula:
[0127] W=A+B+C
[0128] Among them, W is the file classification score, B is the file content score, and C is the file confidentiality score;
[0129] The archive confidentiality score is calculated based on the confidentiality level of the archive, and the archive confidentiality score for confidential documents is 30 points.
[0130] The calculation of the file content score includes:
[0131] Extract keywords from the archive content to obtain a preset number of archive keywords;
[0132] Finding the keyword score of the archive keyword based on the keyword score preset in the keyword database;
[0133] The keyword scores are summed to obtain the archive content score.
[0134] Preferably, the archive classification module is used to:
[0135] Based on the archive classification score and the preset classification score standard, the archive filing category is obtained, including:
[0136] The archival categories include: general archives, confidential archives, and top secret archives;
[0137] When the file classification score exceeds the preset top secret file score threshold, the file filing category is top secret file.
[0138] Preferably, the factor calculation module is used to:
[0139] The importance rating factors are calculated based on the basic data of the archive and the archive category, including:
[0140] The importance rating factors include category factor, completeness factor, and confidentiality factor;
[0141] The categorical factor is calculated using the following formula:
[0142]
[0143] Among them, α is the classification factor, m is the number of categories of archived data, A i is the corresponding importance weight of the i-th category, C i is the classification factor of the i-th category;
[0144] The confidentiality factor is calculated using the following formula:
[0145]
[0146] Where S is the confidentiality factor, L is the current confidentiality level of the archive, and M is the maximum limit of the preset confidentiality level. The integrity factor is calculated using the following formula:
[0147]
[0148] Where η is the completeness factor, T is the completeness of the archive data, and μ is the amount of archived data.
[0149] Preferably, the scoring calculation module is used to:
[0150] The importance score of the archive is calculated based on the importance rating factors, including:
[0151] The importance score of the archive is calculated using the following formula:
[0152] I=α×(S×β-η) / t
[0153] Among them, I is the archive importance score, α is the category factor, η is the completeness factor, S is the confidentiality factor, t is the normalization factor, and β is the importance parameter.
[0154] Preferably, the archive filing module is used to:
[0155] According to the importance score of the archive and the preset retention period standard, select the retention period for archiving, including:
[0156] The retention period standard classifies the retention period as: permanent, long-term or short-term;
[0157] Among them, when the importance score of the archive is greater than 80 points, the retention period is permanent.
[0158] It should be noted that the archive management system provided in the embodiment of the present invention is used to execute all the process steps of the archive management method in the above embodiment. The working principles and beneficial effects of the two correspond one to one, and thus will not be described in detail.
[0159] An embodiment of the present invention further provides an electronic device. The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a data acquisition program. When the processor executes the computer program, the steps of the above-mentioned various archive management method embodiments are implemented, such as Figure 1 Alternatively, when the processor executes the computer program, the functions of the modules / units in the above-mentioned device embodiments are realized, such as the archive acquisition module.
[0160] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.
[0161] The electronic device may be a computing device such as a desktop computer, notebook, PDA, or smart tablet. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will appreciate that the aforementioned components are merely examples of electronic devices and do not constitute a limitation of the electronic device. The electronic device may include more or fewer components than those described above, or a combination of certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, and the like.
[0162] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the electronic device, connecting various parts of the entire electronic device using various interfaces and lines.
[0163] The memory can be used to store the computer programs and / or modules, and the processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created based on the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory can include a high-speed random access memory and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0164] Wherein, if the module / unit integrated in the electronic device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of each of the above-mentioned method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0165] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive effort.
[0166] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A method for managing archives, characterized in that: include: Obtaining basic data of the archive, including: the time when the archive was created, the content of the archive, and the confidentiality period of the archive; Perform classification scoring calculation based on the basic data of the archive to obtain the archive classification score; Based on the archive classification score and the preset classification score standard, the archive filing category is obtained; Calculating an importance rating factor based on the basic data of the archive and the archiving category of the archive; Calculating the importance score based on the importance rating factor to obtain the file importance score; Select a retention period for archiving based on the importance score of the archive and the preset retention period standard; The calculation of the importance rating factor based on the basic archive data and the archive category includes: The importance rating factors include category factor, completeness factor, and confidentiality factor; The categorical factor is calculated using the following formula: Among them, α is the classification factor, m is the number of categories of archived data, A i is the corresponding importance weight of the i-th category, C i is the classification factor of the i-th category; The confidentiality factor is calculated using the following formula: Among them, S is the confidentiality factor, L is the current file confidentiality level, and M is the maximum limit of the preset confidentiality level; The integrity factor is calculated using the following formula: Where η is the completeness factor, T is the completeness of the archive data, and μ is the amount of archived data; The importance score calculation based on the importance rating factor to obtain the archive importance score includes: The importance score of the archive is calculated using the following formula: I=α×(S×β-η) / t Among them, I is the archive importance score, α is the category factor, η is the completeness factor, S is the confidentiality factor, t is the normalization factor, and β is the importance parameter.
2. The file management method according to claim 1, characterized in that: The classification score calculation is performed based on the basic data of the archive to obtain the archive classification score, including: The file creation time score is calculated using the following formula: A=100-Y now +Y Start Among them, A is the score of the time when the file was formed, Y now is the current year, Y Start The year in which the archive was created; The file classification score is calculated using the following formula: W=A+B+C Among them, W is the file classification score, B is the file content score, and C is the file confidentiality score; The archive confidentiality score is calculated based on the confidentiality level of the archive.
3. The file management method according to claim 2, characterized in that: The calculation of the file content score includes: Extract keywords from the archive content to obtain a preset number of archive keywords; Finding the keyword score corresponding to the archive keyword based on the keyword score preset in the keyword database; The keyword scores are summed to obtain the archive content score.
4. The file management method according to claim 1, characterized in that: The file classification scoring and the preset classification scoring standard are used to obtain the file filing category, including: The archival categories include: general archives, confidential archives, and top secret archives; When the file classification score exceeds the preset top secret file score threshold, the file classification category is top secret file.
5. The file management method according to claim 1, characterized in that: The selection of a retention period for archiving based on the importance score of the archive and the preset retention period standard includes: The retention period standard classifies the retention period as: permanent, long-term or short-term; When the importance score of the archive is greater than a first score threshold, the retention period is permanent.
6. A file management system, characterized in that: A method for implementing the archive management method according to any one of claims 1 to 5, comprising: The file acquisition module is used to obtain basic file data, including: the time when the file was created, the content of the file and the confidentiality period of the file; A classification scoring module is used to calculate the classification score based on the basic data of the archive to obtain the archive classification score; An archive classification module, configured to obtain an archive filing category based on the archive classification score and a preset classification scoring standard; A factor calculation module, used to calculate an importance rating factor based on the basic data of the archive and the archive category of the archive; A scoring calculation module, configured to calculate an importance score based on the importance rating factor to obtain an archive importance score; The archive archiving module is used to select a retention period for archiving based on the archive importance score and the preset retention period standard.
7. An electronic device, characterized in that: The invention comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the archive management method according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the archive management method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Electronic file management method and system, electronic equipment and storage medium
CN118796760A