Intelligent Management Method and System for Digital Archives Based on Big Data

Through a big data-based method, combined with the security requirements and complexity of paper archives, and adopting automated or manual digitalization strategies, the problem of unclear digital order in digital archive management is solved, management efficiency and information security are improved, and file management is optimized and upgraded.

CN118673196BActive Publication Date: 2025-07-08HEBEI YUANYING INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410835416.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-26
Publication Date
2025-07-08
Estimated Expiration
2044-06-26

AI Technical Summary

Technical Problem

In the process of digital archive management, there is a lack of determination of whether paper archives need to be digitized and the order of digitization of archives, resulting in resource waste and information sharing being limited, affecting the organization's information management and work efficiency.

Method used

Through a big data-based approach, based on the security requirements and complexity of paper archives, automated or manual digitalization strategies are adopted, combined with fuzzy rules and Bayesian methods, digitalization methods are determined, and intelligently managed through data acquisition, processing and storage modules.

Benefits of technology

It improves the efficiency of archive management, reduces risks, and realizes comprehensive optimization and upgrading of digital archive management to ensure information security and availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118673196B_ABST
    Figure CN118673196B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent management method and system for digital archives based on big data, which relates to the technical field of archive management and is used to solve the problems of the lack of determination of whether paper archives need to be digitized and the sequence of priority for archive digitization; it includes adjusting and determining a digitalization method selection strategy according to the security requirements and complexity of paper archives; collecting analysis data required for the security requirements and complexity of paper archives, calculating to obtain the security requirements and complexity of paper archives; determining the digitalization method according to the security requirements and complexity of paper archives, and starting to implement; through intelligent management of digital archives and combining the security requirements and complexity of paper archives to select a method for digitizing paper archives, a flexible and efficient selection method is provided for digitization in the intelligent management of digital archives. It improves the archive management efficiency and reduces risks, thereby realizing the comprehensive optimization and upgrade of digital archive management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of file management. More specifically, the present invention relates to an intelligent management method and system for digital files based on big data. Background Art

[0002] Big Data refers to a collection of data that is large in scale, diverse in type, and fast in processing speed, such that conventional software tools are unable to capture, manage, and process it. These data collections can include structured data (such as data in relational databases), unstructured data (such as text, images, videos, etc.), and semi-structured data (such as XML files, JSON files, etc.). The core characteristics of big data are usually summarized as 3V: Volume (large quantity), Velocity (high speed), and Variety (diversity).

[0003] Intelligent management of digital files refers to the use of advanced technologies and intelligent systems to manage the digital files and information of an organization or an individual. This management method can improve work efficiency, reduce costs, enhance security, and make information more accessible and useful.

[0004] In today's digital age, the management and utilization of information have become crucial. However, many organizations or institutions face a common problem in the digital transformation process of digital file management: the lack of determination of whether paper files need to be digitized and the sequence of file digitization. The existence of this problem may have various impacts on the organization. First, the lack of a clear definition of the digitalization requirements for paper files may lead to waste of resources and reduced efficiency. Second, the cumbersome nature of paper file management may limit the effectiveness of information sharing and collaboration, thus affecting communication and collaboration within the organization. Therefore, establishing a clear digitalization strategy for paper files and carrying out digital transformation in a reasonable sequence is crucial for an organization to achieve informatization management, improve work efficiency, and ensure information security.

[0005] In view of the above problems, the present invention proposes a solution. Summary of the Invention

[0006] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide an intelligent management method and system for digital files based on big data, XX to solve the problems raised in the above background art.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] An intelligent management method for digital files based on big data, comprising the following steps:

[0009] Step 1: Adjust and determine the digitalization method selection strategy according to the security requirements and complexity of paper files;

[0010] Step 2: Collect the analysis data required for the security requirements and complexity of the paper archives, and perform calculations to obtain the security requirements and complexity of the paper archives.

[0011] Step 3: Determine the digitization method based on the security requirements and complexity of the paper archives, and start implementation.

[0012] In a preferred embodiment, in Step 1, it is mainly determined whether to perform automated digitization or manual digitization on the paper archives according to the security requirements and complexity of the paper archives.

[0013] In a preferred embodiment, in Step 2, the specific process of calculating the security requirements of the paper archives is as follows:

[0014] Y1: Collect objective data; record the types and quantities of sensitive information in the paper archives, and conduct regular integrity audits on the paper archives.

[0015] Y2: Calculate confidentiality and integrity; Let N be the number of samples, i.e., the number of paper archives, Nc be the number of times an event occurs, and the confidentiality data be C. Then Integrity is the proportion of the content of the paper archives to the original data.

[0016] Y3: Obtain the security requirement index data; perform weighted calculations on the encryption and integrity of the paper archives to obtain the security requirement index data of the paper archives, label it as X, and make a judgment on X. If X is greater than or equal to the requirement index threshold, then let X = 1, indicating that the sample meets the security requirements; if X is less than the requirement index threshold, then let X = 0, indicating that the sample does not meet the security requirements.

[0017] Y4: Calculate the confidence level of the security requirement index X; specifically include the following steps:

[0018] Y4.1: Label the confidence level of the security requirement index as A(X), and collect N samples D = {d1, d2, d3,..., d N}; Each sample represents the observation result of the security requirement index X, and use the Bayesian method to calculate A(X).

[0019] Y4.2: Select the prior probability; select the prior probability distribution P(X).

[0020] Y4.3: Calculate the likelihood function; determine the likelihood function according to the sample data and the security requirement index; use the binomial distribution as the likelihood function; Let p be the probability of meeting the security requirements, then the likelihood function can be expressed as P(D|X) = p k ×(1 - p) N-k : where k is the number of samples that meet the security requirements, and N is the total number of samples.

[0021] Y4.4: Calculate the posterior probability; according to Bayes' theorem, calculate the posterior probability of the security requirement index.

[0022] Y4.5: Calculate the confidence level based on the posterior probability; select the mean value of the posterior probability as the estimated value of the confidence level.

[0023] In a preferred embodiment, in step 3, define the security requirements and complexity of the obtained paper files as input variables, and divide them into different fuzzy sets; define the selected digitization method as the output variable, and label automated digitization as P1 and manual digitization as P2.

[0024] Formulate fuzzy rules to describe the influence of different input variables on the output variable.

[0025] Perform fuzzy inference according to the fuzzy rules to determine the digitization method.

[0026] In a preferred embodiment, in step Y2, the confidentiality data C of the paper file is obtained. The confidentiality data involves the types and quantities of sensitive information in the paper file. Normalize the confidentiality data so that it is mapped to the interval [0, 1], and then convert it into binary data containing 0 and 1 by comparing it with the confidentiality threshold; repeat steps Y4, that is, Y4.1 to Y4.5, to obtain the confidence level of the confidentiality data and label it as A(C), that is, only calculate the confidence level of the confidentiality data; if A(C) is relatively high, the data importance of the digital file is relatively high; otherwise, the data importance of the digital file is relatively low.

[0027] In a preferred embodiment, the comprehensive analysis of the data importance and backup storage cost of the digital file is as follows:

[0028] First, sort the observed values of data security and backup storage cost according to their magnitudes respectively to obtain their ranks; if there are observed values with the same numerical value, take the average of their ranks; calculate the difference between the ranks of each pair of data points to obtain di, and the specific formula is rank (Ui) - (Vi), where Ui is the observed value of data security of the i-th data point; Vi is the observed value of backup storage cost of the i-th data point, and the observed value is the specific numerical value of each variable in the data; then calculate its Spearman correlation coefficient, and the specific formula is In the formula, ∑di 2 is the sum of the squares of all rank differences di, and n is the number of data points.

[0029] If ρ is close to 1, more frequent backups are required, and at this time, the time of regular backups needs to be reduced; if ρ is close to -1, the backup frequency is reduced, that is, the time of regular backups is increased; if ρ is close to 0, the time of regular backups remains unchanged.

[0030] The intelligent management system for digital archives based on big data includes a data acquisition module, a data processing module, and a data storage module;

[0031] The data acquisition module is used to obtain the security requirement data and complexity data of paper archives and send them to the data processing module to ensure the operation of the subsequent data processing module;

[0032] The data processing module is used to determine whether to perform automatic digitization or manual digitization on paper archives according to the data collected by the data acquisition module;

[0033] The data storage module is used to store all the data generated during the processing of the intelligent management system for digital archives

[0034] The technical effects and advantages of the intelligent management method and system for digital archives based on big data of the present invention:

[0035] Through the intelligent management of digital archives and the selection of methods for digitizing paper archives in combination with the security requirements and complexity of paper archives, the present invention provides a flexible and efficient selection method for digitization in the intelligent management of digital archives. Selecting an appropriate digitization method can improve the efficiency of archive management and reduce risks, thereby realizing the comprehensive optimization and upgrading of digital archive management. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a logical schematic diagram of the intelligent management system for digital archives based on big data of the present invention;

[0037] Figure 2 It is a flowchart of the intelligent management method for digital archives based on big data of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0039] The present invention intelligently manages digital archives based on big data and selects methods for digitizing paper archives in combination with the security requirements and complexity of paper archives, providing a flexible and efficient selection method for digitization in the intelligent management of digital archives. Selecting an appropriate digitization method can improve the efficiency of archive management and reduce risks, thereby realizing the comprehensive optimization and upgrading of digital archive management.

[0040] Embodiment 1

[0041] Figure 1 The logical schematic diagram of the intelligent management system for digital archives based on big data of the present invention is given. Figure 2 The flowchart of the intelligent management method for digital archives based on big data is given, specifically including the following steps:

[0042] Step 1: Adjust and determine the digitalization method selection strategy according to the security requirements and complexity of paper archives.

[0043] Step 2: Collect the analysis data required for the security requirements and complexity of paper archives, and perform calculations to obtain the security requirements and complexity of paper archives.

[0044] Step 3: Obtain a suitable digitalization method according to the security requirements and complexity of paper archives, and start implementation.

[0045] Specifically in Step 1, it is mainly determined whether to perform automated digitalization or manual digitalization on paper archives according to the security requirements and complexity of paper archives. Automated methods may have higher security risks because they usually rely on computer networks and software and may be vulnerable to cyberattacks or data leaks. While manual methods may be more secure because manual processing can better control access and data flow. When the complexity of paper archives is relatively high or low, it affects the selection of the digitalization method for paper archives. If the security requirements of paper archives are high and the complexity of paper archives is relatively high, that is, complex, at this time, manual digitalization of paper archives needs to be carried out to improve higher security and accuracy and better control access and data flow; otherwise, automated digitalization is carried out to process a large number of documents at a faster speed and improve the completion speed of the project. Automated methods may be more suitable for a large number of documents with similar formats, while manual methods may be more suitable for documents with complex layouts or special formats.

[0046] That is, the digitalization method selection strategy for paper archives includes automated digitalization and manual digitalization of paper archives.

[0047] Specifically, in step 2, the security requirements of paper archives refer to the security requirements for the storage, protection, and management of paper archives, such as confidentiality, integrity, availability, and disaster recovery. Confidentiality refers to the degree to which the information contained in paper archives is kept confidential from the outside world. Integrity means whether the content of the archives is complete, accurate, and not tampered with. Availability indicates whether the archives can be used in a timely manner when needed. Disaster recovery refers to the ability of the archive system to quickly recover after suffering a disaster. It is necessary to collect the security requirement data of paper archives, including confidentiality data, integrity data, availability data, and disaster recovery data of paper archives. The security requirement data of the paper archives in this application includes multiple categories of security requirements. Specifically, which categories of security requirements are included are determined according to the actual situation. For example, the security requirement data of paper archives may only be confidentiality data and integrity data. In this case, only the confidentiality and integrity of the paper archives need to be calculated. This application provides a method for evaluating the security requirements of paper archives based on the confidentiality and integrity of paper archives. The specific process of calculating the security requirements of paper archives is as follows:

[0048] Y1: Collect objective data. Record the types and quantities of sensitive information in paper archives, such as personal identity information and financial data, etc. Conduct regular integrity audits on paper archives, and use relevant systems or software to compare the consistency between the archives and the original data.

[0049] Y2: Calculate confidentiality and integrity. Assume that N is the number of samples, that is, the number of paper archives, Nc is the number of times an event occurs, and the confidentiality data is C. Then there is It means that among N paper archives, Nc archives contain sensitive information. Integrity is the proportion of the content of paper archives to the original data, which is labeled as B, and the value ranges from 0 to 1.

[0050] Y3: Obtain the security requirement index data. Calculate the weighted sum of the encryption and integrity of paper archives to obtain the security requirement index data of paper archives, which is labeled as X. Then there is the formula X = wb×B + wc×C; where wb and wc are the weight coefficients of integrity and confidentiality respectively, and wb + wc = 1. Determine X. If X is greater than or equal to the requirement index threshold, then let X = 1, indicating that the sample meets the security requirements; if X is less than the requirement index threshold, then let X = 0, indicating that the sample does not meet the security requirements. The specific values of the weight coefficients and the requirement index threshold are determined according to the professional knowledge and actual situation in the relevant field and are not limited here.

[0051] Y4: Calculate the confidence level of the security requirement index X. Specifically, it includes the following steps:

[0052] Y4.1: Label the confidence level of the security requirement index as A(X), and collect N samples D = {d1, d2, d3..., d N}; Each sample represents the observation result of the security requirement indicator X, and A(X) is calculated using the Bayesian method.

[0053] Y4.2: Select the prior probability. Select an appropriate prior probability distribution P(X) that reflects the prior belief about the security requirement indicator X. Common choices include the Beta distribution, normal distribution, etc. The selection of the prior probability can be based on past experience, expert opinions, or domain knowledge.

[0054] Y4.3: Calculate the likelihood function. Determine the likelihood function based on the sample data and the security requirement indicator. The likelihood function represents the probability of observing the sample data given the security requirement indicator. For binary security requirement indicators, the binomial distribution is usually used as the likelihood function; for continuous indicators, the normal distribution or other appropriate distributions can be used. In this method, the binomial distribution is used as the likelihood function for the binary security requirement indicator X. Assuming p is the probability of meeting the security requirement, the likelihood function can be expressed as P(D|X) = p k ×(1 - p) N-k ; where k is the number of samples that meet the security requirements, and N is the total number of samples.

[0055] Y4.4: Calculate the posterior probability. Calculate the posterior probability of the security requirement indicator according to Bayes' theorem. Bayes' theorem shows that the posterior probability is proportional to the product of the prior probability and the likelihood function. The specific calculation formula is where P(X|D) is the posterior probability of the security requirement indicator X when the sample data D is observed, P(D|X) is the likelihood function, P(X) is the prior probability, and P(D) is the normalization constant, representing the probability of observing the sample data D.

[0056] Y4.5: Calculate the confidence level based on the posterior probability. Select the mean of the posterior probability as the estimated value of the confidence level. The specific formula is A(X) = E[X|D] = P(X = 1|D); where E[X|D] is the expected value of the posterior probability of the security requirement indicator X when the sample data D is observed. P(X = 1|D) represents the probability of meeting the security requirement X = 1 when the sample data D is observed.

[0057] The confidence level represents the degree of confidence in the security requirements of the paper archives. If the confidence level is high, it means a high confidence in the system's ability to meet the security requirements. In this case, one may be more inclined to accept the current security state and not digitize it; if the confidence level is low, it indicates insufficient confidence in the system's ability to meet the security requirements, and further security measures need to be considered to reduce risks, so digitization is required. That is, a high confidence level means a lower security requirement, and a low confidence level means a higher security requirement.

[0058] It should be noted that the security requirements of paper archives are only exemplified in this embodiment by calculating confidentiality and integrity. In fact, it may include other data, such as availability and disaster recovery, etc., which will not be elaborated here.

[0059] The complexity of paper archives includes content complexity. Obviously, if the content complexity of paper archives is high, that is, complex, manual digitization of paper archives is required; if the content complexity of paper archives is low, that is, simple, automated digitization of paper archives is required.

[0060] Furthermore, the complexity of paper archives can also include more types of data according to actual situations, such as structural complexity. Different structural complexities result in different digitization methods. The content complexity exemplified in this application is only for illustration, and it can be specifically evaluated and calculated through images, icons, and special symbols of paper archives, which will not be elaborated in detail here.

[0061] Specifically, in step 3, the security requirements and complexity of the obtained paper archives are defined as input variables and divided into different fuzzy sets. For example, "Low", "Medium", "High" for the security requirements of paper archives, and "Complex", "Average", "Simple" for the complexity of paper archives. The selected digitization method is defined as the output variable, and automated digitization is labeled as P1, and manual digitization is labeled as P2.

[0062] Formulate fuzzy rules to describe the influence of different input variables on the output variable. The definition of the rules can be based on the professional knowledge of this industry or obtained through data analysis and experiments. For example, the security requirements of paper archives are marked as D, the complexity of paper archives is marked as F, and the selected digitization method is marked as R_select. It can be defined as:

[0063] Rule 1: IF (D is High) AND (F is Complex) THEN (R_select is P2)

[0064] Rule 2: IF (D is Low) AND (F is Simple) THEN (R_select is P1) ......

[0066] Perform fuzzy reasoning according to the fuzzy rules to obtain a suitable vulnerability scanning method.

[0067] It should be noted that the division of fuzzy sets can be adjusted according to the actual situation. For example, in this embodiment, three fuzzy sets are taken as an example. In fact, the security requirements, complexity, and selection of digitalization methods of paper archives can be divided into more than three sets, so as to more conveniently select digitalization methods.

[0068] Furthermore, for the judgment of high, medium, and low security requirements and complex, general, and simple complexity, the demand index threshold can be set according to the actual situation for judgment. For example, when the confidence level A(X) in Y4.5 is ≥0.8, the corresponding security requirement is low, and it is labeled as "Low"; when the number of images, icons, and special symbols in the paper archive is greater than or equal to 30, the complexity is labeled as "Complex", etc., which will not be elaborated here.

[0069] During the implementation process, the specific process of automatic digitalization of paper archives is as follows:

[0070] First, sort and classify the paper archives to be converted, which helps with subsequent scanning and management. Second, ensure there is appropriate scanning equipment, such as a high-quality scanner or multifunctional printer, to ensure the scanning quality. Finally, prepare digital storage devices, such as computer hard drives, network servers, or cloud storage, for storing the scanned digital files.

[0071] Second, place the paper archives on the scanning equipment. Ensure the documents are neatly arranged and adjust the position and orientation of the paper as needed. Then, use the scanning equipment to scan, and select appropriate scanning settings, including resolution, color mode, and file format, etc. During the entire scanning process, ensure the clarity and integrity of the documents, and avoid omission or scanning errors.

[0072] After the scanning is completed, process and edit the digital files to improve their quality and usability. This may include image processing, such as cropping, rotating, adjusting contrast and clarity, etc., to improve the image quality. Additionally, if the scanned images need to be converted into editable text files, OCR (Optical Character Recognition) processing is required for subsequent text editing and searching.

[0073] Then, name and archive the scanned files, and establish a unified file naming and directory structure. The naming should be clear and can accurately reflect the file content and important information. At the same time, establish an appropriate file directory structure, classify and organize according to the file type, date, project name, etc., for convenient subsequent management and retrieval.

[0074] Then store the digitized files in a suitable storage device and perform regular backups. Select a suitable storage device, such as a computer hard drive, network server, or cloud storage, to ensure its security and reliability. Regular backups are very important to prevent data loss or damage, especially in case of unexpected events.

[0075] Finally, if more advanced document management functions are needed, consider using a document management system. These systems provide functions such as search, version control, permission management, and workflow management, which help improve the efficiency and controllability of file management. Select a suitable document management system and customize the configuration according to the organization's needs and scale to maximize its value.

[0076] The specific process of manually digitizing paper archives is as follows:

[0077] T1: Preparation. Take out the paper archives to be converted from the filing cabinet and stack them up. Then, open the lid of the scanning device or raise the scanning table to prepare a place for placing the papers.

[0078] T2: Set scanning parameters. Adjust the parameters of the scanning device, including resolution, color mode, and file format, etc. This process may need to be operated on the control panel of the device, such as turning the knob or pressing the button.

[0079] T3: Scan paper archives. Place the first page of the paper archive on the scanning board of the scanning device to ensure they are neatly arranged. Then, close the lid of the scanning device or lower the scanning table to a suitable position to make the paper contact the scanning head. Finally, press the scan button on the scanning device to start the scanning process. It may be necessary to manually guide the device to scan the paper to ensure that each page is scanned completely.

[0080] T4: Check the scanning results. After scanning is completed, remove the scanned paper and check the quality of the scanning results. This may require rearranging or flipping the paper to ensure image clarity and integrity. If the scanning results are not satisfactory, adjust the scanning parameters or rescan the paper.

[0081] T5: Continuous scanning. If there are multiple pages of paper archives, they need to be scanned page by page. Place each page on the scanning device in turn and scan according to the above steps. It is necessary to remove and place the paper between each scan to maintain a continuous scanning process.

[0082] T6: Save digital files. After all paper archives are scanned, save the digital files to the computer. Enter the file name and select the save path in the scanning software, and then click the save button to complete the save process.

[0083] T7: Backup and archiving. After digitization, it is necessary to back up the saved digital files to ensure data security. Then, archive the digital files into appropriate folders and classify and organize them for subsequent management and retrieval.

[0084] Embodiment 2

[0085] In Embodiment 1 of the present invention, emphasis is placed on illustrating how to select different digitization methods according to the security requirements and complexity of paper archives, and how to perform automated digitization and manual digitization is described in detail. However, during the digitization process, no detailed explanation on adjusting the regular backup time is given for data backup. This embodiment provides a method for adjusting the regular backup time to ensure data security and reliability and reduce backup costs.

[0086] In step Y2 of Embodiment 1, the confidentiality data C of the paper archive is obtained. The confidentiality data involves the types and quantities of sensitive information in the paper archive, such as personal identity information and financial data, etc. In this embodiment, it is used as an evaluation index for the importance of digital archive data. The confidentiality data is normalized to map it to the [0, 1] interval, and then converted into binary data containing 0 and 1 by comparing with the confidentiality threshold. Repeat the entire Y4 step in Embodiment 1, namely Y4.1 to Y4.5, to obtain the confidence level of the confidentiality data and label it as A(C), that is, only the confidence level of the confidentiality data is calculated. If A(C) is relatively high, it reflects a relatively high level of information on the confidentiality of sensitive information in the digital archive, meaning that the data importance of this digital archive is relatively high; conversely, the data importance of the digital archive is relatively low.

[0087] It should be noted that during the process of processing confidentiality data, normalization can use the general formula The setting of the requirement index threshold is determined according to professional knowledge in the relevant field and specific actual situations, and will not be elaborated here.

[0088] Furthermore, using the confidentiality data of paper archives to evaluate the data importance of the digital archives obtained after digitization of these archives is only an illustrative example of this application. It is also possible to use multiple data for comprehensive analysis, including the business value and usability of the data, etc., which will not be described in detail here.

[0089] Adjust the time of regular backups according to the data importance of digital archives and the backup storage cost. The backup storage cost can be obtained through existing processes, such as querying the pricing information of cloud service providers or referring to the prices of local storage devices. If you choose to use cloud storage as the backup storage solution, you can directly go to the official website or management console of the cloud service provider to query the pricing information of its storage services. On the pricing page or price calculator, enter the expected storage volume and usage period to obtain the corresponding storage cost. If you choose to use local storage devices (such as hard disks, tapes, etc.) as the backup storage solution, you can directly query the prices of relevant devices. You can find the prices of the required devices on the websites of online retailers, IT device suppliers, or local electronics stores, and calculate the storage cost based on the expected storage requirements. This will not be elaborated here.

[0090] The comprehensive analysis of the data importance of digital archives and the backup storage cost is as follows:

[0091] First, sort the observed values of data security and backup storage cost separately according to their magnitudes to obtain their ranks. If there are observed values with the same numerical value, take the average of their ranks. Calculate the difference di for the ranks of each pair of data points. The specific formula is rank(Ui) - (Vi), where Ui is the observed value of data security for the i-th data point. Vi is the observed value of backup storage cost for the i-th data point, and the observed value is the specific numerical value of each variable in the data. Then calculate its Spearman correlation coefficient. The specific formula is In the formula, ∑di 2 is the sum of the squares of all rank differences di, and n is the number of data points. The value range of the finally obtained Spearman correlation coefficient is between -1 and 1.

[0092] If ρ is close to 1, it indicates a strong monotonic positive correlation between data importance and backup storage cost. In this case, more frequent backups may be required, and the time of regular backups needs to be reduced. If ρ is close to -1, it indicates a strong monotonic negative correlation between data importance and backup storage cost. In this case, the backup frequency can be considered to be reduced, that is, the time of regular backups is increased. If ρ is close to 0, it indicates that there is almost no monotonic relationship between data importance and backup storage cost, and the time of regular backups can be kept unchanged. Specifically, a backup threshold can be set. For example, when ρ is greater than or equal to 0.5, it is marked as positive correlation, and at this time, the time of regular backups needs to be reduced; when ρ is less than or equal to -0.5, it is marked as negative correlation, and at this time, the time of regular backups needs to be increased; when -0.5 < ρ < 0.5, it is marked as no correlation, and at this time, the time of regular backups is kept unchanged.

[0093] By backing up important data in a timely manner and flexibly adjusting the backup frequency according to data changes, the risk of data loss can be minimized, and data security can be ensured. At the same time, dynamically adjusting the backup time can also avoid unnecessary storage costs and save resources. This flexibility enables organizations to better adapt to the changing business environment, improve backup efficiency, and effectively manage data security and cost-effectiveness.

[0094] Embodiment 3, a digital archive intelligent management system based on big data, includes a data acquisition module, a data processing module, and a data storage module;

[0095] The data acquisition module is used to obtain the security requirement data and complexity data of paper archives and send them to the data processing module to ensure the operation of the subsequent data processing module;

[0096] The data processing module is used to determine whether to automate the digitization of paper archives or perform manual digitization based on the data collected by the data acquisition module;

[0097] The data storage module is used to store all the data generated during the processing of the digital archive intelligent management system.

[0098] The above formulas are all dimensionless and take their numerical values for calculation. The formula is obtained by collecting a large amount of data and performing software simulation to obtain a formula that is closest to the actual situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.

[0099] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product.

[0100] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and invention constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0101] In addition, the functional modules in each embodiment of the present application can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0102] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the said claims.

[0103] Finally: The above description is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A digital archive intelligent management method based on big data, characterized in that, It includes the following steps: Step 1: Adjust and determine the digitalization method selection strategy according to the security requirements and complexity of paper archives; Step 2: Collect the analysis data required for the security requirements and complexity of paper archives, and perform calculations to obtain the security requirements and complexity of paper archives; Step 3: Determine the digitalization method according to the security requirements and complexity of paper archives, and start implementation; In Step 2, the specific process of calculating the security requirements of paper archives is as follows: Y1: Collect objective data; record the types and quantities of sensitive information in paper archives, and conduct regular integrity audits on paper archives; Y2: Computer confidentiality and integrity; Let N be the number of samples, i.e., the number of paper archives, Nc be the number of times an event occurs, and the confidential data be C. Then Integrity is the proportion of the content of the paper archives to the original data; Y3: Obtain security requirement index data; perform weighted calculations on the encryption and integrity of paper archives to obtain the security requirement index data of paper archives, label it as X, and judge X. If X is greater than or equal to the requirement index threshold, then let X = 1, indicating that the sample meets the security requirements; if X is less than the requirement index threshold, then let X = 0, indicating that the sample does not meet the security requirements; Y4: Calculate the confidence level of the security requirement index X; specifically, it includes the following steps: Y4.1: Calibrate the confidence level of the security requirement indicator as A(X), and collect N samples D = {d1, d2, d3,..., d N}; Each sample represents the observation result of the security requirement indicator X, and use the Bayesian method to calculate A(X); Y4.2: Select the prior probability; select the prior probability distribution P(X); Y4.3: Calculate the likelihood function; determine the likelihood function according to the sample data and the security requirement indicators; use the binomial distribution as the likelihood function; let p be the probability of meeting the security requirements, then the likelihood function can be expressed as P(D|X) = p k ×(1 - p) N-k ; where k is the number of samples that meet the security requirements, and N is the total number of samples; Y4.4: Calculate the posterior probability; calculate the posterior probability of the security requirement index according to Bayes' theorem; Y4.5: Calculate the confidence level according to the posterior probability; select the mean value of the posterior probability as the estimated value of the confidence level; In Step 3, define the obtained security requirements and complexity of paper archives as input variables, and divide them into different fuzzy sets; define the selected digitalization method as the output variable, and label automated digitalization as P1 and manual digitalization as P2; Formulate fuzzy rules to describe the influence of different input variables on the output variable; Perform fuzzy reasoning according to the fuzzy rules to determine the digitalization method.

2. The intelligent management method for digital archives based on big data according to claim 1, characterized in that: In Step 1, mainly determine whether to perform automated digitalization or manual digitalization on paper archives according to the security requirements and complexity of paper archives.

3. The intelligent management method for digital archives based on big data according to claim 1, characterized in that: In Step Y2, the confidentiality data C of paper archives is obtained. The confidentiality data involves the types and quantities of sensitive information in paper archives. Normalize the confidentiality data so that it is mapped to the interval [0,1], and then convert it into binary data containing 0 and 1 by comparing with the confidentiality threshold; repeat Step Y4, that is, Y4.1 to Y4.5, to obtain the confidence level of the confidentiality data and label it as A(C), that is, only calculate the confidence level of the confidentiality data; if A(C) is relatively high, the data importance of the digital archives is relatively high; otherwise, the data importance of the digital archives is relatively low.

4. The intelligent management method for digital archives based on big data according to claim 3, characterized in that: The comprehensive analysis of the data importance and backup storage cost of digital archives is as follows: First, sort the observed values of data security and backup storage costs by size respectively to obtain their ranks; if there are observed values with the same numerical value, take the average of their ranks; calculate the difference of the ranks of each pair of data points to obtain di, and the specific formula is rank(Ui) - (Vi), where Ui is the observed value of data security of the i-th data point; Vi is the backup storage cost observation value of the i-th data point, and the observation value is the specific value of each variable in the data; then calculate its Spearman correlation coefficient, and the specific formula is where ∑di 2 is the sum of the squares of all rank differences di, and n is the number of data points; If ρ is close to 1, more frequent backups are required, and at this time, the time of regular backups needs to be reduced; If ρ is close to -1, reduce the backup frequency, that is, increase the time of regular backups; if ρ is close to 0, keep the time of regular backups unchanged.

5. A digital archive intelligent management system based on big data, which is used to implement any one of the digital archive intelligent management methods based on big data described in claims 1-4, and is characterized in that: Including a data acquisition module, a data processing module, and a data storage module; The data acquisition module is used to obtain the security requirement data and complexity data of paper archives and send them to the data processing module to ensure the operation of the subsequent data processing module; The data processing module is used to determine whether to automate the digitization of paper archives or perform manual digitization according to the data collected by the data acquisition module; The data storage module is used to store all data generated during the processing of the digital archive intelligent management system.

Citation Information

Patent Citations

  • Automatic archive editing method

    CN104361111A

  • Scientific and technological archive authentication method based on traceable digital signature technology and intelligent management system

    CN117454440A