Electronic archive pre-archiving method based on big data
By collecting and analyzing parameters such as the keyword TF-IDF weighted value, content standardization rate, and document archiving time difference of electronic archives, and combining archiving integrity and manual intervention rate, a dynamic evaluation mechanism is established. This solves the problems of insufficient adaptability and high maintenance cost of existing electronic archive pre-archiving methods, and achieves efficient and accurate archive archiving.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU XIEZHENG INFORMATION TECH CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-28
AI Technical Summary
Existing electronic record pre-archiving methods lack data-driven quantitative evaluation and dynamic optimization mechanisms, leading to subjective and blind document selection, uncontrollable archiving quality, low record utilization value, low pre-archiving efficiency, and resource waste.
By collecting file characteristic parameters from the target data source, including keyword TF-IDF weighted values, content standardization rate, and file archiving time difference, the file characteristic characterization values are analyzed and compared with predetermined thresholds. Combined with the archiving integrity detection rate and the manual intervention rate, a dynamic evaluation mechanism is established to determine the preservation quality and reliability of pre-archiving, and automatically adjust the file characteristic thresholds to improve archiving accuracy and stability.
It enables precise initial screening and continuous monitoring, improves archiving accuracy and automation, reduces reliance on external manual intervention and rule maintenance, and ensures the accuracy and stability of archiving.
Smart Images

Figure CN121935211A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of document organization technology, and in particular to a method for pre-archiving electronic archives based on big data. Background Technology
[0002] Electronic document pre-archiving technology refers to the key technology that uses automated means to automatically identify, analyze, and assign preliminary classification tags and metadata to massive amounts of electronic documents before formal archiving. Its core application scenarios include document management systems and digital archives in various organizations, as well as cross-platform business data governance processes, aiming to achieve standardized, structured, and efficient retrieval of archives. However, existing technologies mainly rely on preset static rules and simple keyword matching. Static rules are difficult to adapt to the high diversity of archival content in terms of semantic depth, presentation, and dynamic business context, resulting in insufficient archiving accuracy and frequent misjudgments and omissions. Judgment models based on fixed thresholds lack self-learning and adaptive capabilities; when business forms or data characteristics change, model performance rapidly degrades, exhibiting poor adaptability. To maintain the basic effectiveness of the method, technical personnel need to continuously invest significant effort in expanding, adjusting, and maintaining the rule base, directly leading to high system operating costs and creating bottlenecks for the large-scale and intelligent application of the technology.
[0003] Chinese Patent Publication No. CN114817676A discloses an archive management system, which includes an electronic archive centralized management system. The electronic archive centralized management system comprises an electronic document pre-archiving system, an electronic document management system, and an electronic archive long-term preservation system. The electronic document pre-archiving system includes a data receiving module, an electronic document conversion module, a metadata capture module, a metadata conversion module, a metadata supplementation module, an electronic document supplementation scanning module, an archived data comparison module, a receiving list generation module, and an archived data submission module. The electronic document management system includes a user management module, an archive receiving module, a receiving detection module, an online archiving module, a data utilization module, and a data sharing module. The electronic archive long-term preservation system includes an archive storage module, a management module, an archive sharing module, and an archive retrieval module. This invention solves the technical problems of existing archive management systems, such as inconsistent archiving standards, low archiving efficiency, data isolation, and insufficient sharing.
[0004] Chinese Patent Publication No. CN118331927A discloses an electronic archive pre-archiving system based on an AI large-scale model, comprising: a format standardization module for converting a massive number of archived documents of permissible document formats into a standard format; a storage parameter determination module for determining the optimal storage architecture and optimal parallel processing threads for the current massive number of standard format documents; a document content error correction module for correcting the document content of the current massive number of standard format documents based on an AI text error correction large-scale model to obtain the current massive number of final documents; a file number generation module for generating an archive file number for each final document; and an archive task execution module for encrypting and transmitting the corresponding final documents of the current massive number of standard format documents to a digital archive system based on the archive file number of each final document, the optimal storage architecture and optimal parallel processing threads for the current massive number of standard format documents; thereby improving the subsequent archiving efficiency and archiving security of electronic documents.
[0005] Therefore, it is evident that the existing technology has the following problems: Existing electronic record pre-archiving methods lack data-driven quantitative evaluation and dynamic optimization mechanisms, leading to subjective and blind document selection and uncontrollable archiving quality. This results in low utilization value of electronic records, low pre-archiving efficiency, and resource waste. Summary of the Invention
[0006] To address this, the present invention provides a big data-based electronic archive pre-archiving method to overcome the problem that existing electronic archive pre-archiving methods lack quantitative evaluation driven by keyword TF-IDF weighted values, content standardization rates, and document archiving time difference data, leading to misjudgments of the timeliness of archive information quality, resulting in differential effects on the detection rate of archive integrity and the rate of manual intervention, and ultimately causing low document archiving efficiency.
[0007] To achieve the above objectives, the present invention provides a method for pre-archiving electronic archives based on big data, comprising: Collect file characteristic parameters of electronic archives to be archived within the historical period of the target data source; Analyze the file feature representation values based on the aforementioned file feature parameters; The difference between the document feature characterization value and the predetermined document feature characterization threshold is used to determine whether the pre-archiving preservation quality of the target electronic archive meets the standard. In response to the requirement that the pre-archiving preservation quality of the target electronic archives meets the standards, pre-archiving quality indicator parameters of the target electronic archives within the historical period are collected. Analyze the characterization values of the pre-archived quality indicators based on the aforementioned pre-archived quality indicator parameters; The reliability of the target electronic archive pre-archiving is determined based on the comparison between the pre-archiving quality index characterization value and the predetermined pre-archiving quality index characterization threshold. In response to the fact that the reliability of the target electronic archive pre-archiving does not meet the standard, a processing strategy is determined based on the difference between the pre-archiving quality indicator characterization value and the predetermined pre-archiving quality indicator characterization threshold. The target data sources include business systems, email systems, and collaboration platforms; The file feature parameters include keyword TF-IDF weighted value, content standardization rate, and file archiving time difference; The pre-archiving quality indicators include the archiving integrity detection rate and the manual intervention rate.
[0008] Furthermore, the process of analyzing file feature representation values using the file feature parameters includes: Collect the keyword TF-IDF weighted values, content standardization rate, and file archiving time difference of the files to be archived within the historical period of the target data source; The ratio of the keyword TF-IDF weighted value to the predetermined keyword TF-IDF weighted threshold is determined as the first data-limited characterization parameter; The ratio of the calculated content standardization rate to the predetermined content standardization rate threshold is determined as the second data constraint characterization parameter; The ratio of the predetermined file archiving time difference threshold to the file archiving time difference is determined as the third data-limited characterization parameter; The summation of the first data-limited characterization parameter, the second data-limited characterization parameter, and the third data-limited characterization parameter is determined as the file feature characterization value.
[0009] Furthermore, the process of determining whether the pre-archiving preservation quality of the target electronic archive meets the standard by using the difference between the document feature characterization value and the predetermined document feature characterization threshold includes: Calculate the difference between the file feature representation value and the predetermined file feature representation threshold; If the difference between the document feature characterization value and the predetermined document feature characterization threshold is greater than the predetermined difference threshold, then the pre-archiving preservation quality of the target electronic document is determined to meet the standard.
[0010] The process of determining whether the pre-archiving preservation quality of the target electronic archive does not meet the standard based on the difference between the document feature characterization value and the predetermined document feature characterization threshold includes: Calculate the difference between the file feature representation value and the predetermined file feature representation threshold; If the difference between the document feature representation value and the predetermined document feature representation threshold is less than or equal to the predetermined difference threshold, then the pre-archiving preservation quality of the target electronic document is determined to be non-compliant with the standard.
[0011] Furthermore, the process of analyzing the pre-archived quality index parameters and their representation values includes: Collect the archival integrity detection rate and manual intervention rate of target electronic archives within the historical period; The ratio of the predetermined archive integrity detection rate threshold to the archive integrity detection rate is the first quality limit characterization parameter. The ratio of the human intervention rate to the predetermined human intervention rate threshold is used as the second quality-limited characterization parameter. The summation of the first quality-limited characterization parameter and the second quality-limited characterization parameter is determined as the pre-archived quality index characterization value.
[0012] Furthermore, the process of determining whether the pre-archiving reliability of the target electronic archive meets the standard by comparing the pre-archiving quality indicator characterization value with the predetermined pre-archiving quality indicator characterization threshold includes: Extract the comparison results between the pre-archived quality indicator characterization values and the predetermined pre-archived quality indicator characterization thresholds; If the pre-archiving quality index value is less than the predetermined pre-archiving quality index threshold, then the pre-archiving reliability of the target electronic archive is determined to meet the standard.
[0013] Furthermore, the process of determining whether the reliability of the pre-archiving of the target electronic archive does not meet the standard by comparing the pre-archiving quality index characterization value with the predetermined pre-archiving quality index characterization threshold includes: Extract the comparison results between the pre-archived quality indicator characterization values and the predetermined pre-archived quality indicator characterization thresholds; If the pre-archiving quality indicator value is greater than or equal to the predetermined pre-archiving quality indicator threshold, then the pre-archiving reliability of the target electronic archive is determined to be non-compliant with the standard.
[0014] Furthermore, the process of determining the processing strategy includes: Extract the comparison results between the pre-archived quality indicator characterization values and the predetermined pre-archived quality indicator characterization thresholds; If the pre-archiving quality index characterization value is greater than or equal to the predetermined pre-archiving quality index characterization threshold, it is determined that the pre-archiving reliability of the target electronic archive does not meet the standard, and a processing strategy is determined to determine the adjustment range of the document feature characterization threshold.
[0015] Furthermore, the process of determining the adjustment range of the file feature representation threshold includes: Calculate the difference between the pre-archived quality indicator characterization value and the predetermined pre-archived quality indicator characterization threshold; The adjustment range of the file feature representation threshold is determined based on the difference between the pre-archived quality indicator representation value and the predetermined pre-archived quality indicator representation threshold.
[0016] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention provides a big data-based electronic archive pre-archiving method. By collecting the keyword TF-IDF weighted values, content standardization rate, and document archiving time difference of the electronic archives to be archived from the target data source, the method analyzes the document characteristic values, quantifies the necessity of archiving, and achieves accurate preliminary screening. It determines whether the pre-archiving preservation quality of the target electronic archives meets the standards. When the pre-archiving preservation quality meets the standards, it collects the archiving integrity detection rate and human intervention rate of the target electronic archives within a historical period to analyze the pre-archiving quality indicator values, continuously monitors the applicability of the pre-archiving method, and determines whether the pre-archiving reliability of the target electronic archives meets the standards. When the pre-archiving reliability of the target electronic archives does not meet the standards, it determines the processing strategy to determine the adjustment range of the document characteristic threshold. This solves the problems of insufficient adaptability and high maintenance costs of traditional static rule models. While improving archiving accuracy, automation level, and long-term system stability, it effectively reduces the dependence on external human intervention and continuous rule maintenance.
[0017] In particular, this invention selects three key parameters that respectively characterize the coreness of the content, the compliance of the format, and the timeliness of the document archiving time, namely the TF-IDF weighted value of the keyword, the content standardization rate, and the difference in document archiving time. After standardizing these parameters into ratios, they are weighted and summed. By comparing the difference between the document feature characterization value and the preset document feature characterization threshold, the invention reduces the occurrence of misjudgments caused by relying on a single rule or subjective experience, and significantly improves the accuracy of the pre-archiving screening process.
[0018] In particular, this invention selects two core indicators that directly reflect the operational effectiveness and efficiency of the pre-archiving method: the archiving integrity detection rate and the manual intervention rate. These indicators are then summed after being compared with their respective quality thresholds. By comparing the calculated pre-archiving quality indicator values with the preset archiving quality indicator thresholds, an applicability judgment standard is established, thereby improving archiving reliability.
[0019] In particular, this invention ensures the accuracy and stability of archiving by analyzing and calculating the difference between the pre-archiving quality index characterization value and the predetermined archiving quality index characterization threshold, and by accurately determining the increase of the document feature characterization threshold based on the difference and a predetermined adjustment coefficient, and reduces the reliance on external manual intervention and maintenance. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating the steps of the electronic archive pre-archiving method based on big data, as described in an embodiment of the present invention. Figure 2 This is a flowchart illustrating the steps involved in analyzing file feature representation values according to an embodiment of the present invention. Figure 3This invention provides a logic diagram for determining whether the pre-archiving preservation quality and effectiveness of target electronic archives meet the standards. Figure 4 This is a logic diagram for determining whether the pre-archiving reliability of target electronic archives meets the standards in an embodiment of the present invention. Detailed Implementation
[0021] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0022] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0023] Please see Figure 1 The diagram shows a flowchart of the steps in the electronic archive pre-archiving method based on big data according to an embodiment of the present invention. The present invention provides an electronic archive pre-archiving method based on big data, comprising: Step S1: Collect file characteristic parameters of the electronic archives to be archived within the historical period of the target data source; Step S2: Analyze the file feature representation value based on the file feature parameters; Step S3: Determine whether the pre-archiving preservation quality of the target electronic archive meets the standard based on the difference between the document feature characterization value and the predetermined document feature characterization threshold. Step S4: In response to the fact that the pre-archiving preservation quality of the target electronic archive meets the standard, collect the pre-archiving quality index parameters of the target electronic archive within the historical period; and analyze the pre-archiving quality index characterization value based on the pre-archiving quality index parameters. Step S5: Determine whether the reliability of the target electronic archive pre-archiving meets the standard based on the comparison result between the pre-archiving quality index characterization value and the predetermined pre-archiving quality index characterization threshold. Step S6: In response to the target electronic archive's pre-archiving reliability not meeting the standard, a processing strategy is determined based on the difference between the pre-archiving quality indicator characterization value and the predetermined pre-archiving quality indicator characterization threshold.
[0024] In this embodiment, the data acquisition process is implemented using a distributed technical architecture to meet the requirements of massive volume, multiple sources, and real-time processing in a big data environment. By employing the Apache Kafka distributed message queue and log collection system, the system captures the raw streams of electronic archive data continuously generated from various business systems, mail servers, and collaboration platforms in real time. This raw stream includes file entities and metadata and is uniformly ingested into the big data processing platform in a high-throughput, low-latency manner. The data is persistently stored in cloud-native object storage within the Hadoop Distributed File System, thereby constructing a horizontally scalable unified data lake capable of accommodating both historical and real-time data, providing a stable and consistent data foundation for subsequent feature calculations and analysis.
[0025] In this embodiment, by collecting multi-dimensional file feature parameters such as the keyword TF-IDF weighted value, content standardization rate, and file archiving time difference of the target data source electronic archives to be archived, the file feature characterization value is analyzed to determine whether the pre-archiving preservation quality effectiveness of the target electronic archives meets the standard. When the pre-archiving preservation quality effectiveness of the target electronic archives meets the standard, the archiving integrity detection rate and human intervention rate of the target electronic archives within the historical period are collected to analyze the pre-archiving quality indicator characterization value and determine whether the pre-archiving reliability of the target electronic archives meets the standard. When the pre-archiving reliability of the target electronic archives does not meet the standard, a processing strategy is determined to determine the adjustment range of the file feature characterization threshold. This improves the accuracy and stability of archiving while effectively reducing the dependence on external human intervention and continuous rule maintenance.
[0026] Please see Figure 2 The diagram shows a flowchart illustrating the steps of analyzing file feature representation values according to an embodiment of the present invention. The process of analyzing file feature representation values based on file feature parameters according to the present invention includes: Step S21: Collect the keyword TF-IDF weighted value, content standardization rate, and file archiving time difference of the files to be archived within the historical period of the target data source; Step S22: Calculate the ratio of the keyword TF-IDF weighted value to the predetermined keyword TF-IDF weighted threshold and determine it as the first data constraint characterization parameter; calculate the ratio of the content standardization rate to the predetermined content standardization rate threshold and determine it as the second data constraint characterization parameter; calculate the ratio of the predetermined file archiving time difference threshold to the file archiving time difference and determine it as the third data constraint characterization parameter. Step S23: The first data-limited characterization parameter, the second data-limited characterization parameter, and the third data-limited characterization parameter are summed to determine the file feature characterization value.
[0027] In this embodiment, heterogeneous parameters are transformed into superimposed standardized representation parameters by calculating the ratios of keyword TF-IDF weighted values, content standardization rate, and file archiving time difference to their respective predetermined thresholds. These three are then summed to generate a single file feature representation value. This method achieves the fusion and standardization of multi-dimensional evaluation indicators, enabling the core content, format compliance, and timeliness to be comprehensively reflected within the same quantitative system. A dynamic evaluation mechanism based on relative ratios is established, ensuring that the evaluation results depend not only on the absolute values of the parameters but also on preset quality benchmarks, thus improving the system's adaptability and consistency. Ultimately, it provides a precise and calculable comprehensive basis for subsequent archiving necessity judgments, supporting the automation and intelligent decision-making of the entire pre-archiving process.
[0028] In this embodiment, the formula for calculating the TF-IDF weighted value of the keyword is: TF-IDF = TF × IDF Where TF = the number of times the keyword appears in the file / the total number of words in the file; IDF = the total number of documents in the log document set / the number of documents containing the keyword + 1 In this embodiment, the formula for calculating the content standardization rate is: Content standardization rate = Number of items that meet standardization requirements / Total number of items inspected The total number of inspection items refers to the sum of all inspection items covered in the standardized inspection checklist prepared in advance based on the business attributes and archiving standards of the archives when conducting accuracy and standardization checks on a single pre-archived document. These inspection items include whether the pre-archived document contains the necessary signature pages, whether the attachments are complete, and whether the text structure and layout conform to the established template.
[0029] In this embodiment, the formula for calculating the file archiving time difference is: File archiving time difference = File archiving timestamp - File creation timestamp Please see Figure 3 As shown, this is a logic diagram for determining whether the pre-archiving preservation quality of a target electronic file meets the standard, according to an embodiment of the present invention. The process of determining whether the pre-archiving preservation quality of a target electronic file meets the standard based on the difference between the file feature characterization value and the predetermined file feature characterization threshold includes: Calculate the difference between the file feature representation value and the predetermined file feature representation threshold; If the difference between the document feature characterization value and the predetermined document feature characterization threshold is greater than the predetermined difference threshold, then the pre-archiving preservation quality of the target electronic archive is determined to meet the standard. If the difference between the document feature representation value and the predetermined document feature representation threshold is less than or equal to the predetermined difference threshold, then the pre-archiving preservation quality of the target electronic document is determined to be non-compliant with the standard.
[0030] In this embodiment, the predetermined document feature representation threshold is obtained in advance. All document feature representation values of the target electronic archive within 3 months are collected, and their average value is calculated as the document feature representation threshold. The document feature representation threshold is selected within the range [3.05, 3.25], and the preferred value in this embodiment is 3.15.
[0031] In this embodiment, the predetermined difference threshold is obtained in advance. The difference between the total document feature representation values of the target electronic file within 3 months and the predetermined document feature representation threshold is calculated, and the average value is calculated as the difference threshold. The predetermined difference threshold is selected within the range [0.15, 0.35], and is preferably 0.20 in this embodiment.
[0032] In this embodiment, by introducing a dual difference comparison judgment logic, the initial difference between the file feature characterization value and a predetermined threshold is compared with a predetermined difference threshold, thereby achieving multiple beneficial effects. This not only completes the upgrade from simple binary judgment to quantitative degree evaluation, providing a data foundation for subsequent differential processing, but also enhances the accuracy of screening decisions and system stability by setting confidence intervals. The precise difference result output by this logic provides a key quantitative input for the adaptive feedback closed loop of the entire pre-archiving system, enabling subsequent parameter optimization and strategy adjustment to be accurately based on objective performance deviations, thus laying the core data foundation for the intelligent evolution of the system.
[0033] Specifically, the process of analyzing the pre-archived quality index parameters and their representation values includes: Collect the archival integrity detection rate and manual intervention rate of target electronic archives within the historical period; The ratio of the predetermined archive integrity detection rate threshold to the archive integrity detection rate is the first quality limit characterization parameter. The ratio of the human intervention rate to the predetermined human intervention rate threshold is used as the second quality-limited characterization parameter. The summation of the first quality-limited characterization parameter and the second quality-limited characterization parameter is determined as the pre-archived quality index characterization value.
[0034] In this embodiment, a standardized comprehensive assessment of the health status of the pre-archiving method is achieved by calculating and summing the ratios of the two core performance indicators, namely the archiving integrity detection rate and the manual intervention rate, to their corresponding thresholds. Through the fusion of the two indicator ratios, a multi-dimensional quantitative diagnostic model that unifies effectiveness and efficiency is established, enabling the assessment results to objectively reflect both the reliability and economy of the method. By adopting the standardization of relative ratios, the dimensional differences of heterogeneous indicators are eliminated, and a performance benchmark that can be compared horizontally is constructed with a baseline as a reference. The generated singular and quantifiable quality indicator characterization values provide accurate, interpretable, and logically consistent data-driven signals for subsequent applicability determination and closed-loop optimization decisions, becoming a key hub for the entire system to evolve from static monitoring to dynamic adaptive governance.
[0035] In this embodiment, the formula for calculating the archive integrity detection rate is: Archive integrity detection rate = Number of archives to be archived correctly identified by the system / Total number of archives that should actually be archived In this embodiment, the formula for calculating the human intervention rate is: Human intervention rate = Number of files requiring human intervention / Total number of files processed by the system Please see Figure 4 As shown, this is a logic diagram for determining whether the pre-archiving reliability of a target electronic archive meets the standard according to an embodiment of the present invention. The process of determining whether the pre-archiving reliability of a target electronic archive meets the standard based on the comparison result of the pre-archiving quality index characterization value and the predetermined pre-archiving quality index characterization threshold includes: Extract the comparison results between the pre-archived quality indicator characterization values and the predetermined pre-archived quality indicator characterization thresholds; If the pre-archiving quality index characterization value is less than the predetermined pre-archiving quality index characterization threshold, then the pre-archiving reliability of the target electronic archive is determined to meet the standard. If the pre-archiving quality indicator value is greater than or equal to the predetermined pre-archiving quality indicator threshold, then the pre-archiving reliability of the target electronic archive is determined to be non-compliant with the standard.
[0036] In this embodiment, the predetermined pre-archiving quality index characterization threshold is obtained in advance. All pre-archiving quality index characterization values of the target electronic archives within 3 months are collected, and their average value is calculated as the pre-archiving quality index characterization threshold. The pre-archiving quality index characterization threshold is selected within the range [2.05, 2.15], and is preferably 2.10 in this embodiment.
[0037] In this embodiment, by directly comparing the comprehensive pre-archiving quality indicator characterization value with the preset pre-archiving quality indicator characterization threshold, if the pre-archiving quality indicator characterization value is less than the preset pre-archiving quality indicator characterization threshold, the business objective of high detection rate and low manual intervention rate is applied, thereby ensuring the core decision-making mechanism that evolves from static operation to dynamic autonomy.
[0038] Specifically, the process of determining the processing strategy includes: Extract the comparison results between the pre-archived quality indicator characterization values and the predetermined pre-archived quality indicator characterization thresholds; If the pre-archiving quality index characterization value is greater than or equal to the predetermined pre-archiving quality index characterization threshold, it is determined that the pre-archiving reliability of the target electronic archive does not meet the standard, and a processing strategy is determined to determine the adjustment range of the document feature characterization threshold.
[0039] In this embodiment, the predetermined difference threshold is obtained in advance. The difference between the pre-archiving quality index characterization values of all target electronic archives within 3 months of stable operation and the predetermined pre-archiving quality index characterization threshold is collected, and the average value is calculated as the difference threshold. The predetermined difference threshold is selected in the range [0.15, 0.35], and is preferably 0.20 in this embodiment.
[0040] In this embodiment, by directly using the difference between the pre-archived quality index characterization value and the preset threshold as the basis for determining the adjustment range of the file feature characterization threshold, a closed-loop execution from performance diagnosis to precise control is achieved, completing a key step in the system's intelligent self-healing. By tightening the screening conditions, the archiving quality of high-value archives is prioritized. At the same time, this quantification mechanism provides a stable and controllable foundation for the entire adaptive process, ensuring that the adjustment intensity matches the degree of performance deviation, thus laying the core foundation for the system's continuous self-optimization and robust operation.
[0041] Specifically, the process of determining the adjustment range of the file feature representation threshold includes: Calculate the difference between the pre-archived quality indicator characterization value and the predetermined pre-archived quality indicator characterization threshold; The adjustment range of the file feature representation threshold is determined based on the difference between the pre-archived quality indicator representation value and the predetermined pre-archived quality indicator representation threshold.
[0042] In this embodiment, the difference between the pre-archived quality index characterization value and the predetermined pre-archived quality index characterization threshold is calculated to quantify the quality deviation; the adjustment range is positively correlated with the quality deviation value, that is, the larger the quality deviation, the greater the upward adjustment range of the document feature characterization threshold.
[0043] In this embodiment, when the pre-archiving quality is determined to be substandard, the adjustment range of the previous file screening threshold is automatically and accurately calculated, thereby realizing the dynamic self-optimization of the pre-archiving screening standard. The quality assessment results are directly and quantitatively transformed into parameter adjustment, which not only enables the system to have the ability to continuously self-calibrate and improve performance, but also fundamentally transforms the management strategy into a stable and repeatable automated operation logic, significantly improving the long-term accuracy and adaptability of pre-archiving management.
[0044] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to the specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A method for pre-archiving electronic records based on big data, characterized in that, include: Collect file characteristic parameters of electronic archives to be archived within the historical period of the target data source; Analyze the file feature representation values based on the aforementioned file feature parameters; The difference between the document feature characterization value and the predetermined document feature characterization threshold is used to determine whether the pre-archiving preservation quality of the target electronic archive meets the standard. In response to the requirement that the pre-archiving preservation quality of the target electronic archives meets the standards, pre-archiving quality indicator parameters of the target electronic archives within the historical period are collected. Analyze the characterization values of the pre-archived quality indicators based on the aforementioned pre-archived quality indicator parameters; The reliability of the target electronic archive pre-archiving is determined based on the comparison between the pre-archiving quality index characterization value and the predetermined pre-archiving quality index characterization threshold. In response to the failure of the target electronic archive pre-archiving reliability to meet the standard, a processing strategy is determined based on the difference between the pre-archiving quality indicator characterization value and the predetermined pre-archiving quality indicator characterization threshold, so as to determine the adjustment range of the document feature characterization threshold. The target data sources include business systems, email systems, and collaboration platforms; The file feature parameters include keyword TF-IDF weighted value, content standardization rate, and file archiving time difference; The pre-archiving quality indicators include the archiving integrity detection rate and the manual intervention rate.
2. The method for pre-archiving electronic archives based on big data according to claim 1, characterized in that, The process of analyzing file feature representation values based on the aforementioned file feature parameters includes: Collect the keyword TF-IDF weighted values, content standardization rate, and file archiving time difference of the files to be archived within the historical period of the target data source; The ratio of the keyword TF-IDF weighted value to the predetermined keyword TF-IDF weighted threshold is determined as the first data-limited characterization parameter; The ratio of the calculated content standardization rate to the predetermined content standardization rate threshold is determined as the second data constraint characterization parameter; The ratio of the predetermined file archiving time difference threshold to the file archiving time difference is determined as the third data-limited characterization parameter; The summation of the first data-limited characterization parameter, the second data-limited characterization parameter, and the third data-limited characterization parameter is determined as the file feature characterization value.
3. The method for pre-archiving electronic archives based on big data according to claim 2, characterized in that, The process of determining whether the pre-archiving preservation quality of the target electronic archive meets the standard based on the difference between the document feature characterization value and the predetermined document feature characterization threshold includes: Calculate the difference between the file feature representation value and the predetermined file feature representation threshold; If the difference between the document feature characterization value and the predetermined document feature characterization threshold is greater than the predetermined difference threshold, then the pre-archiving preservation quality of the target electronic document is determined to meet the standard.
4. The method for pre-archiving electronic archives based on big data according to claim 3, characterized in that, The process of determining whether the pre-archiving preservation quality of a target electronic archive does not meet the standard based on the difference between the document feature characterization value and the predetermined document feature characterization threshold includes: Calculate the difference between the file feature representation value and the predetermined file feature representation threshold; If the difference between the document feature representation value and the predetermined document feature representation threshold is less than or equal to the predetermined difference threshold, then the pre-archiving preservation quality of the target electronic document is determined to be non-compliant with the standard.
5. The method for pre-archiving electronic archives based on big data according to claim 4, characterized in that, The process of analyzing the pre-archived quality indicator characterization values based on the pre-archived quality indicator parameters includes: Collect the archival integrity detection rate and manual intervention rate of target electronic archives within the historical period; The ratio of the predetermined archive integrity detection rate threshold to the archive integrity detection rate is the first quality limit characterization parameter. The ratio of the human intervention rate to the predetermined human intervention rate threshold is used as the second quality-limited characterization parameter. The summation of the first quality-limited characterization parameter and the second quality-limited characterization parameter is determined as the pre-archived quality index characterization value.
6. The method for pre-archiving electronic archives based on big data according to claim 5, characterized in that, The process of determining whether the pre-archiving reliability of the target electronic archives meets the standard based on the comparison results between the pre-archiving quality indicator characterization value and the predetermined pre-archiving quality indicator characterization threshold includes: Extract the comparison results between the pre-archived quality indicator characterization values and the predetermined pre-archived quality indicator characterization thresholds; If the pre-archiving quality index value is less than the predetermined pre-archiving quality index threshold, then the pre-archiving reliability of the target electronic archive is determined to meet the standard.
7. The method for pre-archiving electronic archives based on big data according to claim 6, characterized in that, The process of determining whether the reliability of the target electronic archive pre-archiving does not meet the standard based on the comparison result of the pre-archiving quality index characterization value and the predetermined pre-archiving quality index characterization threshold includes: Extract the comparison results between the pre-archived quality indicator characterization values and the predetermined pre-archived quality indicator characterization thresholds; If the pre-archiving quality indicator value is greater than or equal to the predetermined pre-archiving quality indicator threshold, then the pre-archiving reliability of the target electronic archive is determined to be non-compliant with the standard.
8. The method for pre-archiving electronic archives based on big data according to claim 7, characterized in that, The process of determining the processing strategy includes: Extract the comparison results between the pre-archived quality indicator characterization values and the predetermined pre-archived quality indicator characterization thresholds; If the pre-archiving quality index characterization value is greater than or equal to the predetermined pre-archiving quality index characterization threshold, it is determined that the pre-archiving reliability of the target electronic archive does not meet the standard, and a processing strategy is determined to determine the adjustment range of the document feature characterization threshold.
9. The method for pre-archiving electronic archives based on big data according to claim 8, characterized in that, The process of determining the adjustment range of the file feature characterization threshold includes: Calculate the difference between the pre-archived quality indicator characterization value and the predetermined pre-archived quality indicator characterization threshold; The adjustment range of the file feature representation threshold is determined based on the difference between the pre-archived quality indicator representation value and the predetermined pre-archived quality indicator representation threshold.
Citation Information
Patent Citations
Archive management system
CN114817676A
Electronic archive pre-filing system based on AI large model
CN118331927A