Digital archive intelligent information mining method
By evaluating data quality through quantitative indicators, automatically screening high-quality data sets and training deep learning models, the problem of low manual screening efficiency is solved and the model accuracy and decision-making accuracy are improved.
Patent Information
- Application Number
- CN202510445523.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-10
AI Technical Summary
In the existing intelligent information mining methods for digital archives, manual screening data is of low quality, low efficiency and cannot effectively ensure data quality, which affects the model output accuracy and decision-making accuracy.
Data quality is evaluated through consistency indicators, integrity indicators and importance indexes, weighted thresholds are automatically generated, high-quality data sets are automatically screened, and deep learning models are used for training and prediction.
It improves the scientificity and accuracy of data screening, significantly improves the accuracy and stability of model output, and ensures the scientific and rationality of business decisions.
Smart Images

Figure CN120353846A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of archive information processing, and in particular to a digital archive intelligent information mining method. Background Art
[0002] Intelligent information mining of digital archives is a process of automatically extracting valuable information from a large number of digital archives using artificial intelligence and data mining technology. This method can help organizations and individuals manage and utilize large data sets more effectively, thereby improving work efficiency, promoting decision-making, and supporting research activities. Existing intelligent information mining of digital archives usually follows a series of steps to process and analyze data to extract valuable information. The following are specific implementation steps and their target results: First, collect raw data from various sources such as databases and file systems, and clean them, such as removing noise, filling missing values, and converting formats to ensure data quality. Then define and extract relevant features according to the problem requirements. For example, in text mining, it may include TF-IDF weights and word frequency statistics; in image processing, it may be color histograms, edge detection, etc. Then, select a suitable machine learning or deep learning model based on the nature of the task, and train the model using a labeled training set. This step usually involves the process of parameter tuning. Then apply the trained model to new data to generate prediction results or complete specific tasks such as classification and clustering. At the same time, evaluate the model performance through methods such as cross-validation, and finally support decision making based on the obtained results, that is, prediction results; Although the existing digital archive intelligent information mining methods have achieved certain results, they still face a series of challenges in practical applications, especially in terms of data quality: Data quality issues: In order to ensure the accuracy of subsequent model output, in addition to reciprocating and iterative training of the model, there is also a feasible measure in the existing technology, that is, to screen the data to be trained to ensure that the application of high-quality data can further improve the accuracy of the model output, so as to get closer to the expected goal. However, when screening data, manual labeling and screening are often used. In addition, the intervention of domain experts can assist in data screening. Although this can meet the expected data quality requirements to a certain extent, it is obvious that this manual screening mode requires a lot of manpower input for manual labeling, so the efficiency is low, and the screening method based on human subjective factors cannot effectively ensure data quality to a certain extent; therefore, the existing method of using manual mode to screen data is not only inefficient, but also still cannot effectively ensure data quality, which is easy to affect the output accuracy of subsequent models, and then easily affect the accuracy and scientificity of the final decision-making.
[0003] Therefore, there is an urgent need in the prior art for a technical solution for an intelligent information mining method for digital archives. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides an intelligent information mining method for digital archives, specifically including the following steps: Step S1: Obtain a digital archive data set to be processed, and preprocess the digital archive data set to generate a preprocessed digital archive data set; Step S2: Extract all feature columns from the preprocessed digital archive data set. Each feature column contains at least two data points. Respectively obtain the consistency index, integrity index of all data points included in each feature column, and the importance index of each data point, and obtain the comprehensive importance index of all data points based on the importance index of each data point. Obtain the comprehensive weight of all data points included in each feature column according to the consistency index, integrity index, and comprehensive importance index; Step S2a1: Calculate the mean value of all data points included in each feature column, and calculate the squared deviation of each data point from the mean value based on the mean value; Step S2a2: Based on the squared deviation of each data point from the mean value, obtain the mean value of the squared deviation; Step S2a3: Perform a square root operation on the mean value of the squared deviation to obtain the standard deviation of all data points included in each feature column; Step S2a4: Convert the standard deviation of all data points included in each feature column into a consistency score, and perform a standardization process on the consistency score to obtain the consistency index of all data points included in each feature column; Step S2b1: Count the number of all data points included in each feature column; Step S2b2: Detect the missing values of each data point included in each feature column, and count the number of data points with missing values in each feature column; Step S2b3: According to the number of all data points included in each feature column and the number of data points with missing values in each feature column, obtain the missing value ratio of all data points included in each feature column; Step S2b4: Convert the missing value ratio of all data points included in each feature column into an integrity score, and perform a standardization process on the integrity score to obtain the integrity index of all data points included in each feature column; Step S2c1: Analyze the importance degree of each data point included in each feature column, and assign an importance score to each data point according to the analysis result; Step S2c2: Perform normalization processing on the importance scores of all data points included in each feature column to obtain the importance index of each data point included in each feature column; Step S2c3: Synthesize the importance indices of each data point included in each feature column and perform averaging processing to obtain the comprehensive importance index of all data points included in each feature column; Step S2d1: Assign initial weights to all data points included in each feature column according to the comprehensive importance index; Step S2d2: Obtain the secondary weights of all data points included in each feature column based on the consistency index and integrity index of all data points included in each feature column; Step S2d3: Adjust the initial weights according to the secondary weights to obtain the comprehensive weights of all data points included in each feature column; Step S3: Obtain the weighted threshold of all data points included in each feature column based on the consistency index, integrity index, importance index, comprehensive weight, and comprehensive importance index; Among them, the calculation formula for obtaining the weighted threshold of all data points included in each feature column is: ; Among them, T i represents the weighted threshold of all data points included in the i-th feature column; represents the consistency index of all data points included in the i-th feature column; represents the integrity index of all data points included in the i-th feature column; E i represents the comprehensive importance index of all data points included in the i-th feature column; O ij represents the importance index of the j-th data point included in the i-th feature column; W i represents the comprehensive weight of all data points included in the i-th feature column; N i represents the number of data points in the i-th feature column; α, β, γ, δ respectively represent the weight coefficients of the consistency index, integrity index, comprehensive importance index, and importance index; Step S4: Use the weighted threshold of all data points included in each feature column to determine the importance index of each data point included in each feature column to obtain a high-quality data set, including: Step S41: Perform normalization processing on the weighted threshold of all data points included in each feature column and the importance index of each data point included in each feature column; Step S42: Use the normalized weighted threshold to determine the importance index of each data point; Step S43: Determine the importance index of the current data point in the current feature column according to the weighted threshold. If the importance index of the current data point in the current feature column is greater than or equal to the weighted threshold, retain the current data point; if the importance index of the current data point in the current feature column is less than the weighted threshold, eliminate the current data point; until each data point included in each feature column has been determined by the weighted threshold; Step S44: Define the current feature column after determination as a new feature column, and integrate all the obtained new feature columns to form a high-quality data set; Step S5: Use the high-quality data set to train the deep learning model, predict the digital archive data to be processed according to the trained deep learning model to obtain a prediction result, and make a business decision according to the prediction result.
[0005] The embodiments of the present invention have the following technical effects: The present invention evaluates data quality through a series of quantitative indicators including consistency indicators, integrity indicators, importance indexes, and comprehensive importance indexes, effectively reducing the subjective judgment in manual screening, making the data screening process more objective, reliable, and scientific. At the same time, by combining the quantitative indicators with the comprehensive weight, a data-driven weighted threshold is automatically generated. The application of this weighted threshold can not only quickly process large-scale data sets, thereby significantly improving the speed and efficiency of data screening, reducing the need for a large amount of manual annotation of human resources, but also effectively improving the scientificity and accuracy of data screening, that is, it can effectively screen out high-quality data sets and apply them to the training of subsequent deep learning models to significantly improve the output accuracy and stability of the model, ensure that more accurate prediction results can be obtained, and based on this more accurate prediction result, more scientific and reasonable business decisions can be made. Description of the Drawings
[0006] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0007] Figure 1 It is a flowchart of a digital archive intelligent information mining method provided by an embodiment of the present invention. Specific Embodiments
[0008] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be described clearly and completely below. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope protected by the present invention.
[0009] Embodiment 1: As Figure 1 shown, the present invention provides a digital archive intelligent information mining method, including the following steps: Step S1, obtain a digital archive data set to be processed, and preprocess the digital archive data set to generate a preprocessed digital archive data set; Regarding the acquisition sources of the digital archive data set, database: extract digital archive data from a relational database inside the organization such as a NoSQL database; file system: read stored digital archive files from a local file system or a network file system such as NAS, SAN, and these files can be in formats such as CSV, Excel, PDF, images, etc.; cloud storage: download digital archive data from a cloud storage service such as Amazon S3, Google Cloud Storage; API interface: obtain real-time or historical digital archive data from an external system through a RESTful API or other Web service interfaces; Regarding the acquisition methods of the digital archive data set, batch export: use a database management tool such as SQLDeveloper to batch export data tables to files; ETL tool: use an ETL (Extract, Transform, Load) tool such as Talend to extract data from multiple data sources and integrate them into a unified data warehouse; manual upload: users upload local files to a specified server directory through a graphical interface; automated script: write a Shell script or use a task scheduling tool such as Cron to regularly download data files from cloud storage; Regarding the steps for preprocessing the digital archive data set, they include: data cleaning, data integration, and data transformation, etc.; among them, data cleaning includes: removing noise, handling missing values, and format conversion; data integration includes: data merging and data deduplication; data transformation includes: feature engineering processing; Regarding removing noise, if it is text data, then remove redundant spaces, special characters, HTML tags, etc.; if it is numerical data, then detect and correct outliers, such as numerical values outside a reasonable range; if it is image data, then remove the noise in the image to improve the image quality; Regarding the handling of missing values, including identifying missing values: checking whether each field has null values or default values such as "N / A", "NULL"; filling missing values: for numerical data, the mean, median or interpolation method can be used for filling; for categorical data, the mode can be used for filling; deleting records with severe missing values: if the proportion of missing values in certain records is too high, these records can be selected for deletion; Regarding format conversion, including unifying all text to lowercase, removing punctuation marks, performing stemming or lemmatization; unifying date and time fields to a standard format such as YYYY-MM-DD HH:MM:SS; ensuring that all numerical fields have the same type such as integer, floating point; converting all images to a unified size and format such as JPEG format of 224x224 pixels.
[0010] Regarding data merging, including table joining, joining data from different tables through a common key such as customer ID, order number to generate a comprehensive table; file merging, merging data from multiple files into one file, for example, merging multiple CSV files into a large CSV file; Regarding data deduplication, including duplicate record detection: identifying and deleting exactly the same records; approximate duplicate record handling: for records with similar but not exactly the same content, deduplication can be performed through fuzzy matching technology; Regarding feature engineering, including feature selection: selecting relevant feature columns according to the problem requirements, removing irrelevant or redundant features for subsequent analysis; feature construction: constructing new features based on existing features, such as calculating derivative features like age, transaction frequency, etc.
[0011] Step S2: Extract all feature columns from the preprocessed digital archive dataset. Each feature column contains at least two data points. Respectively obtain the consistency index, integrity index of all data points contained in each feature column and the importance index of each data point, and obtain the comprehensive importance index of all data points based on the importance index of each data point. Obtain the comprehensive weight of all data points contained in each feature column according to the consistency index, integrity index and comprehensive importance index; Steps S2a1 to S2a4 mainly expound on how to obtain the consistency index of all data points contained in each feature column; Step S2a1: Calculate the mean of all data points contained in each feature column, and calculate the square of the deviation of each data point from the mean based on the mean; Calculate the mean of all data points contained in each feature column. The calculation formula is as follows: ; Among them, represents the mean of all data points contained in the i-th feature column; represents the number of data points in the \(i\)-th feature column; represents the value of the \(j\)-th data point in the \(i\)-th feature column; Calculate the squared deviation of each data point from the mean, and the calculation formula is as follows: ; where, represents the squared deviation of each data point from the mean; Step S2a2, based on the squared deviation of each data point from the mean, obtain the mean of the squared deviations; Obtain the mean of the squared deviations, and the calculation formula is as follows: ; where, represents the average value of the squared deviations of all data points; Step S2a3, take the square root of the mean of the squared deviations to obtain the standard deviation of all data points included in each feature column; Step S2a4, convert the standard deviation of all data points included in each feature column into a consistency score, and perform standardization processing on the consistency score to obtain the consistency index of all data points included in each feature column ; When performing the conversion of the consistency score, convert the standard deviation into a consistency score. Generally, the smaller the standard deviation, the higher the consistency. The purpose of the standardization processing is mainly to normalize the consistency score to the interval [0,1]; the formula used is as follows: ; where, represents the standard deviation of the \(i\)-th feature column, and represent the minimum value and the maximum value of the standard deviation respectively; Steps S2b1 to S2b4 mainly elaborate on how to obtain the integrity index of all data points included in each feature column; Step S2b1, count the number of all data points included in each feature column; Step S2b2, detect the missing values of each data point included in each feature column, and count the number \(M\) of data points with missing values in each feature column i ; Step S2b3, according to the number of all data points included in each feature column and the number of data points with missing values in each feature column, obtain the missing value ratio of all data points included in each feature column; the calculation formula is as follows: ; where \(P\) iRepresents the missing value ratio of the i-th feature column; It should be noted that when calculating the missing value ratio, first, the missing value information recorded in the missing value processing step during the preprocessing process should be used to count the number of missing values in each feature column.
[0012] Step S2b4: Convert the missing value ratio of all data points included in each feature column into an integrity score, and perform standardization processing on the integrity score to obtain the integrity index of all data points included in each feature column; Convert the missing value ratio into an integrity score. Generally, the lower the missing value ratio, the higher the integrity, and the integrity score needs to be normalized to the interval [0, 1]; the calculation formula is as follows: ; Where, Represents the integrity index of all data points included in the i-th feature column, I i Represents the integrity score of all data points included in the i-th feature column, I min And I max Represent the minimum and maximum values of the integrity score respectively; Steps S2c1 to S2c3 mainly focus on elaborating how to obtain the comprehensive importance index of all data points included in each feature column; Step S2c1: Analyze the importance degree of each data point included in each feature column, and assign an importance score L to each data point according to the analysis result ij ; It should be noted that when analyzing the importance degree of each data point included in each feature column, a feature selection algorithm such as random forest is used to analyze the importance degree of each data point and assign an importance score. This analysis process is a relatively existing analysis method and will not be elaborated one by one here; Step S2c2: Perform standardization processing on the importance scores of all data points included in each feature column to obtain the importance index of each data point included in each feature column; Normalize the importance score of each data point to the interval [0, 1], and the calculation formula is as follows: ; Where, O ij Represents the importance index of the j-th data point included in the i-th feature column, L max 、L min Represent the maximum and minimum values of the importance scores of all data points respectively; Step S2c3: Integrate the importance indexes of each data point included in each feature column and perform averaging processing to obtain the comprehensive importance index of all data points included in each feature column; Steps S2d1 to S2d3 mainly expound how to obtain the comprehensive weight of all data points included in each feature column according to the consistency index, integrity index and comprehensive importance index; Step S2d1: According to the comprehensive importance index, assign an initial weight to all data points in the i-th feature column included in each feature column ; Step S2d2: According to the consistency index of all data points included in each feature column and the integrity index of all data points included in each feature column, obtain the secondary weight of all data points included in each feature column; The calculation formula is as follows: ; Among them, represents the secondary weight of all data points in the i-th feature column; Step S2d3: Adjust the initial weight according to the secondary weight to obtain the comprehensive weight of all data points included in each feature column ; The calculation formula is as follows: ; Step S3: According to the consistency index, integrity index, importance index, comprehensive weight and comprehensive importance index, obtain the weighted threshold of all data points included in each feature column; Among them, the calculation formula for obtaining the weighted threshold of all data points included in each feature column is: ; Among them, T i represents the weighted threshold of all data points included in the i-th feature column; C i represents the consistency index of all data points included in the i-th feature column; represents the integrity index of all data points included in the i-th feature column; E i represents the comprehensive importance index of all data points included in the i-th feature column; O ij represents the importance index of the j-th data point included in the i-th feature column; W i represents the comprehensive weight of all data points included in the i-th feature column; N i represents the number of data points in the i-th feature column. α, β, γ, δ respectively represent the weight coefficients of the consistency index, integrity index, comprehensive importance index, and importance index; Step S4: Use the weighted threshold of all data points included in each feature column to determine the importance index of each data point included in each feature column, and obtain a high-quality data set, including: Step S41: Perform normalization processing on the weighted thresholds of all data points included in each feature column and the importance indices of each data point included in each feature column; Step S42: Use the normalized weighted thresholds to determine the importance indices of each data point; Step S43: Determine the importance index of the current data point in the current feature column according to the weighted threshold. If the importance index of the current data point in the current feature column is greater than or equal to the weighted threshold, retain the current data point; if the importance index of the current data point in the current feature column is less than the weighted threshold, eliminate the current data point; until the importance indices of each data point included in each feature column are all determined by the weighted threshold; Step S44: Define the current feature column after determination as a new feature column, and integrate all the obtained new feature columns to form a high-quality data set; Step S5: Use the high-quality data set to train the deep learning model, and predict the digital archive data to be processed according to the trained deep learning model to obtain a prediction result, and make a business decision according to the prediction result; The prediction process is prior art and will not be elaborated here one by one; the main point is to illustrate how to make a business decision according to the prediction result. Specifically, first, the prediction result should be explained to ensure its interpretability. For example, for an anomaly detection task, explain which features lead to the anomaly judgment, and it is necessary to deeply understand the meaning of the prediction result and link it to the business goal. Then, make a business decision based on the explained prediction result, including: customer segmentation: segment customers based on the prediction result, identify high-value customers, potential churn customers, etc.; personalized service: provide personalized services and marketing strategies for different customer groups; for example, provide exclusive offers and services for high-value customers; customer retention: take targeted retention measures for customers predicted to churn, such as providing discounts, improving service quality, etc.; even an operation optimization strategy can be made based on the prediction result, which will not be elaborated here one by one.
[0013] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A method for intelligent information mining of digital archives, characterized in that, Including the following steps: Step S1: Obtain the digital archive dataset to be processed, and preprocess the digital archive dataset to generate a preprocessed digital archive dataset; Step S2: Extract all feature columns from the preprocessed digital archive dataset, respectively obtain the consistency index, integrity index of all data points included in each feature column, and the importance index of each data point, and obtain the comprehensive importance index of all data points based on the importance index of each data point, and obtain the comprehensive weight of all data points included in each feature column according to the consistency index, integrity index and comprehensive importance index; Step S3: Obtain the weighted threshold of all data points included in each feature column according to the consistency index, integrity index, importance index, comprehensive weight and comprehensive importance index; Step S4: Use the weighted threshold of all data points included in each feature column to judge the importance index of each data point included in each feature column to obtain a high-quality dataset; Step S5: Use the high-quality dataset to train the deep learning model, predict the digital archive data to be processed according to the trained deep learning model to obtain a prediction result, and make a business decision according to the prediction result.
2. The digital file intelligent information mining method according to claim 1, wherein The obtaining of the consistency index of all data points included in each feature column includes: Step S2a1: Calculate the mean value of all data points included in each feature column, and calculate the squared deviation of each data point from the mean value based on the mean value; Step S2a2: Based on the squared deviation of each data point from the mean value, obtain the mean value of the squared deviation; Step S2a3: Perform a square root operation on the mean value of the squared deviation to obtain the standard deviation of all data points included in each feature column; Step S2a4: Convert the standard deviation of all data points included in each feature column into a consistency score, and perform a standardization process on the consistency score to obtain the consistency index of all data points included in each feature column.
3. A digital archive intelligent information mining method according to claim 1, characterized in that, The obtaining of the integrity index of all data points included in each feature column includes: Step S2b1: Count the number of all data points included in each feature column; Step S2b2: Detect the missing values of each data point included in each feature column, and count the number of data points with missing values in each feature column; Step S2b3: Obtain the missing value ratio of all data points included in each feature column according to the number of all data points included in each feature column and the number of data points with missing values in each feature column; Step S2b4: Convert the missing value ratio of all data points included in each feature column into an integrity score, and perform a standardization process on the integrity score to obtain the integrity index of all data points included in each feature column.
4. A digital archive intelligent information mining method according to claim 1, characterized in that, The using of the weighted threshold of all data points included in each feature column for the importance index of each data point included in each feature column includes: Step S2c1: Analyze the importance degree of each data point included in each feature column, and assign an importance score to each data point according to the analysis result; Step S2c2: Perform normalization processing on the importance scores of all data points included in each feature column to obtain the importance index of each data point included in each feature column; Step S2c3: Integrate the importance indices of each data point included in each feature column and perform averaging processing to obtain the comprehensive importance index of all data points included in each feature column.
5. A method for intelligent information mining of digital archives according to claim 1, characterized in that, The obtaining of the comprehensive weight of all data points included in each feature column according to the consistency index, integrity index and comprehensive importance index includes: Step S2d1: Assign initial weights to all data points included in each feature column according to the comprehensive importance index; Step S2d2: Obtain the secondary weights of all data points included in each feature column according to the consistency index of all data points included in each feature column and the integrity index of all data points included in each feature column; Step S2d3: Adjust the initial weights according to the secondary weights to obtain the comprehensive weights of all data points included in each feature column.
6. The digital file intelligent information mining method according to claim 1, characterized in that The calculation formula of the weighted threshold is: ; Among them, T i represents the weighted threshold of all data points included in the i-th feature column; C i represents the consistency index of all data points included in the i-th feature column; represents the integrity index of all data points included in the i-th feature column; E i represents the comprehensive importance index of all data points included in the i-th feature column; O ij represents the importance index of the j-th data point included in the i-th feature column; W i represents the comprehensive weight of all data points included in the i-th feature column; N i represents the number of data points in the i-th feature column; α, β, γ, δ represent the weight coefficients of the consistency index, the weight coefficient of the integrity index, the weight coefficient of the comprehensive importance index, and the weight coefficient of the importance index, respectively.
7. A digital archive intelligent information mining method according to claim 1, characterized in that, Determine the importance index of each data point included in each feature column by using the weighted threshold of all data points included in each feature column to obtain a high-quality data set, including: Step S41: Perform normalization processing on the weighted threshold of all data points included in each feature column and the importance index of each data point included in each feature column; Step S42: Determine the importance index of each data point by using the normalized weighted threshold; Step S43: Determine the importance index of the current data point in the current feature column according to the weighted threshold. If the importance index of the current data point in the current feature column is greater than or equal to the weighted threshold, retain the current data point; if the importance index of the current data point in the current feature column is less than the weighted threshold, eliminate the current data point; until the importance index of each data point included in each feature column has been determined; Step S44: Define the current feature column after determination as a new feature column, and integrate all the obtained new feature columns to form a high-quality data set.
Citation Information
Patent Citations
Remote monitoring and fault diagnosis method and system for environment simulation system
CN119728452A
Cement and coking enterprise atmospheric pollutant emission monitoring method and device based on electric power data
CN119761713A
Other Solution Automation & Interface Analysis Implementations
US20230044564A1