Intelligent archive content evaluation system based on deep learning
By building a deep learning evaluation system with multi-model fusion, the problem of single evaluation dimensions and poor dynamic adaptability of the archive management system is solved, and multi-dimensional comprehensive evaluation and intelligent dynamic evaluation are realized, which improves the comprehensiveness and scientificity of the evaluation and enhances the practicality and reliability of the system.
Patent Information
- Application Number
- CN202510575842.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The existing archive management system has a single evaluation dimension, poor dynamic adaptability, difficulty in conducting multi-dimensional comprehensive evaluation, limited intelligence level, low warning accuracy, lack of dynamic confidence analysis, and cannot effectively capture complex semantic and timing behavior patterns.
Build a deep learning evaluation system that integrates multi-models, including archive data collection, content preprocessing, model construction, indicator setting and intelligent evaluation modules. Through semantic understanding, risk assessment and value classification models, combined with evaluation rule decision tree, dynamically optimize split properties, generate evaluation confidence matrix, trigger exception marking and hierarchical early warning.
It realizes a multi-dimensional comprehensive assessment of archive content, improves the comprehensiveness and scientificity of the assessment, can promptly detect abnormalities and provide repair suggestions, and enhances the practicality and reliability of the system.
Smart Images

Figure CN120492405A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to an intelligent archive content evaluation system based on deep learning. Background Art
[0002] With the accelerated advancement of informatization, the archival data generated by various institutions has experienced explosive growth, and its content has become increasingly complex, encompassing a variety of formats including text, tables, images, and multimedia. At the same time, the demands of archival management have gradually shifted from traditional storage and retrieval to the dynamic assessment of content quality, security, and utilization value.
[0003] Currently, natural language processing and machine learning technologies have been gradually introduced into the field of archival management, such as classification systems based on keyword matching and risk warning tools based on statistical models. Some studies have attempted to use deep learning models, such as convolutional neural networks and recurrent neural networks, to perform semantic analysis of archival texts, or to construct thematic datasets through clustering algorithms. However, existing technologies are mostly limited to single tasks, lack multi-dimensional evaluation capabilities, and model training relies on static rules, making it difficult to dynamically optimize evaluation indicators. In addition, existing systems still have significant shortcomings in complex semantic understanding, multi-source feature fusion, and cross-domain indicator supplementation, resulting in insufficient comprehensiveness and credibility of evaluation results.
[0004] The main problems of existing technologies are reflected in the following aspects: the evaluation dimension is single and cannot comprehensively measure the integrity, security and utilization value of archives; the dynamic adaptability is poor, and traditional rule engines find it difficult to dynamically adjust the evaluation rules according to the archive type and confidentiality level, and there is a lack of effective supplementary mechanism when the indicator coverage is insufficient; the level of intelligence is limited, the semantic understanding of unstructured text is not deep enough, and risk assessment relies on manually preset sensitive word libraries, which makes it difficult to capture the potential correlation between temporal behavior patterns and permission violations; the early warning accuracy is low, and anomaly detection is mostly based on static thresholds, lacking dynamic confidence analysis that combines semantic features with temporal features.
[0005] Therefore, it is necessary to invent an intelligent archive content evaluation system based on deep learning to solve the above problems. Summary of the Invention
[0006] The purpose of the present invention is to provide an intelligent evaluation system for archive content based on deep learning to solve the problems raised in the above background technology.
[0007] To achieve the above objectives, the present invention provides the following technical solutions: a deep learning-based intelligent evaluation system for archive content, comprising the following modules:
[0008] Archive data collection module: obtains archive basic information and content data through the API interface, identifies the archive type and storage format, and constructs the archive content original data set including original content, type label, and storage format;
[0009] Archive content preprocessing module: This module performs text standardization and density clustering-based data cleaning on the original archive content dataset, extracts vocabulary vectors, time series features, and semantic distribution features, generates standardized archive content feature vectors, and outputs thematic dataset labels.
[0010] Model building module: This module is used to construct a multi-model fusion deep learning assessment model. It trains a semantic understanding model, a risk assessment model, and a value classification model based on standardized archival content feature vectors. The outputs of these models are weightedly fused to generate a comprehensive assessment result. An assessment rule decision tree with archival type and confidentiality level as root nodes dynamically optimizes split attributes using the C4.5 algorithm to generate a rule matching benchmark.
[0011] Archives indicator setting module: Based on the evaluation rule decision tree, an evaluation indicator system including integrity, security, and utilization value is established. When the indicator coverage rate is lower than the preset threshold, the adjacent archive type indicators are retrieved through BERT semantic similarity retrieval for weighted supplementation;
[0012] Intelligent Archive Content Assessment Module: This module inputs standardized feature vectors into a deep learning model, combines them with rule matching benchmarks to calculate the time series waveform factor and semantic kurtosis factor, and generates an assessment confidence matrix, including cosine similarity and Spearman correlation coefficient. The module then compares indicator thresholds to trigger abnormality marking and graded warnings.
[0013] Assessment result output module: Based on the anomaly detection results, a structured assessment report is generated through a predefined report template, marking the risk item priority, value level and repair suggestions.
[0014] Preferably, the archive content preprocessing module includes:
[0015] Text processing unit: Builds an archive text library, uses natural language processing technology to perform word segmentation, part-of-speech tagging, and named entity recognition on archive texts, and generates a sequence of vocabulary vectors with entity labels;
[0016] Data cleaning unit: Filter out duplicate and invalid data through cosine similarity matching of vocabulary vectors, and generate thematic data sets by file type and theme through DBSCAN density clustering algorithm;
[0017] Feature fusion unit: extracts timestamps, confidentiality levels, and version information metadata, generates time series features and semantic distribution features, and maps them with vocabulary vectors to form multi-dimensional feature vectors.
[0018] Preferably, the process of establishing the thematic dataset includes:
[0019] Extract institution names, business terms, and event keywords from archive content, calculate the cosine similarity of vocabulary vectors, and for archival application scenarios, classify them into the same thematic dataset when the similarity is greater than a preset threshold;
[0020] According to the archive formation time, storage period and business category, the thematic dataset is time series labeled and weighted.
[0021] Preferably, in the model building module:
[0022] The semantic understanding model uses the Transformer architecture. After inputting the standardized archive content feature vector, it generates a semantic similarity matrix and topic distribution vector through a multi-head attention mechanism.
[0023] The risk assessment model uses a long short-term memory (LSTM) network combined with a conditional random field (CRF). It inputs sensitive word frequency statistics, access log sequence analysis, and modification trace comparison, and outputs leakage risk probability values, tampering risk probability values, and permission violation risk probability values.
[0024] The value classification model uses a multi-layer perceptron (MLP) to input utilization frequency, historical evaluation results, and industry standards, optimizes the weight matrix through the back-propagation algorithm, and generates value classification labels and scoring rules.
[0025] Preferably, the construction process of the evaluation rule decision tree includes:
[0026] Taking the file type and confidentiality level as the root node, the C4.5 algorithm is used to combine the information gain ratio of the temporal waveform factor and the semantic kurtosis factor to dynamically optimize the splitting attributes;
[0027] Split the intermediate nodes and leaf nodes in descending order of information gain ratio. When the information gain of the leaf node is less than the preset threshold, the split will stop. When the number of samples is insufficient, the split will also stop.
[0028] The temporal waveform factor and semantic kurtosis factor output by the leaf node are standardized, and the mean and standard deviation are calculated as the rule matching benchmark.
[0029] Preferably, the implementation of the file indicator setting module includes:
[0030] Extracting archive content integrity indicators from the output of the semantic understanding model, including field missing rate and content consistency score;
[0031] Extract security indicators from the output of the risk assessment model, including sensitive information leakage risk value and access permission compliance score;
[0032] Extracting utilization value indicators from the output of the value classification model, including academic reference value, business guidance value, and historical voucher value scores;
[0033] When the indicator coverage is lower than the threshold, the semantic similarity of adjacent file types is calculated through the BERT model, and a weighted supplementary indicator is generated according to the similarity weight. The weight distribution rule is: for every 0.1 increase in similarity, the supplementary weight increases by 10%.
[0034] Preferably, the file content intelligent evaluation module includes:
[0035] Decision factor calculation unit: Calculates the time series waveform factor composed of content update frequency, access frequency, and related business frequency based on time series characteristics; calculates the semantic kurtosis factor composed of keyword concentration, semantic sentiment intensity, and topic consistency based on keyword distribution;
[0036] Intelligent reasoning unit: performs rule matching by evaluating the decision tree and generates an evaluation confidence matrix including cosine similarity and Spearman correlation coefficient;
[0037] Anomaly detection unit: compares the indicator threshold and triggers an early warning signal according to the risk level when the difference exceeds the threshold. The risk level includes high, medium and low.
[0038] Preferably, the assessment result output module includes a repair suggestion generation unit, which is used to trigger the following operations according to the risk level:
[0039] Level 1 risk: Generates mandatory encryption storage instructions and permission restriction instructions;
[0040] Second level risk: Generate content verification instructions and log audit instructions;
[0041] Level 3 risk: Generate periodic review marking instructions.
[0042] The technical effects and advantages of the present invention are as follows:
[0043] 1. The present invention, through the archive data acquisition module and archive content preprocessing module, can perform comprehensive data collection, standardization processing and feature extraction on archive content, generate multi-dimensional feature vectors, and provide a high-quality data foundation for subsequent evaluation model training and intelligent evaluation, effectively improving the efficiency and accuracy of archive content processing;
[0044] 2. The present invention constructs a multi-model fusion deep learning evaluation model, including a semantic understanding model, a risk assessment model, and a value classification model. Combined with the evaluation rule decision tree to dynamically optimize the split attributes, it can comprehensively evaluate the archive content from multiple dimensions such as semantic understanding, risk assessment, and value classification, generate accurate comprehensive evaluation results and rule matching benchmarks, and improve the comprehensiveness and scientificity of archive content evaluation;
[0045] 3. The present invention establishes an evaluation index system including integrity, security, and utilization value through the archive index setting module. When the index coverage is insufficient, it is supplemented by weighted BERT semantic similarity retrieval. The archive content intelligent evaluation module calculates the time series waveform factor and semantic kurtosis factor, generates an evaluation confidence matrix, and triggers abnormal marking and graded warning. It can realize dynamic and intelligent evaluation of archive content, timely detect abnormalities and provide corresponding repair suggestions, thereby enhancing the practicality and reliability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a schematic diagram of the overall structure of the present invention. DETAILED DESCRIPTION
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0048] The present invention provides Figure 1 The deep learning-based intelligent evaluation system for archive content shown in the figure includes the following modules:
[0049] Archive data collection module: obtains archive basic information and content data through the API interface, identifies the archive type and storage format, and constructs the archive content original data set including original content, type label, and storage format;
[0050] It is important to note that the API interface supports access to multiple data sources and is compatible with data format standards of different systems, ensuring the universality and compatibility of data collection; the basic information includes the file number, creation time, confidentiality level, storage period, affiliated organization and responsible person; the content data includes text content, tables, attachments, images and multimedia files; the original content is the actual content data of the file;
[0051] Archive content preprocessing module: This module performs text standardization and density clustering-based data cleaning on the original archive content dataset, extracts vocabulary vectors, time series features, and semantic distribution features, generates standardized archive content feature vectors, and outputs thematic dataset labels.
[0052] Furthermore, in the above technical solution, the archive content preprocessing module includes:
[0053] Text processing unit: Builds an archive text library, uses natural language processing technology to perform word segmentation, part-of-speech tagging, and named entity recognition on archive texts, and generates a sequence of vocabulary vectors with entity labels;
[0054] Data cleaning unit: Filter out duplicate and invalid data through cosine similarity matching of vocabulary vectors, and generate thematic data sets by file type and theme through DBSCAN density clustering algorithm;
[0055] Feature fusion unit: extracts timestamps, confidentiality levels, and version information metadata, generates time series features and semantic distribution features, and maps them with vocabulary vectors to form multi-dimensional feature vectors.
[0056] Furthermore, in the above technical solution, the process of establishing the thematic dataset includes:
[0057] Extract institution names, business terms, and event keywords from archive content, calculate the cosine similarity of vocabulary vectors, and for archival application scenarios, classify them into the same thematic dataset when the similarity is greater than a preset threshold;
[0058] According to the archive formation time, storage period and business category, the thematic dataset is time series labeled and weighted.
[0059] It should be noted that the core task of the text processing unit is to convert unstructured archival text into a machine-understandable semantic vector sequence. The processing steps are to first collect special vocabulary in the archival field, such as "classification", "retention period", "filing unit", etc., to form an archival field text library to ensure the accuracy of word segmentation and semantic analysis. Combined with natural language processing technology, the continuous text is first divided into independent words, such as "XXXX financial statements" is divided into "XXXX year", "finance", and "report", and then each word is marked with a part of speech, such as noun, verb, time word, etc. For example, "archive" is marked as a verb and "XXXX year" is marked as a time word. Then, proper nouns are identified, such as the name of the institution "XX Archives Bureau", the name of the person "AA", the time "XXXX year X month", etc.), and a vocabulary vector sequence with entity labels is generated, such as [("XX Archives Bureau", institution), ("AA", name)], and finally the vocabulary vector sequence with entity labels is output;
[0060] The core task of the data cleaning unit is to filter invalid data, cluster by file type and theme, and form a thematic data set. The processing steps are to determine the degree of duplication of file content by calculating the cosine similarity of word vectors. If the similarity is ≥ a preset threshold of 0.8, which is set according to the actual situation, it is regarded as duplicate data and eliminated. For example, only one of two meeting minutes with highly similar content is retained. Based on the word vector sequence, combined with file types such as "personnel files" and "financial files" and theme keywords such as "onboarding", "reimbursement", and "audit" as clustering features, similar file contents are clustered into the same theme according to the DBSCAN density clustering algorithm. For example, all financial files containing the keywords "reimbursement process" and "travel expenses" are classified as "travel reimbursement theme". Finally, the thematic data sets are divided by type and theme, and each data set is labeled, such as "personnel-onboarding files" and "finance-reimbursement vouchers".
[0061] The core task of the feature fusion unit is to integrate text semantic features, time attributes and metadata to generate a standardized feature vector. The processing steps are to extract timestamps from the archive, such as the formation time "XXXX-XX-XX", confidentiality level identifiers, such as "secret", "confidential", "top secret", version information, such as "V1.0", "revised version" and other metadata, and calculate the archive formation time and update frequency based on the timestamp, such as "retention period of 10 years", "updated 3 times in the past year", and at the same time analyze the distribution density of keywords in the text, such as "sensitive word frequency" and "theme word concentration". Then, the vocabulary vector sequence is combined with the time series feature and the semantic distribution feature mapping to form a multi-dimensional feature vector containing text semantics, time attributes, confidentiality level and other information, such as a 100-dimensional vector, where the first 50 dimensions are lexical semantic features, the middle 30 dimensions are time and confidentiality level features, and the last 20 dimensions are semantic distribution features. Finally, a standardized archive content feature vector, such as [0.8, 0.3, 0.9, ...] and a thematic data set are output;
[0062] The specific construction process of the thematic dataset is as follows: the first step is semantic clustering based on keywords and cosine similarity, first performing keyword extraction, using named entity recognition technology to extract institution names from the archive content, such as "XX City Archives Bureau", "Finance Department", business terms such as "confidentiality appraisal", "filing", "storage period", event keywords such as "audit rectification", "contract renewal", "personnel transfer", and then performing vocabulary vector conversion, inputting keywords into pre-trained word embedding models such as Word2Vec and BERT to generate high-dimensional semantic vectors, such as 100-dimensional vectors, where the vector values reflect the semantic features of the vocabulary, and finally performing cosine similarity calculation and clustering, calculating the cosine similarity of keyword vectors of different archives, with the preset threshold set to 0.8. When the average similarity of the keyword vectors of two archives is ≥0.8, they are classified into the same thematic dataset; the second step is time series labeling and weight assignment. Time series annotation first extracts the time when the archives are generated, and divides the time intervals by year and quarter, such as "XXXXQ1" and "XXXXQ4", where XXXX is the specific year, Q1 is spring, and Q4 is winter. The retention period is marked according to the archive type, such as "permanent", "10 years", and "short-term", and mapped to a numerical value. For example, permanent is mapped to a value of 5, 10 years is mapped to a value of 3, and short-term is mapped to a value of 2. Archives are classified according to their fields, such as "finance", "personnel", "administration", and "infrastructure", and category labels are generated, such as "C01-finance" and "C02-personnel". The weight distribution rule is: recent archives have a higher weight; permanent archives have a higher weight than short-term archives; core business archives have a higher weight, which is set according to the agency's business priority; the weighted sum of the three is calculated, such as time weight × 0.5 + retention period × 0.3 + business category × 0.2, to generate the weight of each archive in the thematic dataset. The final thematic dataset structure is shown in the table:
[0063] File ID Keyword collection Formation time Storage period Business Category Comprehensive weight D001 ["Financial Audit of XXXX Year",...] XXXX-XX-XX 10 years Finance-Audit 0.3
[0064] Model building module: This module is used to construct a multi-model fusion deep learning assessment model. It trains a semantic understanding model, a risk assessment model, and a value classification model based on standardized archival content feature vectors. The outputs of these models are weightedly fused to generate a comprehensive assessment result. An assessment rule decision tree with archival type and confidentiality level as root nodes dynamically optimizes split attributes using the C4.5 algorithm to generate a rule matching benchmark.
[0065] Furthermore, in the above technical solution, in the model building module:
[0066] The semantic understanding model uses the Transformer architecture. After inputting the standardized archive content feature vector, it generates a semantic similarity matrix and topic distribution vector through a multi-head attention mechanism.
[0067] The risk assessment model uses a long short-term memory (LSTM) network combined with a conditional random field (CRF). It inputs sensitive word frequency statistics, access log sequence analysis, and modification trace comparison, and outputs leakage risk probability values, tampering risk probability values, and permission violation risk probability values.
[0068] The value classification model uses a multi-layer perceptron (MLP) to input utilization frequency, historical evaluation results, and industry standards. It optimizes the weight matrix through a back-propagation algorithm to generate value classification labels and scoring rules.
[0069] It should be noted that the input layer of the semantic understanding model converts vocabulary vectors into Token embeddings and generates input tensors in combination with position encoding. For example, "XXXX-XX-XX" is encoded as a position vector to retain temporal information. The multi-head attention layer uses 8 attention heads, and each head independently calculates the Query, Key, and Value matrices. The feedforward neural network layer performs nonlinear transformations on the multi-head attention output and uses the GeLU activation function to enhance the semantic representation capability. The output layer is a semantic similarity matrix and a topic distribution vector. The semantic similarity matrix calculates the semantic cosine similarity between the current archive and the historical archive, and the topic distribution vector generates a topic probability distribution, such as a procurement probability of 0.7, a financial probability of 0.2, and other probabilities of 0.1, which are used for subsequent value classification model input.
[0070] The risk assessment model inputs sensitive word frequency statistics, such as the frequency of occurrence of sensitive words such as "ID number" and "bank account number" in the archive, access log sequences, such as the format of "[timestamp][user ID][operation type, such as read, modify and delete][authority level]" and modification trace comparison, such as recording the text difference fragments before and after the modification, the modification user and time; the LSTM layer converts the input access log into a time series vector, and the timestamp into a time interval feature, such as the number of hours since the last visit; the network configuration is a 2-layer bidirectional LSTM with a hidden layer dimension of 256, which captures the dependency between the previous and next time steps, such as the time series pattern of multiple consecutive abnormal permission accesses; the hidden state vector of each time step is output to represent the risk characteristics of the moment, such as the change amplitude of sensitive words in a certain modification operation; the sequence in the CRF layer Column labeling is based on the hidden state of LSTM output, and the transition probability of adjacent time step operations is calculated. For example, whether the read-to-delete operation complies with the permission rules, the loss function uses CRF log-likelihood loss, combined with preset risk rules, such as confidential files are prohibited from being modified by unauthorized users, to optimize the risk label sequence prediction; the output results of the output layer are leakage risk probability values, tampering risk probability values, and permission violation risk probability values, which are used to trigger graded warnings; the leakage risk probability value is calculated based on the propagation path of sensitive words, such as the probability of being sent through email attachments, and the range is [0, 1]; the tampering risk probability value is based on the semantic consistency of the modification traces, such as the probability of the modified content logically conflicting with the original text; the permission violation risk probability value compares the matching degree of the operation permission and the file classification, such as the probability of ordinary users accessing confidential files;
[0071] The value classification model inputs the utilization frequency, historical evaluation results and industry standards. The utilization frequency is the number of times the file is retrieved and downloaded, and the format is a continuous numerical value. The historical evaluation results are past value classification labels such as high value, medium value, etc. and scores, such as 1-5 points. The input layer is to perform one-hot encoding on discrete features, such as historical evaluation labels, and normalize continuous features, such as utilization frequency, to the interval [0, 1]. The hidden layer is a two-layer fully connected layer, with 128 neurons in the first layer and ReLU activation function; 64 neurons in the second layer and Sigmoid activation function. For example, input "utilization frequency 0.8 + historical score 4 points + industry standard The weight of medium business guidance is 0.4", and the initial score of each value dimension is calculated through the weight matrix; the output layer outputs the value classification label and scoring rule. The value classification label generates multi-classification probability through the Softmax function, such as high value is 0.6, medium value is 0.3, and low value is 0.1. The scoring rule is a linear combination of scores of each dimension, such as academic reference (0.3×4) + business guidance (0.4×5) + historical voucher (0.3×3) = 4.1 points, retaining 1 decimal place, and the final output format is a binary, such as value classification label: high value, comprehensive score: 4.1, which is used for utilization value evaluation in the archive indicator setting module.
[0072] Furthermore, in the above technical solution, the process of constructing the evaluation rule decision tree includes:
[0073] Taking the file type and confidentiality level as the root node, the C4.5 algorithm is used to combine the information gain ratio of the temporal waveform factor and the semantic kurtosis factor to dynamically optimize the splitting attributes;
[0074] Split the intermediate nodes and leaf nodes in descending order of information gain ratio. When the information gain of the leaf node is less than the preset threshold, the split will stop. When the number of samples is insufficient, the split will also stop.
[0075] The temporal waveform factor and semantic kurtosis factor output by the leaf node are standardized, and the mean and standard deviation are calculated as the rule matching benchmark.
[0076] It should be noted that the preset threshold is set based on common industry practices.
[0077] Archives indicator setting module: Based on the evaluation rule decision tree, an evaluation indicator system including integrity, security, and utilization value is established. When the indicator coverage rate is lower than the preset threshold, the adjacent archive type indicators are retrieved through BERT semantic similarity retrieval for weighted supplementation;
[0078] Furthermore, in the above technical solution, the implementation of the file indicator setting module includes:
[0079] Extracting archive content integrity indicators from the output of the semantic understanding model, including field missing rate and content consistency score;
[0080] Extract security indicators from the output of the risk assessment model, including sensitive information leakage risk value and access permission compliance score;
[0081] Extracting utilization value indicators from the output of the value classification model, including academic reference value, business guidance value, and historical voucher value scores;
[0082] When the indicator coverage is lower than the threshold, the semantic similarity of adjacent file types is calculated through the BERT model, and a weighted supplementary indicator is generated according to the similarity weight. The weight distribution rule is: for every 0.1 increase in similarity, the supplementary weight increases by 10%.
[0083] It should be noted that the calculation formula for the field missing rate is field missing rate = number of missing required fields ÷ total number of required fields; the content consistency score is calculated using a multi-head attention mechanism based on the Transformer model, which generates a semantic similarity matrix for the archive text, calculates the mean cosine similarity between paragraphs, and combines the topic distribution vector to check whether the topic consistency is lower than the preset threshold. If it is lower, points are deducted, and the output range is 0-100 points. The higher the score, the stronger the content consistency. The academic reference value score is calculated based on the number of times the archive is cited by the academic retrieval system and the frequency of reference by research projects, with an output range of 0-5 points. The business guidance value score is calculated based on the frequency of access to the archive by the business department, the number of downloads, and the number of triggering of related business processes, with an output range of 0-5 points. The historical voucher value score is calculated based on the accuracy of the archive formation time, that is, the integrity of the timestamp and the non-tamperability of the content modification traces, with an output range of 0-5 points.
[0084] Intelligent Archive Content Assessment Module: This module inputs standardized feature vectors into a deep learning model, combines them with rule matching benchmarks to calculate the time series waveform factor and semantic kurtosis factor, and generates an assessment confidence matrix, including cosine similarity and Spearman correlation coefficient. The module then compares indicator thresholds to trigger abnormality marking and graded warnings.
[0085] Furthermore, in the above technical solution, the file content intelligent evaluation module includes:
[0086] Decision factor calculation unit: Calculates the time series waveform factor composed of content update frequency, access frequency, and related business frequency based on time series characteristics; calculates the semantic kurtosis factor composed of keyword concentration, semantic sentiment intensity, and topic consistency based on keyword distribution;
[0087] Intelligent reasoning unit: performs rule matching by evaluating the decision tree and generates an evaluation confidence matrix including cosine similarity and Spearman correlation coefficient;
[0088] Anomaly detection unit: compares indicator thresholds and triggers warning signals based on risk levels, including high, medium, and low, when the difference exceeds the threshold.
[0089] It should be noted that the content update frequency is the number of times the archive is modified within a set time period, such as monthly or quarterly, and the mean and standard deviation are calculated by sliding the time window; the access frequency is to parse the access log, count the number of visits per unit time, such as daily or weekly, and perform weighted processing in combination with the visitor's authority level; the associated business frequency is the number of associations between the archive content and the business database analyzed by keyword matching, and classified by business category; the keyword concentration is to calculate the keyword weight based on the TF-IDF algorithm, and count the cumulative proportion of the top N high-weight keywords; the semantic sentiment intensity uses a pre-trained sentiment analysis model, such as BERT-Emotion, to perform sentiment scoring on the archive text and output the probability value of positive, negative or neutral sentiment; the topic consistency extracts the topic distribution of the archive content through the LDA topic model, calculates the KL divergence between each topic, and if the divergence is lower than a threshold, such as 0.1, the topic consistency is determined to be high;
[0090] The comparison method of the indicator threshold is to obtain the dynamic thresholds of integrity, security, and utilization value from the archive indicator setting module, such as integrity threshold = 0.9, security threshold = 0.85, and utilization value threshold = 0.8, and then calculate the degree of difference from the threshold based on the comprehensive score in the confidence matrix, and the degree of difference is the standardized Z value;
[0091] Assessment result output module: Based on the anomaly detection results, a structured assessment report is generated using a predefined report template, with risk item priority, value level, and remediation suggestions marked;
[0092] Furthermore, in the above technical solution, the assessment result output module includes a repair suggestion generation unit, which is used to trigger the following operations according to the risk level:
[0093] Level 1 risk: Generates mandatory encryption storage instructions and permission restriction instructions;
[0094] Second level risk: Generate content verification instructions and log audit instructions;
[0095] Level 3 risk: Generate periodic review mark instructions;
[0096] It is important to know that the condition for the first level of risk is that the absolute value of the standardized Z value is greater than 1.5, the condition for the second level of risk is that the absolute value of the standardized Z value is between 0.5 and 1.5, and the condition for the third level of risk is that the absolute value of the standardized Z value is less than 0.5.
[0097] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A deep learning-based intelligent evaluation system for archive content, characterized in that: Includes the following modules: Archive data collection module: obtains archive basic information and content data through the API interface, identifies the archive type and storage format, and constructs the archive content original data set including original content, type label, and storage format; Archive content preprocessing module: This module performs text standardization and density clustering-based data cleaning on the original archive content dataset, extracts vocabulary vectors, time series features, and semantic distribution features, generates standardized archive content feature vectors, and outputs thematic dataset labels. Model building module: This module is used to construct a multi-model fusion deep learning assessment model. It trains a semantic understanding model, a risk assessment model, and a value classification model based on standardized archival content feature vectors. The outputs of these models are weightedly fused to generate a comprehensive assessment result. An assessment rule decision tree with archival type and confidentiality level as root nodes dynamically optimizes split attributes using the C4.5 algorithm to generate a rule matching benchmark. Archives indicator setting module: Based on the evaluation rule decision tree, an evaluation indicator system including integrity, security, and utilization value is established. When the indicator coverage rate is lower than the preset threshold, the adjacent archive type indicators are retrieved through BERT semantic similarity retrieval for weighted supplementation; Intelligent Archive Content Assessment Module: This module inputs standardized feature vectors into a deep learning model, combines them with rule matching benchmarks to calculate the time series waveform factor and semantic kurtosis factor, and generates an assessment confidence matrix, including cosine similarity and Spearman correlation coefficient. The module then compares indicator thresholds to trigger abnormality marking and graded warnings. Assessment result output module: Based on the anomaly detection results, a structured assessment report is generated through a predefined report template, marking the risk item priority, value level and repair suggestions.
2. The deep learning-based intelligent evaluation system for archive content according to claim 1, characterized in that: The archive content preprocessing module includes: Text processing unit: Builds an archive text library, uses natural language processing technology to perform word segmentation, part-of-speech tagging, and named entity recognition on archive texts, and generates a sequence of vocabulary vectors with entity labels; Data cleaning unit: Filter out duplicate and invalid data through cosine similarity matching of vocabulary vectors, and generate thematic data sets by file type and theme through DBSCAN density clustering algorithm; Feature fusion unit: extracts timestamps, confidentiality levels, and version information metadata, generates time series features and semantic distribution features, and maps them with vocabulary vectors to form multi-dimensional feature vectors.
3. The deep learning-based intelligent evaluation system for archive content according to claim 2, characterized in that: The process of establishing the thematic dataset includes: Extract institution names, business terms, and event keywords from archive content, calculate the cosine similarity of vocabulary vectors, and for archival application scenarios, classify them into the same thematic dataset when the similarity is greater than a preset threshold; According to the archive formation time, storage period and business category, the thematic dataset is time series labeled and weighted.
4. The deep learning-based intelligent evaluation system for archive content according to claim 1, characterized in that: In the model building module: The semantic understanding model uses the Transformer architecture. After inputting the standardized archive content feature vector, it generates a semantic similarity matrix and topic distribution vector through a multi-head attention mechanism. The risk assessment model uses a long short-term memory (LSTM) network combined with a conditional random field (CRF). It inputs sensitive word frequency statistics, access log sequence analysis, and modification trace comparison, and outputs leakage risk probability values, tampering risk probability values, and permission violation risk probability values. The value classification model uses a multi-layer perceptron (MLP) to input utilization frequency, historical evaluation results, and industry standards, optimizes the weight matrix through the back-propagation algorithm, and generates value classification labels and scoring rules.
5. The deep learning-based intelligent evaluation system for archive content according to claim 1, characterized in that: The construction process of the evaluation rule decision tree includes: Taking the file type and confidentiality level as the root node, the C4.5 algorithm is used to combine the information gain ratio of the temporal waveform factor and the semantic kurtosis factor to dynamically optimize the splitting attributes; Split the intermediate nodes and leaf nodes in descending order of information gain ratio. When the information gain of the leaf node is less than the preset threshold, the split will stop. When the number of samples is insufficient, the split will also stop. The temporal waveform factor and semantic kurtosis factor output by the leaf node are standardized, and the mean and standard deviation are calculated as the rule matching benchmark.
6. The deep learning-based intelligent evaluation system for archive content according to claim 1, characterized in that: The implementation of the file indicator setting module includes: Extracting archive content integrity indicators from the output of the semantic understanding model, including field missing rate and content consistency score; Extract security indicators from the output of the risk assessment model, including sensitive information leakage risk value and access permission compliance score; Extracting utilization value indicators from the output of the value classification model, including academic reference value, business guidance value, and historical voucher value scores; When the indicator coverage is lower than the threshold, the semantic similarity of adjacent file types is calculated through the BERT model, and a weighted supplementary indicator is generated according to the similarity weight. The weight distribution rule is: for every 0.1 increase in similarity, the supplementary weight increases by 10%.
7. The deep learning-based intelligent evaluation system for archive content according to claim 1, characterized in that: The file content intelligent evaluation module includes: Decision factor calculation unit: Calculates the time series waveform factor composed of content update frequency, access frequency, and related business frequency based on time series characteristics; calculates the semantic kurtosis factor composed of keyword concentration, semantic sentiment intensity, and topic consistency based on keyword distribution; Intelligent reasoning unit: performs rule matching by evaluating the decision tree and generates an evaluation confidence matrix including cosine similarity and Spearman correlation coefficient; Anomaly detection unit: compares the indicator threshold and triggers an early warning signal according to the risk level when the difference exceeds the threshold. The risk level includes high, medium and low.
8. The deep learning-based intelligent evaluation system for archive content according to claim 1, characterized in that: The assessment result output module includes a repair suggestion generation unit, which is used to trigger the following operations based on the risk level: High-priority risks: Generate mandatory encryption storage instructions and permission restriction instructions; Medium-priority risks: Generate content verification instructions and log audit instructions; Low priority risks: Generate periodic review markup instructions.
Citation Information
Patent Citations
An intelligent reasoning method of archival data based on semantic ontology
CN109271484A
Social media event detection method combining deep learning classification and graph clustering
CN117974340A
Personnel file intelligent management method and system
CN118840087A
Moxibustion decision-making method and system based on Bert pre-training and deep reinforcement learning
CN119694493A
Semantic cluster formation in deep learning intelligent assistants
US20210374168A1
Cited By
Multi-modal data driven general report generation method and system based on large model
CN121031543A
Deep learning-based pollution site multi-source heterogeneous data intelligent analysis and standardization system and method, electronic equipment and storage medium
CN121095038A
Intelligent search system based on deep learning
CN121256072A
An intelligent search system based on deep learning
CN121256072B
Scientific and technological achievement intelligent evaluation method and system based on multi-source data fusion
CN121481296A