Archive information management method and system based on artificial intelligence
By comprehensively utilizing multi-dimensional feature vectors and adaptively adjusting the neighborhood radius of the DBSCAN algorithm, the fragmentation and merging problems caused by single features in automatic archive classification are solved, achieving high accuracy and efficient management of archive classification.
Patent Information
- Application Number
- CN202511059956.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-14
AI Technical Summary
Existing automatic archival classification technologies rely on a single feature to construct a feature space, which cannot adapt to differences in the density of multiple features. This results in high-density archives being fragmented into small clusters, low-density archives being forcibly merged, and the differences in location and time of the documents being ignored.
By acquiring multi-dimensional features such as word vectors, sentence vectors, time vectors, and location vectors, the accuracy and information entropy of the feature vectors are calculated. The initial neighborhood radius of the DBSCAN algorithm is adaptively adjusted to optimize the clustering process and achieve accurate classification of archives.
It significantly improves the accuracy of archival classification and clustering quality, enhances the efficiency of archival management and retrieval, strengthens the system's intelligence level, and adapts to archival classification tasks in various complex scenarios.
Smart Images

Figure CN120950465A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing, and more particularly to an artificial intelligence-based method and system for managing archival information. Background Technology
[0002] Archival classification and coding are fundamental tasks in archival management, crucial for the orderly management and efficient retrieval of archives. A scientific and rational archival classification system not only enhances the standardization and normalization of archival management but also promotes the integration, sharing, and utilization of archival information.
[0003] In the context of information-based archival management, traditional methods relying on manual archival classification have gradually revealed numerous problems, failing to meet the demands of rapidly developing information technology. To address this challenge, existing technologies are beginning to employ artificial intelligence algorithms to achieve information-based archival classification by acquiring word and phrase vectors. These automatic classification technologies based on clustering algorithms play a crucial role in improving the efficiency of archival management.
[0004] However, existing automatic classification technologies also have some limitations. For example, various high-frequency words are commonly found in different archives, which makes the traditional DBSCAN algorithm prone to erroneously merging policy documents and implementation documents when relying on a single feature (such as text keywords) to construct the feature space. This reliance on a single feature not only ignores the differences in location and time among different documents, but its fixed parameter settings also cannot adapt to the density differences of multiple features. This results in archives being incorrectly fragmented into small clusters in high-density feature regions, while being forcibly merged in low-density feature regions. Summary of the Invention
[0005] To address the limitations of existing classification technologies that rely on a single feature to construct the feature space, which cannot adapt to differences in the density of multiple features, resulting in high-density area files being fragmented into small clusters and low-density area files being forcibly merged, while ignoring the differences in location and time among different files, this invention provides solutions in the following aspects.
[0006] In the first aspect, the artificial intelligence-based archival information management method includes: acquiring information data of classified historical archives and archives to be classified in the information management system, and performing word and sentence segmentation processing to obtain features and feature vectors of the information data of historical archives and archives to be classified, wherein the feature vectors include: word vectors, sentence vectors, time vectors, and location vectors; clustering the feature vectors corresponding to each feature of the classified historical archives, calculating the difference between the classification result of the historical archives and the clustering result to obtain the number of misclassified historical archives, and using the ratio of 1 minus the number of misclassified historical archives to the total number of historical archives as the accuracy of the classification of the corresponding feature vectors, and obtaining the attention weight of the corresponding feature based on the information entropy and accuracy of each feature vector; using the cosine similarity of the feature vectors between the features of the archives to be classified to calculate the density of each feature, fusing different feature vectors according to the normalized attention weights to obtain the target vector and the comprehensive density, and using the feature density and the comprehensive density to adaptively adjust the initial neighborhood radius during the clustering process to obtain the adjustment radius; using DBSCAN to cluster the archives to be classified based on the adjustment radius and Euclidean distance to achieve classification, thus completing the classification and archiving of the archives to be classified.
[0007] The benefits are as follows: By comprehensively utilizing multi-dimensional features such as word vectors, sentence vectors, time vectors, and location vectors, the accuracy of archival classification and the quality of clustering are significantly improved. Attention weights are determined by calculating the classification accuracy and information entropy of the feature vectors, and then the initial neighborhood radius of the clustering algorithm is adaptively adjusted based on these weights and feature density, optimizing the DBSCAN clustering process and effectively avoiding over-segmentation or under-segmentation. Furthermore, this method improves the efficiency of archival management and retrieval, enhances the system's intelligence level, and thus provides users with a more accurate and efficient archival information management solution.
[0008] Preferably, the method for calculating the attention weights includes:
[0009] Using any feature in the classified historical archives as a marker feature, calculate the information entropy of the marker feature, and calculate the ratio between the information entropy of the marker feature and the largest information entropy value among the features to obtain the normalized information entropy value;
[0010] The normalized accuracy is obtained by calculating the ratio between the clustering accuracy of the feature vectors of the labeled features and the sum of the clustering accuracy of the feature vectors of all features.
[0011] The product of the normalized information entropy value and the normalized accuracy is used as the attention weight for the labeled features.
[0012] Its effectiveness lies in combining accuracy and information entropy values to calculate attention weights, effectively avoiding over-reliance on single features, fully utilizing the information richness of each feature, and reducing the interference of unimportant features on classification results. This not only improves classification accuracy but also enhances the model's robustness and adaptability, making it suitable for archival classification tasks in various complex scenarios.
[0013] Preferably, the method for calculating the density of the feature includes:
[0014] Using any feature in the files to be classified as the target feature, sum the cosine similarity of the target features of all files to be classified to obtain the density of the feature vector of the target feature for all files to be classified.
[0015] Preferably, the target vector includes:
[0016] The target vector is obtained by weighting and concatenating the attention weights of all files to be classified with the feature vectors of each feature. One target vector corresponds to one file to be classified.
[0017] Preferably, the overall density includes the following steps:
[0018] Using any feature in the files to be classified as the target feature, calculate the cosine similarity between the target vectors of any two files to be classified to measure their similarity in the feature space. Sum the cosine similarities of all files to be classified and divide by the total number of files to be classified to obtain the comprehensive density.
[0019] Preferably, the adjustment radius includes the following steps:
[0020] Using any feature in the archives to be classified as the target feature, calculate the difference between the density of the target feature and the overall density, and sum the weighted summation of the product of the target feature and the attention weight of the target feature. Use an exponential function to map the weighted summation result, and use the mapping result as the adjustment weight of the initial neighborhood radius to obtain the adjustment radius.
[0021] Preferably, the classification and archiving of the files to be classified includes:
[0022] If the ratio of the number of files in each cluster to the total number of files is less than a preset threshold, the cluster is considered an abnormal cluster; conversely, if the ratio is greater than the preset threshold, the cluster is considered a normal cluster. Abnormal clusters are manually labeled and classified a second time to facilitate management and retrieval.
[0023] Secondly, an artificial intelligence-based archival information management system includes: a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the aforementioned artificial intelligence-based archival information management method is implemented.
[0024] The present invention has the following effects:
[0025] 1. This invention overcomes the limitations of traditional single-feature-dependent methods by comprehensively utilizing multiple features such as word vectors, sentence vectors, time vectors, and location vectors. It can adapt to differences in feature density, preventing high-density archives from being incorrectly fragmented into small clusters and low-density archives from being improperly merged. Furthermore, this method specifically considers differences in document location and time, thereby improving the accuracy of archive classification and its adaptability to diverse data.
[0026] 2. This invention achieves adaptive adjustment of the initial neighborhood radius during clustering by calculating the attention weight and comprehensive density of features. This radius adjustment method can more accurately reflect the local structure of the data, thereby optimizing the performance of clustering algorithms such as DBSCAN, reducing over-segmentation or under-segmentation problems, and making the clustering results more consistent with the actual data distribution. Attached Figure Description
[0027] Figure 1 This is a flowchart of steps S1-S4 in the artificial intelligence-based archival information management method of this invention.
[0028] Figure 2 This is a structural block diagram of an artificial intelligence-based archival information management system according to an embodiment of the present invention. Detailed Implementation
[0029] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0030] Reference Figure 1 The artificial intelligence-based archival information management method includes steps S1-S4, as detailed below:
[0031] S1: Obtain information data of classified historical archives and unclassified archives in the information management system, and perform word segmentation and sentence segmentation processing to obtain the features and feature vectors of the information data of historical archives and unclassified archives. The feature vectors include: word vectors, sentence vectors, time vectors and location vectors.
[0032] It should be noted that in the information management system, to extract useful information from classified historical archives and unclassified archives, the text data is first processed through word segmentation and sentence segmentation. Then, the TF-IDF algorithm is used to extract word vectors, time vectors, and location vectors. These vectors are mainly obtained from the archive covers, but may also be extracted from the main text as needed. Simultaneously, the BERT model is used to obtain sentence vectors from the archive covers and the main text to capture deeper semantic information. To eliminate the influence of different feature vectors on their dimensions, these vectors are standardized, and may be selected and fused based on feature importance to form a comprehensive feature vector. Furthermore, other pre-trained word embedding models such as Word2Vec and GloVe, as well as rule-based methods or specific NLP tools, can be considered to further optimize the feature extraction process.
[0033] S2: Cluster the feature vectors corresponding to each feature of the classified historical archives, calculate the difference between the classification result and the clustering result of the historical archives, and obtain the number of misclassified historical archives. Use the ratio of 1 minus the number of misclassified historical archives to the total number of historical archives as the accuracy of the classification of the corresponding feature vector. Obtain the attention weight of the corresponding feature based on the information entropy and accuracy of each feature vector.
[0034] Specifically, the accuracy of clustering classification for each feature satisfies the following relationship:
[0035]
[0036] Among them, A k C' represents the accuracy of clustering the feature vector of the k-th feature. k Let C represent the number of historical archives that were misclassified based solely on the feature vector of feature k, and let C represent the total number of historical archives.
[0037] In other words, higher accuracy indicates that the features perform better in the classification task, contribute more to the classification results, help identify key features, optimize the feature selection process, and improve the overall accuracy and efficiency of classification.
[0038] Methods for calculating attention weights include:
[0039] Using any feature in the classified historical archives as a marker feature, calculate the information entropy of the marker feature, and calculate the ratio between the information entropy of the marker feature and the largest information entropy value among the features to obtain the normalized information entropy value;
[0040] The normalized accuracy is obtained by calculating the ratio between the clustering accuracy of the feature vectors of the labeled features and the sum of the clustering accuracy of the feature vectors of all features.
[0041] The product of the normalized information entropy value and the normalized accuracy is used as the attention weight for the labeled features.
[0042] Specifically, the attention weights satisfy the following relationship:
[0043]
[0044] In the formula, ω k A represents the attention weight of the k-th feature, N represents the number of features, and A k H represents the accuracy of clustering the feature vector of the k-th feature. k H represents the information entropy value of the k-th feature, and max(H) represents the maximum information entropy value.
[0045] It should be noted that this embodiment uses four features: words, sentences, time, and location, but it is not limited to these four. In practical applications, the types of features can be added or reduced according to specific needs. For example, author features, document type features, etc., can be added to further enrich the feature system.
[0046] If a feature has a high accuracy rate, it means that the feature performs well in the classification task and can effectively distinguish between different categories of documents. For example, if the accuracy rate of word vectors is high, it means that the semantic information of the words is very helpful for document classification. The higher the accuracy rate, the better the feature distinguishes documents.
[0047] By calculating the information entropy of each feature, we can determine whether each feature contains a wealth of information, and thus further determine whether it can provide more classification basis. The larger the information entropy value, the richer the information contained in the feature. For example, the information entropy of the time vector is very high, indicating that time information is very important in the classification of archives, and archives from different times have obvious differences.
[0048] When calculating the attention weights for each feature, the accuracy and information entropy values of features can vary significantly. Directly using these weights may result in some features having excessively high or low weights, thus affecting the balance of the classification results. Normalization adjusts the attention weights of all features to the same scale, ensuring that the weight of each feature is within the range [0,1] and that the sum of all feature weights is 1. This avoids classification bias caused by excessive weight differences and avoids the classification blind spots of a single feature, enabling cross-validation of spatiotemporal semantics and thus improving the accuracy and robustness of classification.
[0049] In practical applications, the importance of different features may vary depending on the dataset and task. Without normalization, the model may over-rely on some features while ignoring others.
[0050] S3: Based on the cosine similarity of feature vectors between features of the archives to be classified, the density of each feature is calculated. The different feature vectors are fused according to the normalized attention weights to obtain the target vector and the comprehensive density. The feature density and the comprehensive density are used to adaptively adjust the initial neighborhood radius during the clustering process to obtain the adjustment radius.
[0051] Specifically, cosine similarity satisfies the following relationship:
[0052]
[0053] In the formula, S i,j V represents the cosine similarity between the feature vectors corresponding to the features of the i-th and j-th files to be classified. i V j Let |V| represent the feature vectors of the i-th and j-th files to be classified on the k-th feature, respectively. i ||、||V j || represent the norms of the eigenvectors corresponding to the k-th feature of the i-th and j-th files to be classified, respectively.
[0054] Methods for calculating feature density include:
[0055] Using any feature in the files to be classified as the target feature, sum the cosine similarity of the target features of all files to be classified to obtain the density of the feature vector of the target feature for all files to be classified.
[0056] Specifically, the density satisfies the following relationship:
[0057]
[0058] In the formula, ρ k S represents the density of all files to be classified on the eigenvector of the k-th feature, where n represents the number of files to be classified. i,j It represents the cosine similarity between the feature vectors corresponding to the features of the i-th and j-th files to be classified.
[0059] To further explain, by calculating the cosine similarity between each pair of archives and summing all the cosine similarities, the density of all archives to be classified on the feature vector of the k-th feature can be obtained. This not only reflects the density of the archives' distribution on that feature but also provides an important basis for subsequent cluster analysis, helping to more accurately adjust clustering parameters and optimize classification results.
[0060] The target vector is obtained by weighting and concatenating the attention weights of all files to be classified with the feature vectors of each feature. One target vector corresponds to one file to be classified.
[0061] It's important to note that, to classify archives more accurately, we first analyze and determine the attention weights for each feature from already classified historical archives. These attention weights reflect the importance of different features in the classification process. These weights are applied to the feature vectors of the archives to be classified, using a weighting operation to emphasize the contribution of important features and suppress the influence of less critical features. Then, we concatenate the weighted feature vectors into a comprehensive target vector. This vector integrates information from all features, and the weight adjustments reflect the importance of each feature. Ultimately, each archive to be classified corresponds to such a target vector, which serves as input to the classification model, helping it make more effective classification decisions. This process ensures that the classification model can comprehensively consider multiple features, thereby improving the accuracy and reliability of the classification.
[0062] For example, the attention weight for word features is 0.4, and the standardized vector values are: [0.12, 0.85, 0.30, 0.42, 0.08, 0.91, 0.23, 0.56, 0.67, 0.15]; the attention weight for sentence features is 0.3, and the standardized vector values are: [0.75, 0.32, 0.61, 0.44, 0.19, 0.83, 0.27, 0.50]; the attention weight for time features is 0.2, and the standardized vector values are: [0.95, 0.10, 0.45]; the attention weight for location features is 0.1, and the standardized vector values are: [0.33, 0.68, 0.22, 0.57, 0.89]. The weighted feature vectors are as follows: Weighted word vector: [0.048, 0.340, 0.120, 0.168, 0.032, 0.364, 0.092, 0.224, 0.268, 0.060]; Weighted sentence vector: [0.225, 0.096, 0.183, 0.132, 0.057, 0.249, 0.081, 0.150]; Weighted time vector: [0.190, 0.020, 0.090]; Weighted location vector: [0.033, 0.068, 0.022] [0.057,0.089]; superimpose all weighted vectors: [0.048,0.340,0.120,0.168,0.032,0.364,0.092,0.224,0.268,0.060,0.225,0.096,0.183,0.132,0.057,0.249,0.081,0.150,0.190,0.020,0.090,0.033,0.068,0.022,0.057,0.089].
[0063] The overall density includes the following steps:
[0064] Using any feature in the files to be classified as the target feature, calculate the cosine similarity between the target vectors of any two files to be classified to measure their similarity in the feature space. Sum the cosine similarities of all files to be classified and divide by the total number of files to be classified to obtain the comprehensive density.
[0065] It should be noted that high overall density indicates that archives are close to each other in the feature space and may belong to similar categories, which is beneficial to improving the accuracy of classification and clustering; while low overall density indicates that archives are scattered and may belong to different categories or contain outliers. This indicator not only helps to identify important features and guide data preprocessing, but also allows for the selection of appropriate machine learning models based on the distribution characteristics of the archives, thereby optimizing the performance of tasks such as classification, clustering, and anomaly detection.
[0066] Using any feature in the archives to be classified as the target feature, calculate the difference between the density of the target feature and the overall density, and sum the weighted summation of the product of the target feature and the attention weight of the target feature. Use an exponential function to map the weighted summation result, and use the mapping result as the adjustment weight of the initial neighborhood radius to obtain the adjustment radius.
[0067] Specifically, the adjustment radius satisfies the following relationship:
[0068]
[0069] In the formula, R' represents the adjustment radius, R0 represents the initial neighborhood radius, N represents the number of features, and ρ k ρ' represents the density of all files to be classified on the eigenvector of the k-th feature, ρ0 represents the overall density of the files to be classified, and ω' represents the density of all files to be classified on the eigenvector of the k-th feature. k Let represent the attention weight after normalization of the k-th feature, and exp() represent the exponential function with the natural number e as the base.
[0070] By considering the density of each feature and the attention weight, the initial neighborhood radius can be adaptively adjusted, which can better adapt to the distribution characteristics of the data and improve the accuracy of classification. By calculating the difference between the density of each feature and the overall density, the differences in data distribution on features of different people can be obtained, thereby adjusting the neighborhood radius more accurately.
[0071] To further explain, in response to ρ k When the difference ρ0 is less than 0, the feature density is less than the overall density, and the radius should be reduced to avoid excessive segmentation due to an overly large radius, in response to ρ. k When the difference between -ρ0 and the overall density is greater than 0, the feature density is greater than the overall density, and the radius should be increased to avoid undersegmentation due to an insufficient radius. The attention weight represents the degree to which the difference between the feature density and the overall density affects the adjustment of the neighborhood radius. The larger the weight, the more attention is paid to the feature density and the more it dominates the adjustment direction.
[0072] S4: Based on the adjustment radius and Euclidean distance, DBSCAN is used to cluster the files to be classified to achieve classification, thus completing the classification and archiving of the files to be classified.
[0073] If the ratio of the number of files in each cluster to the total number of files is less than a preset threshold, the cluster is considered an abnormal cluster; conversely, if the ratio is greater than the preset threshold, the cluster is considered a normal cluster. Abnormal clusters are manually labeled and classified a second time to facilitate management and retrieval.
[0074] For example, the preset threshold is 30%, which can be adjusted according to specific circumstances. For instance, professional files are a special case and must be categorized separately. Since they are core project files and there are few of them, they also need to be categorized separately. This is not an abnormal situation. This step depends on the situation, and the threshold can be adjusted according to the specific scenario.
[0075] This invention also provides an artificial intelligence-based archival information management system. For example... Figure 2 As shown, the system includes a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement the artificial intelligence-based archival information management method according to the first aspect of the present invention. The system also includes other components well known to those skilled in the art, such as a communication bus and a communication interface, the setup and functions of which are known in the art and will not be described further here.
[0076] It should be noted that those skilled in the art can make various modifications and improvements without departing from the inventive concept, and these all fall within the scope of protection of this invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. An artificial intelligence-based method for managing archival information, characterized in that: include: The system acquires information data from classified historical archives and unclassified archives in the information management system, performs word and sentence segmentation processing, and obtains the features and feature vectors of the information data from the historical archives and unclassified archives. The feature vectors include word vectors, sentence vectors, time vectors, and location vectors. Cluster the feature vectors corresponding to each feature of the classified historical archives, calculate the difference between the classification result and the clustering result of the historical archives to obtain the number of misclassified historical archives, and use the ratio of 1 minus the number of misclassified historical archives to the total number of historical archives as the classification accuracy of the corresponding feature vector. Obtain the attention weight of the corresponding feature based on the information entropy and accuracy of each feature vector. Based on the cosine similarity of feature vectors among the features of the archives to be classified, the density of each feature is calculated. The different feature vectors are fused according to the normalized attention weights to obtain the target vector and the comprehensive density. The feature density and the comprehensive density are used to adaptively adjust the initial neighborhood radius during the clustering process to obtain the adjustment radius. DBSCAN is used to cluster the files to be classified based on the adjustment radius and Euclidean distance, thus completing the classification and archiving of the files to be classified.
2. The artificial intelligence-based archival information management method according to claim 1, characterized in that, The method for calculating the attention weights includes: Using any feature in the classified historical archives as a marker feature, calculate the information entropy of the marker feature, and calculate the ratio between the information entropy of the marker feature and the largest information entropy value among the features to obtain the normalized information entropy value; The normalized accuracy is obtained by calculating the ratio between the clustering accuracy of the feature vectors of the labeled features and the sum of the clustering accuracy of the feature vectors of all features. The product of the normalized information entropy value and the normalized accuracy is used as the attention weight for the labeled features.
3. The artificial intelligence-based archival information management method according to claim 1, characterized in that, The method for calculating the density of the feature includes: Using any feature in the files to be classified as the target feature, sum the cosine similarity of the target features of all files to be classified to obtain the density of the feature vector of the target feature for all files to be classified.
4. The artificial intelligence-based archival information management method according to claim 1, characterized in that, The target vector includes: The target vector is obtained by weighting and concatenating the attention weights of all files to be classified with the feature vectors of each feature. One target vector corresponds to one file to be classified.
5. The artificial intelligence-based archival information management method according to claim 1, characterized in that, The overall density includes the following steps: Using any feature in the files to be classified as the target feature, calculate the cosine similarity between the target vectors of any two files to be classified to measure their similarity in the feature space. Sum the cosine similarities of all files to be classified and divide by the total number of files to be classified to obtain the comprehensive density.
6. The artificial intelligence-based archival information management method according to claim 1, characterized in that, The adjustment radius includes the following steps: Using any feature in the archives to be classified as the target feature, calculate the difference between the density of the target feature and the overall density, and sum the weighted summation of the product of the target feature and the attention weight of the target feature. Use an exponential function to map the weighted summation result, and use the mapping result as the adjustment weight of the initial neighborhood radius to obtain the adjustment radius.
7. The artificial intelligence-based archival information management method according to claim 1, characterized in that, The classification and archiving of the files to be classified includes: If the ratio of the number of files in each cluster to the total number of files is less than a preset threshold, the cluster is considered an abnormal cluster; conversely, if the ratio is greater than the preset threshold, the cluster is considered a normal cluster. Abnormal clusters are manually labeled and classified a second time to facilitate management and retrieval.
8. An artificial intelligence-based archival information management system, characterized in that: include: A processor and a memory, the memory storing computer program instructions that, when executed by the processor, implement the artificial intelligence-based archival information management method according to any one of claims 1-7.
Citation Information
Cited By
Construction engineering project file information management and storage method and system
CN121785995A