Intelligent computer file classification management system based on artificial intelligence
By using multimodal feature fusion and adaptive learning of dynamic knowledge graphs, the problems of low efficiency and poor adaptability in traditional methods are solved, achieving efficient and accurate classification of computer files and discovery of unknown categories, thus improving the adaptability and accuracy of the classification model.
Patent Information
- Application Number
- CN202511067048.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional computer file classification methods rely on manual or rule-based methods, which are inefficient and difficult to adapt to massive amounts of files, semantic ambiguity, and new file types. Machine learning-based methods, on the other hand, lack global perception and dynamic learning capabilities, making them difficult to handle multimodal data.
Employing multimodal feature fusion, dynamic knowledge graphs, and adaptive learning methods, features are extracted using BERT, ViT-Base, graph neural networks, and gated attention mechanisms. Combined with knowledge graph construction and user feedback optimization, cross-modal association and discovery of unknown categories are achieved.
It achieves efficient, accurate, and flexible classification of files, supports cross-modal association and discovery of unknown categories, reduces the cost of manual rule maintenance, and improves classification accuracy and adaptability.
Smart Images

Figure CN120929427A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and data management technology, and in particular to an intelligent classification and management system for computer files based on artificial intelligence. Background Technology
[0002] With the acceleration of digital transformation, the types and sizes of computer files are growing exponentially, including multimodal data such as documents (Word / Excel / PDF), images (JPG / PNG), videos (MP4 / AVI), code (Python / Java), and emails (EML).
[0003] Traditional file classification methods mainly rely on two modes: manual classification and rule-based classification. Manual classification is done by users manually labeling or organizing files, which is inefficient and easily affected by subjective factors, making it difficult to handle massive file scenarios. Rule-based classification achieves classification by pre-setting keywords, file extensions, or simple regular expression matching, but it has problems such as high rule maintenance costs, inability to handle semantic ambiguity (files with the same name but different meanings), and difficulty in adapting to new file types (such as emerging formats or cross-modal mixed files).
[0004] In recent years, machine learning-based classification methods have been extensively studied, such as using SVM and random forests for text feature classification. Although these methods have made breakthroughs, they still have shortcomings. They rely only on single-modal features such as text and images, ignoring auxiliary information such as file metadata and contextual relationships. The classification models are trained on fixed training sets and cannot dynamically learn new file types or user-defined classification rules. They also lack a global awareness of the deep semantics and implicit intentions of file content. Therefore, this invention proposes an intelligent computer file classification management system based on artificial intelligence to solve the problems existing in the prior art. Summary of the Invention
[0005] To address the aforementioned problems, the present invention aims to propose an AI-based intelligent classification and management system for computer files. This AI-based intelligent classification and management system completes classification decisions through multimodal feature fusion, semantic association information in dynamic knowledge graphs, and adaptive learning, achieving efficient, accurate, and flexible classification of files, and supporting cross-modal association, discovery of unknown categories, and user intent perception.
[0006] To achieve the objectives of this invention, the invention is implemented through the following technical solution: an intelligent computer file classification management system based on artificial intelligence, comprising a data acquisition and processing module, a multimodal feature extraction module, a knowledge graph construction module, an adaptive classification module, and a dynamic optimization module. The data acquisition and processing module is used to acquire multi-source files and perform preprocessing and metadata extraction. The multimodal feature extraction module extracts multimodal data and performs cross-modal fusion based on the BERT pre-trained model, ViT-Base, graph neural network, and gated attention mechanism. The knowledge graph construction module is used to construct and update the knowledge graph based on file features and external knowledge bases. The adaptive classification module completes classification decisions and identifies unknown categories based on global features and knowledge graph association information. The dynamic optimization module is used to optimize the classification model and knowledge graph through user feedback and incremental learning.
[0007] A further improvement is that the data acquisition and processing module includes an acquisition cache unit, a data processing unit, and an information extraction unit. The acquisition cache unit is used to acquire and cache batch files in real time. The data processing unit is used to perform format standardization and noise reduction on the files. The information extraction unit is used to extract the metadata information of the processed files.
[0008] Further improvements are made in that: the data collection and caching unit collects data in the following ways: local disk scanning and monitoring, cloud storage file download and mail server file information extraction; the metadata information includes file name, creation time, modification history, author, project, and tag information.
[0009] Further improvements are made in that: the multimodal feature extraction module includes a text feature extraction unit, a visual feature extraction unit, a metadata feature extraction unit, and a cross-modal fusion unit. The text feature extraction unit encodes the text content based on the BERT pre-trained model to generate a 768-dimensional semantic vector and extracts it by combining TF-IDF features to capture local keyword information. The visual feature extraction unit uses ViT-Base to extract features from image and video frames and outputs a 768-dimensional visual embedding vector. The metadata feature extraction unit generates a 512-dimensional structured association vector based on graph neural networks to model the relationship between metadata. The cross-modal fusion unit performs weighted fusion of text, visual, and metadata features based on a gating attention mechanism.
[0010] The further improvement lies in: the weighted fusion calculation formula is as follows:
[0011]
[0012] in W represents textual, visual, and metadata feature vectors, respectively. g W v Wm The weight matrix is a learnable matrix, σ is the activation function, and ⊙ represents element-wise multiplication, ultimately generating a 1536-dimensional global feature vector.
[0013] Further improvements are made in that: the knowledge graph construction module includes a knowledge extraction unit, a knowledge reasoning unit, and a dynamic update unit. The knowledge extraction unit is used to extract entity, entity relationship, and entity attribute data from the file text. The knowledge reasoning unit performs knowledge completion based on the TransE model and derives the implicit relationships between entities. The dynamic update unit integrates the feature vectors of new files and user-annotated data based on an incremental learning mechanism and updates the entities, relationships, and attributes in the knowledge graph.
[0014] Further improvements are made in that: the adaptive classification module includes a basic classification unit, an unknown category identification unit, and a user intent perception unit. The basic classification unit uses a multi-task Transformer model to classify files into preset categories. The unknown category identification unit uses the DBSCAN clustering algorithm combined with a confidence threshold to discover potential new categories and generate new category feature vectors. The user intent perception unit uses a proximal policy optimization reinforcement learning algorithm combined with user modification behavior of classification results to update the basic classification parameters.
[0015] Further improvements are made in the following aspects: The dynamic optimization module includes an incremental learning unit, a knowledge graph optimization unit, and a performance monitoring unit. The incremental learning unit updates the low-rank matrix of the attention layer of the basic classification based on LoRA technology. The knowledge graph optimization unit processes homonymous and similar relationship information based on entity alignment and relation disambiguation. The performance monitoring unit is used to statistically analyze the classification accuracy, recall, and F1 score in real time, and issue reminders to initiate manual intervention.
[0016] The beneficial effects of this invention are as follows: By integrating multi-dimensional features such as text semantics, visual content, and metadata association, this invention improves the classification accuracy of complex files; through knowledge extraction, reasoning, and dynamic updates, it constructs a knowledge system of files, entities, and relationships, solving the "rule solidification" problem of traditional methods and supporting rapid adaptation to new file types; at the same time, by combining reinforcement learning and incremental learning based on user feedback, it achieves continuous optimization of the classification model and knowledge graph, avoiding the problem of model obsolescence; through cluster analysis and knowledge graph association, it proactively discovers new file categories not defined by users, reducing the cost of manual rule maintenance. Attached Figure Description
[0017] Figure 1 This is a system architecture diagram of the present invention. Detailed Implementation
[0018] To enhance understanding of the present invention, the present invention will be further described in detail below with reference to embodiments. These embodiments are only used to explain the present invention and do not constitute a limitation on the scope of protection of the present invention.
[0019] according to Figure 1 As shown, this embodiment provides an intelligent classification and management system for computer files based on artificial intelligence, including a data acquisition and processing module, a multimodal feature extraction module, a knowledge graph construction module, an adaptive classification module, and a dynamic optimization module:
[0020] The data acquisition and processing module is used to acquire multi-source files and perform preprocessing and metadata extraction. It includes an acquisition caching unit, a data processing unit, and an information extraction unit. The acquisition caching unit is used to acquire batch files in real time and cache them. The data processing unit is used to standardize the format and denoise the files. The information extraction unit is used to extract the metadata information of the processed files.
[0021] Format standardization includes operations such as converting PDFs into editable text and videos into frame sequences; noise reduction includes processes such as deleting redundant pages in documents and restoring blurred images.
[0022] The data collection and caching unit collects data through local disk scanning and monitoring, cloud storage file downloads, and mail server file information extraction. Metadata information includes filename, creation time, modification history, author, project, and tag information.
[0023] The multimodal feature extraction module extracts and fuses multimodal data based on a BERT pre-trained model, ViT-Base, graph neural networks, and a gated attention mechanism. It includes a text feature extraction unit, a visual feature extraction unit, a metadata feature extraction unit, and a cross-modal fusion unit. The text feature extraction unit encodes text content using a BERT pre-trained model, generating a 768-dimensional semantic vector. The first 200 dimensions are then combined with TF-IDF features to capture local keyword information. The visual feature extraction unit uses ViT-Base to extract features from image and video frames, outputting a 768-dimensional visual embedding vector. The metadata feature extraction unit uses a graph neural network to model the relationships between metadata, generating a 512-dimensional structured association vector. These relationships include ternary relationships such as author-department-project. The cross-modal fusion unit uses a gated attention mechanism to perform weighted fusion of text, visual, and metadata features; the weighted fusion calculation formula is as follows:
[0024]
[0025] in W represents textual, visual, and metadata feature vectors, respectively. g W v Wm The weight matrix is a learnable matrix, σ is the activation function, and ⊙ represents element-wise multiplication, ultimately generating a 1536-dimensional global feature vector.
[0026] The knowledge graph construction module is used to build and update knowledge graphs based on file features and external knowledge bases. It includes a knowledge extraction unit, a knowledge reasoning unit, and a dynamic update unit. The knowledge extraction unit is used to extract entity, entity relationship, and entity attribute data from file text. Specifically, it uses the SpanBERT entity recognition model to extract entities such as contracts and R&D reports from the file text content, the PCNN+ATTENTION relationship extraction model to extract relationships such as association and belonging between entities, and the BiLSTM-CRF attribute extraction model to extract entity attributes such as security level and urgency.
[0027] The knowledge reasoning unit performs knowledge completion based on the TransE model and derives implicit relationships between entities, such as inferring the subsequent relationship between project weekly reports and project summaries; the dynamic update unit integrates feature vectors of new documents and user-annotated data based on an incremental learning mechanism and updates entities, relationships and attributes in the knowledge graph.
[0028] The adaptive classification module completes classification decisions and identifies unknown categories based on global features and knowledge graph association information, including basic classification units, unknown category identification units, and user intent perception units;
[0029] The basic classification unit uses a multi-task Transformer model to classify files into preset categories. The encoder of the multi-task Transformer model is BERT-large, and the decoder includes a category prediction head and a confidence prediction head. Its input is a global feature vector and a knowledge graph association vector, and its output is the probability distribution of the file belonging to each preset category.
[0030] The unknown category identification unit uses the DBSCAN clustering algorithm combined with a confidence threshold to discover potential new categories and generate new category feature vectors. Specifically, the confidence threshold is set to 0.7. For files with a probability lower than the threshold, the DBSCAN clustering algorithm with eps=0.5 and min_samples=3 is used to discover potential new categories and generate new category feature vectors.
[0031] The user intent perception unit updates the basic classification parameters based on the near-end policy optimization reinforcement learning algorithm and the user's modification behavior of the classification result, where the optimization objective is to minimize the number of user corrections.
[0032] The reinforcement learning training process of the user intent perception unit includes: defining the state space and the user's historical classification correction record; defining the action space and adjusting the weights of each modality feature of the base classifier; defining the reward function, with the reduction in the number of user corrections as the reward; and using the PPO algorithm to optimize the policy network, updating the policy parameters once every 50 interactions.
[0033] The dynamic optimization module is used to optimize the classification model and knowledge graph through user feedback and incremental learning. It includes an incremental learning unit, a knowledge graph optimization unit, and a performance monitoring unit. The incremental learning unit updates the low-rank matrix of the attention layer of the basic classification based on LoRA technology.
[0034] The knowledge graph optimization unit handles homonymous and similar relationship information based on entity alignment and relation disambiguation. Specifically, it solves the homonymous problem through entity alignment based on a weighted fusion of string similarity and embedding cosine similarity, and distinguishes similar relationships through relation disambiguation using the cross-entropy loss function.
[0035] The performance monitoring unit is used to statistically analyze classification accuracy, recall, and F1 score in real time, and issue reminders to initiate manual intervention. Specifically, when any indicator is below the threshold for three consecutive days (e.g., accuracy <90%), an alarm is triggered and the manual intervention process is initiated.
[0036] The intelligent file classification method of this AI-based computer file intelligent classification management system is as follows:
[0037] S1. File Acquisition and Preprocessing
[0038] The data acquisition and processing module acquires files from multiple storage media and uses OCR technology such as Tesseract to convert unstructured files into editable text.
[0039] Standardize file formats: convert video files to MP4 format, unify the resolution to 1080P, and convert document files to .docx format;
[0040] Extracting metadata: The creation time, modification time, and author information are obtained through the file system API of Python's os module, and the sender and recipient information are obtained through email header parsing.
[0041] S2: Multimodal Feature Extraction and Fusion
[0042] Text feature extraction: Input the text content into the BERT pre-trained model (maximum sequence length 512), take the output at the [CLS] position as a 768-dimensional semantic vector, and calculate the TF-IDF features at the same time. The stop word list is the NLTK English stop word list, and the ngram range is 1-2.
[0043] Visual feature extraction: Adjust the image and video frames to 224×224 pixels, input them into the ViT-Base model, with a patch size of 16×16 and a hidden layer dimension of 768. Take the [CLS] vector output by the last Transformer layer as the 768-dimensional visual embedding vector.
[0044] Metadata feature extraction: Construct a metadata graph where nodes are entities and edges are relations. Use a graph neural network model to perform graph convolution, and output node embeddings of 512-dimensional vectors. The graph neural network model has 8 heads and 256 hidden layer dimensions.
[0045] Cross-modal fusion: The weighted sum of features from each modality is calculated through a gated attention mechanism to generate a 1536-dimensional global feature vector.
[0046] S3: Knowledge Graph Association Reasoning
[0047] Query the dynamic knowledge graph, calculate the cosine similarity between the global feature vector of the current file and the 512-dimensional vectors embedded by all entities in the knowledge graph, and select the top 3 entities with the highest similarity as associated entities;
[0048] The confidence level of the relationship between entities is calculated based on the TransE model, and a 512-dimensional vector of association is generated.
[0049] When the confidence scores of all related entities retrieved are less than 0.5, the knowledge graph expansion process is triggered: relevant entities are retrieved from an external knowledge base and added to the knowledge graph.
[0050] During entity alignment, if entities with the same name but different meanings are found, they are distinguished by the "department" attribute in the metadata, and the entity attribute information is updated.
[0051] S4: Classification Decision and Confidence Assessment
[0052] The 1536-dimensional global feature vector and the 512-dimensional association vector are concatenated and then input into the multi-task Transformer model. The model has 12 encoder layers and 12 attention heads, and outputs the probability distribution of each preset category.
[0053] If the highest probability is greater than or equal to the confidence threshold of 0.7, it is determined to be the corresponding category;
[0054] If the confidence threshold is less than 0.7, the DBSCAN clustering algorithm is used to group similar files and generate new category candidates.
[0055] S5: User Feedback and Dynamic Optimization
[0056] Receive user instructions to correct the classification results and record correction information including the original classification, the corrected classification, and the timestamp.
[0057] The corrected samples were added to the incremental learning dataset, and the base classifier was fine-tuned using the LoRA technique with a learning rate of 1e-4 for 3 rounds of training.
[0058] Extract feature vectors and associated entities from new category files to update the knowledge graph;
[0059] The performance monitoring unit calculates the classification accuracy using the formula "number of correctly classified samples / total number of samples". If the accuracy is less than 90% for three consecutive days, it triggers a manual check of the knowledge graph and classifier parameters.
[0060] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A computer file intelligent classification and management system based on artificial intelligence, characterized in that: The system includes a data acquisition and processing module, a multimodal feature extraction module, a knowledge graph construction module, an adaptive classification module, and a dynamic optimization module. The data acquisition and processing module is used to acquire multi-source files and perform preprocessing and metadata extraction. The multimodal feature extraction module extracts multimodal data and performs cross-modal fusion based on the BERT pre-trained model, ViT-Base, graph neural network, and gated attention mechanism. The knowledge graph construction module is used to build and update the knowledge graph based on file features and external knowledge bases. The adaptive classification module completes classification decisions and identifies unknown categories based on global features and knowledge graph association information. The dynamic optimization module is used to optimize the classification model and knowledge graph through user feedback and incremental learning.
2. The computer file intelligent classification and management system based on artificial intelligence according to claim 1, characterized in that: The data acquisition and processing module includes an acquisition caching unit, a data processing unit, and an information extraction unit. The acquisition caching unit is used to acquire and cache batch files in real time. The data processing unit is used to perform format standardization and noise reduction on the files. The information extraction unit is used to extract the metadata information of the processed files.
3. The computer file intelligent classification and management system based on artificial intelligence according to claim 2, characterized in that: The data collection and caching unit collects data through local disk scanning and monitoring, cloud storage file download, and mail server file information extraction. The metadata information includes filename, creation time, modification history, author, project, and tag information.
4. The computer file intelligent classification and management system based on artificial intelligence according to claim 1, characterized in that: The multimodal feature extraction module includes a text feature extraction unit, a visual feature extraction unit, a metadata feature extraction unit, and a cross-modal fusion unit. The text feature extraction unit encodes the text content based on the BERT pre-trained model to generate a 768-dimensional semantic vector and extracts local keyword information by combining TF-IDF features. The visual feature extraction unit uses ViT-Base to extract features from image and video frames and outputs a 768-dimensional visual embedding vector. The metadata feature extraction unit generates a 512-dimensional structured association vector based on graph neural networks to model the relationship between metadata. The cross-modal fusion unit performs weighted fusion of text, visual, and metadata features based on a gating attention mechanism.
5. The computer file intelligent classification and management system based on artificial intelligence according to claim 4, characterized in that: The weighted fusion calculation formula is as follows: in W represents textual, visual, and metadata feature vectors, respectively. g W v W m The weight matrix is a learnable matrix, σ is the activation function, and ⊙ represents element-wise multiplication, ultimately generating a 1536-dimensional global feature vector.
6. The computer file intelligent classification and management system based on artificial intelligence according to claim 1, characterized in that: The knowledge graph construction module includes a knowledge extraction unit, a knowledge reasoning unit, and a dynamic update unit. The knowledge extraction unit is used to extract entity, entity relationship, and entity attribute data from the file text. The knowledge reasoning unit performs knowledge completion based on the TransE model and derives the implicit relationships between entities. The dynamic update unit integrates the feature vectors of new files and user-annotated data based on an incremental learning mechanism and updates the entities, relationships, and attributes in the knowledge graph.
7. The computer file intelligent classification and management system based on artificial intelligence according to claim 1, characterized in that: The adaptive classification module includes a basic classification unit, an unknown category identification unit, and a user intent perception unit. The basic classification unit uses a multi-task Transformer model to classify files into preset categories. The unknown category identification unit uses the DBSCAN clustering algorithm combined with a confidence threshold to discover potential new categories and generate new category feature vectors. The user intent perception unit uses a proximal policy optimization reinforcement learning algorithm combined with user modification behavior of classification results to update the basic classification parameters.
8. The computer file intelligent classification and management system based on artificial intelligence according to claim 1, characterized in that: The dynamic optimization module includes an incremental learning unit, a knowledge graph optimization unit, and a performance monitoring unit. The incremental learning unit updates the low-rank matrix of the attention layer of the basic classification based on LoRA technology. The knowledge graph optimization unit processes synonyms and similarity information based on entity alignment and relation disambiguation. The performance monitoring unit is used to statistically analyze classification accuracy, recall, and F1 score in real time and issue reminders to initiate manual intervention.
Citation Information
Cited By
Network additional storage intelligent data management method and system
CN122045162A