Multi-modal file content automatic classification and theme combination method based on artificial intelligence
By processing multimodal files through AI-based deep learning models and clustering algorithms, the problems of single file format and low content retrieval efficiency in the existing system are solved, intelligent management and efficient merging of multimodal files are realized, and the level of intelligent information management is improved.
Patent Information
- Application Number
- CN202510712016.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-12
AI Technical Summary
Existing personal knowledge management and enterprise archive management systems have problems such as single file format, low content retrieval efficiency, and inability to automatically classify and merge. The management difficulty increases especially when multimodal data is widely used.
Adopting AI-based methods, we process files of various formats through deep learning models, extract text and image information, use NLP and CNN models for feature extraction, combine K-means and DBSCAN algorithms for topic clustering and content merging, and realize intelligent management of multimodal files.
It achieves efficient subject classification and content merging of multi-format and multi-modal files, improves the intelligence level of information management, and is suitable for personal and corporate scenarios.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and information management technology, and specifically to a method and system for automatic classification and topic merging of multimodal file content based on AI. Background Art
[0002] Existing personal knowledge management and enterprise archive management systems often suffer from monotonous file formats, low content retrieval efficiency, and an inability to automatically classify and merge data. With the widespread use of multimodal data (text, images, documents, etc.), efficient management and intelligent classification have become urgent technical challenges that need to be addressed. Summary of the Invention
[0003] The present invention proposes an AI-based method and system for automatic classification and topic merging of multimodal file content, which can automatically process files of various formats, extract text and image information therein, and realize semantic understanding of content, topic clustering and merging of similar content through deep learning models, greatly improving the intelligence and automation level of information management. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] · Figure 1 :System structure diagram
[0005] · Figure 2 : Multimodal content processing flow chart
[0006] · Figure 3 :Flowchart of topic clustering and merging algorithm
[0007] · Figure 4 :Display of topic clustering and merging results DETAILED DESCRIPTION
[0008] 1. File upload and content extraction
[0009] Users upload PDF, Word, pictures and other files through the front-end interface, and the system automatically identifies the file type.
[0010] Tools such as pdfplumber, docx, and OCR are used to extract text and image information.
[0011] 2. Multimodal feature extraction
[0012] NLP models such as BERT are used to vectorize the text content, and CNN models are used to extract visual features from the image content. Finally, the two are spliced into a unified multimodal feature vector.
[0013] The multimodal feature fusion formula is as follows:
[0014] v i =α·f NLP(T i )+β·f CNN (I i )
[0015] Among them, α and β are weight coefficients.
[0016] 3. AI classification and topic merging
[0017] Clustering algorithms such as K-means and DBSCAN are used to classify all memory entries by topic, and similar content is merged. The merged results are stored in a structured JSON format.
[0018] The topic clustering objective function is as follows:
[0019]
[0020] 4. Front-end display and interaction
[0021] The front-end page supports horizontal arrangement of multiple pictures, click-to-enlarge preview, content editing and deletion, visualization of the import process, and friendly prompts for empty states.
[0022] 5. Automatically clean up invalid image paths
[0023] The backend regularly checks the validity of image paths in the database and automatically deletes invalid paths to ensure that the images rendered by the frontend are real.
[0024] Beneficial effects
[0025] The present invention can automatically process multi-format and multi-modal file contents, realize efficient subject classification and content merging, greatly improve the intelligent level of information management, and is suitable for various scenarios such as individuals and enterprises, with broad application prospects and market value.
Claims
1. A method for automatic classification and topic merging of multimodal document contents based on artificial intelligence, characterized in that: The steps include: a) Receive multi-format files uploaded by users, including but not limited to PDF, Word, images, etc.; b) extracting the contents of the file and obtaining text and image information using text parsing and image recognition technology; c) Based on natural language processing (NLP) and multimodal deep learning models, semantic understanding and feature vectorization of the extracted content are performed. The formula is as follows: v i =f NLP (T i )+f CNN (I i ) Among them, v i is the multimodal feature vector of the i-th memory, T i For text content, I i is the image content, f NLP is the text feature extraction function, f CNN is the image feature extraction function; d) The feature vectors of all memory items are clustered and similar content is merged using the following clustering algorithm: Among them, C k is the kth topic cluster, μ k as its center; e) The classified and merged results are stored in a structured manner, and support front-end horizontal display of multiple pictures, content editing and deletion.
2. The method according to claim 1, characterized in that The content extraction step includes performing OCR recognition on the image file to automatically extract text information from the image.
3. The method according to claim 1, characterized in that The system has an automatic cleanup function for invalid image paths, regularly checks the validity of image paths in the database, and deletes invalid paths.
4. The method according to claim 1, wherein The front-end interface supports visualization of the import progress and displays the file processing progress in real time.
5. A system for implementing the above method includes a front-end user interaction module, a back-end data processing module and an AI intelligent classification module, which work together through network communication.