Multi-modal file content automatic classification and theme combination method based on artificial intelligence

By processing multimodal files through AI-based deep learning models and clustering algorithms, the problems of single file format and low content retrieval efficiency in the existing system are solved, intelligent management and efficient merging of multimodal files are realized, and the level of intelligent information management is improved.

CN120632096APending Publication Date: 2025-09-12林文博
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510712016.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing personal knowledge management and enterprise archive management systems have problems such as single file format, low content retrieval efficiency, and inability to automatically classify and merge. The management difficulty increases especially when multimodal data is widely used.

Method used

Adopting AI-based methods, we process files of various formats through deep learning models, extract text and image information, use NLP and CNN models for feature extraction, combine K-means and DBSCAN algorithms for topic clustering and content merging, and realize intelligent management of multimodal files.

Benefits of technology

It achieves efficient subject classification and content merging of multi-format and multi-modal files, improves the intelligence level of information management, and is suitable for personal and corporate scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a multi-modal file content automatic classification and theme combination method and system based on artificial intelligence, and belongs to the technical field of intelligent information management and human-computer interaction. The system comprises a front-end user interaction module, a rear-end data processing module and an AI intelligent classification module. A user can upload files (including PDF, Word, pictures and the like) in various formats through a front-end interface, and a system automatically extracts text and picture information in the files and performs semantic understanding, theme classification and similar content combination on contents by utilizing an AI algorithm. A classification result is stored in a memory library in a structured mode, and multi-picture transverse display, content editing, deletion and unified management of multi-format files are supported. The system also has the functions of invalid picture path automatic cleaning, import progress visualization, exception handling and the like, and the management efficiency and the user experience of the multi-source heterogeneous information are remarkably improved. The method can be widely applied to personal knowledge management, enterprise archive arrangement, intelligent assistant and other scenes, and has high innovativeness and practical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and information management technology, and specifically to a method and system for automatic classification and topic merging of multimodal file content based on AI. Background Art

[0002] Existing personal knowledge management and enterprise archive management systems often suffer from monotonous file formats, low content retrieval efficiency, and an inability to automatically classify and merge data. With the widespread use of multimodal data (text, images, documents, etc.), efficient management and intelligent classification have become urgent technical challenges that need to be addressed. Summary of the Invention

[0003] The present invention proposes an AI-based method and system for automatic classification and topic merging of multimodal file content, which can automatically process files of various formats, extract text and image information therein, and realize semantic understanding of content, topic clustering and merging of similar content through deep learning models, greatly improving the intelligence and automation level of information management. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] · Figure 1 :System structure diagram

[0005] · Figure 2 : Multimodal content processing flow chart

[0006] · Figure 3 :Flowchart of topic clustering and merging algorithm

[0007] · Figure 4 :Display of topic clustering and merging results DETAILED DESCRIPTION

[0008] 1. File upload and content extraction

[0009] Users upload PDF, Word, pictures and other files through the front-end interface, and the system automatically identifies the file type.

[0010] Tools such as pdfplumber, docx, and OCR are used to extract text and image information.

[0011] 2. Multimodal feature extraction

[0012] NLP models such as BERT are used to vectorize the text content, and CNN models are used to extract visual features from the image content. Finally, the two are spliced ​​into a unified multimodal feature vector.

[0013] The multimodal feature fusion formula is as follows:

[0014] v i =α·f NLP(T i )+β·f CNN (I i )

[0015] Among them, α and β are weight coefficients.

[0016] 3. AI classification and topic merging

[0017] Clustering algorithms such as K-means and DBSCAN are used to classify all memory entries by topic, and similar content is merged. The merged results are stored in a structured JSON format.

[0018] The topic clustering objective function is as follows:

[0019]

[0020] 4. Front-end display and interaction

[0021] The front-end page supports horizontal arrangement of multiple pictures, click-to-enlarge preview, content editing and deletion, visualization of the import process, and friendly prompts for empty states.

[0022] 5. Automatically clean up invalid image paths

[0023] The backend regularly checks the validity of image paths in the database and automatically deletes invalid paths to ensure that the images rendered by the frontend are real.

[0024] Beneficial effects

[0025] The present invention can automatically process multi-format and multi-modal file contents, realize efficient subject classification and content merging, greatly improve the intelligent level of information management, and is suitable for various scenarios such as individuals and enterprises, with broad application prospects and market value.

Claims

1. A method for automatic classification and topic merging of multimodal document contents based on artificial intelligence, characterized in that: The steps include: a) Receive multi-format files uploaded by users, including but not limited to PDF, Word, images, etc.; b) extracting the contents of the file and obtaining text and image information using text parsing and image recognition technology; c) Based on natural language processing (NLP) and multimodal deep learning models, semantic understanding and feature vectorization of the extracted content are performed. The formula is as follows: v i =f NLP (T i )+f CNN (I i ) Among them, v i is the multimodal feature vector of the i-th memory, T i For text content, I i is the image content, f NLP is the text feature extraction function, f CNN is the image feature extraction function; d) The feature vectors of all memory items are clustered and similar content is merged using the following clustering algorithm: Among them, C k is the kth topic cluster, μ k as its center; e) The classified and merged results are stored in a structured manner, and support front-end horizontal display of multiple pictures, content editing and deletion.

2. The method according to claim 1, characterized in that The content extraction step includes performing OCR recognition on the image file to automatically extract text information from the image.

3. The method according to claim 1, characterized in that The system has an automatic cleanup function for invalid image paths, regularly checks the validity of image paths in the database, and deletes invalid paths.

4. The method according to claim 1, wherein The front-end interface supports visualization of the import progress and displays the file processing progress in real time.

5. A system for implementing the above method includes a front-end user interaction module, a back-end data processing module and an AI intelligent classification module, which work together through network communication.