Intelligent literature management and classification system based on big data
By using a big data-based intelligent document management and classification system, and by employing document interpretation and search models, a document classification framework diagram and associated catalog table are established. This solves the problem of excessive and low-relevance search results in traditional document retrieval methods, and achieves efficient and accurate document management and classification.
Patent Information
- Application Number
- CN202511193201.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional document retrieval methods produce too many results with low relevance when processing large-scale documents, making it difficult to meet user needs. Existing document classification and management systems are not precise enough and cannot quickly find relevant documents.
We adopt a big data-based intelligent document management and classification system. We interpret document data through a document interpretation model, establish a document classification framework diagram and related catalog table, use a document search model to search and download documents from online document databases, and combine document management and classification modules for efficient management and classification.
It achieves high-precision document classification, improves document relevance and search efficiency, enables quick retrieval of relevant documents, and meets users' precise search needs.
Smart Images

Figure CN120804320A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of document management and classification, and particularly relates to a document intelligent management and classification system based on big data. BACKGROUND
[0002] With the rapid growth of document information, users have higher and higher demands for accurate retrieval. Traditional document retrieval methods have too many retrieval results and low relevance when dealing with large-scale documents, and it is difficult to meet the needs of users. In the fields of academic research and technical intelligence mining, it is crucial to efficiently and accurately retrieve and screen relevant documents. When a large number of documents related to a target theme are needed, traditional document retrieval methods mainly rely on keyword matching or rule-based classification systems. However, these methods have obvious limitations. Therefore, it is necessary to intelligently manage and classify documents in order to quickly find relevant documents and analyze documents. However, existing document classification management is only a simple field and theme classification, and there is no fine classification management for document relevance and research value. Most importantly, the classification information is not clear enough, and it is difficult to quickly find the required document information. The data used for analyzing and managing and classifying documents is small, and the classification is not accurate enough.
[0003] Therefore, we provide a document intelligent management and classification system based on big data to solve the above problems. SUMMARY
[0004] The purpose of the present application is to provide a document intelligent management and classification system based on big data, which interprets document data from each document through a document interpretation model, so as to construct a document classification framework graph and generate a document correlation directory table based on a document classification framework model architecture, thereby solving the problems of inconvenience in searching after classification, inability to quickly find, poor correlation after classification, small amount of document data for classification, and low classification accuracy.
[0005] To solve the above technical problems, the present application is realized by the following technical scheme:
[0006] The present application is a document intelligent management and classification system based on big data, which comprises a document interpretation model, a document management and classification module, a document search model, a database, and a document correlation directory table.
[0007] The document interpretation model is used to interpret the documents input into the system, and interpret the document data of the theme, author information, research direction, research method, and research conclusion from the documents.
[0008] The literature management and classification module comprises a literature information storage management unit and a literature classification framework model, the literature information storage management unit is used for analyzing and saving literature data interpreted by the literature interpretation model, the literature information storage management unit is used for saving the literature and the literature data to the database or deleting the literature and only saving the literature data to the database, and the literature classification framework model is used for constructing a literature classification framework graph for the literature associated with the literature data in the database and searching the literature data in the database and establishing the association of the literature data;
[0009] The literature search model is used for searching and downloading the literature with the literature data from various network literature libraries;
[0010] The database is used for storing the literature and the literature data;
[0011] The literature association directory table is used for saving the association information of the cross-field associated literature.
[0012] The literature information storage management unit is further used for analyzing the research value of the literature, the literature information storage management unit compares and analyzes the similar literature stored in the database, sets a score for the research value, if the score of the analysis of the research value of the literature is lower than a set value, the literature is not stored in the database, if the score is higher than or equal to the set value, the literature is stored in the database.
[0013] The literature classification framework model selects whether to save the original text of the literature in the database according to the score of the research value of the literature, only the original text of the literature with a high research value is selected to be saved, and only the literature data interpreted from the literature is saved for the literature with a research value lower than the set value.
[0014] In the process of searching for the literature, when the found literature only has the literature data, the literature search model is used to search and download the literature from the network.
[0015] The literature interpretation model is further used for interpreting and analyzing the literature data from each piece of literature when the database for managing the literature is just established.
[0016] The literature classification framework model sorts and classifies each piece of literature data, classifies the literature data according to the research field, the research theme and the research direction, and then constructs a literature classification framework graph.
[0017] The application is further configured to structure the subject, field, research direction and research method of a document on the document classification framework, and other document data of the document is structured to the document classification framework in a hyperlink manner, and the research method on the document classification framework only displays the method category, and the research method specific operation details and research results are also set with hyperlinks at the research method position.
[0018] The application is further configured to perform correlation searching on each document after the document classification framework model finishes classifying each document data, to form a correlation document table of each document, and then compile the correlation document table of each document into a document correlation directory table.
[0019] The application is further configured to classify the new document data according to the research field, research subject and research direction, and restructure and layout the new document to the similar document structure position on the document classification framework graph each time a new document is input.
[0020] If there is no similar document on the document classification framework graph, a new section is established on the document classification framework graph to structure the new document information.
[0021] The application is further configured to have a function of searching related documents from each document library, and when more related document information is needed for comparison after a document is interpreted, document storage processing and document framework processing, the related documents are searched from each document library through the document search model, the related documents searched from the document library are interpreted, and the document data interpreted from the related documents searched from the document library are structured through the document classification framework model.
[0022] The application is further configured to set a hyperlink for the correlation document table of each document in the document correlation directory table on the document classification framework graph established by the document classification framework model, and the document correlation directory table can be opened and the position of the correlation document table is displayed by clicking the hyperlink of the correlation document table of the corresponding document on the document classification framework graph.
[0023] The application has the following beneficial effects:
[0024] 1、The present application interprets each document, interprets the document data of the theme, author information, research direction, research method and research conclusion, and the document classification framework model uses the document data of each document to construct a document classification framework graph, which can clearly and concisely display the information of each document and the information of the associated documents, and the same block position displays the document data and the document data of the associated documents, which can effectively classify the information between the documents and the associated documents, and also analyze and interpret the research value of the documents, and through the research value of the documents, high-value documents can be efficiently found.
[0025] 2、The document classification of the present application is based on a large amount of documents for classification and management, and the classification precision is high, and the present application can also search a plurality of existing document libraries to associate a certain document to be classified, so that the document classification association value is higher and the association effect is better.
[0026] Of course, any product implementing the present application does not necessarily need to achieve all the advantages described above. BRIEF DESCRIPTION OF DRAWINGS
[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments, and obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings.
[0028] Figure 1 It is a principle diagram of a literature intelligent management and classification system based on big data.
[0029] Figure 2 It is a principle diagram of a literature classification framework graph. DETAILED DESCRIPTION
[0030] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application, and obviously, the described embodiments are only some embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0031] Please refer to Figure 1-2 , the present application is a literature intelligent management and classification system based on big data, which comprises a document interpretation model, a document management and classification module, a document search model, a database and a document association directory table;
[0032] The literature interpretation model is used for interpreting the literature input into the system, interpreting the literature data of the literature into topics, author information, research direction, research method and research conclusion; the author information includes author name (including corresponding author), publishing unit and publishing date.
[0033] The literature management and classification module includes a literature information warehouse management unit and a literature classification framework model, the literature information warehouse management unit is used for analyzing and saving the literature data interpreted by the literature interpretation model, the literature information warehouse management unit is used for saving the literature and the literature data to the database or deleting the literature and only saving the literature data to the database, and the literature classification framework model is used for constructing a literature classification framework graph for the literature data associated in the database and searching the literature data in the database and establishing the association of the literature data;
[0034] The literature information warehouse management unit needs to audit the literature data interpreted by the literature, and the literature data can be stored in the warehouse after meeting the basic requirements, but whether the text of the literature is saved to the database needs to be evaluated for the value of the literature or manually selected whether to save.
[0035] The literature to be saved to the database will be a large literature library with large memory occupation, therefore, part of the literature can only save the literature data (interpreted literature data) of the literature, and the text of the literature itself is not saved, because the literature search model can be used to search and download the literature for later search, and only the literature with high value or needing to be saved to the database is selected to be saved.
[0036] The literature classification framework graph can clearly show the position of each literature and the related literature data, opening the literature classification framework model interface in the system, inputting the related information, marking the literature position related to the information on the literature classification framework graph, and searching for the required literature position by switching the position.
[0037] The literature search model is used for searching and downloading the literature with literature data from various network literature libraries; some literature does not save the literature text when stored in the warehouse, in order to save the database capacity, and the literature search model can be used to search the literature (download from the literature library, literature journal, etc. on the network) for later search.
[0038] The database is used for storing literature and literature data; it is also used for saving the programs of the literature interpretation model, the literature management and classification module and the literature search model, and the literature association directory table.
[0039] The literature association directory table is used for saving the association information of the cross-field associated literature.
[0040] The literature information warehousing management unit is also used for analyzing the research value of the literature, and the literature information warehousing management unit compares and analyzes similar literatures stored in the database, sets a score for the research value, and if the score of the analysis of the research value of the literature is lower than the set value, the literature is not warehoused, and if the score is higher than or equal to the set value, the literature is stored.
[0041] The literature value is mainly evaluated by the research direction, research method and research conclusion (result), for example, the research direction is common, the method is commonly used, and the conclusion has no outstanding change (compared with related literatures), and the literature is evaluated as a low-value literature, and the evaluation points are formed in multiple directions, and the evaluation scores are set for each evaluation point, and then the scores of the multiple evaluation points are added to obtain the value score.
[0042] The literature classification framework model screens whether the original text of the literature is stored in the database according to the literature research value score, and only the original text of the literature and the literature data interpreted from the literature are selected to be stored for the literature with a high research value, and only the literature data interpreted from the literature are kept for the literature with a low research value but higher than the set value.
[0043] In the process of searching for the literature, the searched literature only has the literature data, and the literature search model is downloaded from the network search.
[0044] Some literatures have a value, but are not particularly high, and are convenient for later downloading, and the text can not be stored to save the database capacity, and some student project texts have a large text memory, and can be selected not to be stored.
[0045] When the literature management database is established, the literature interpretation model analyzes and interprets the literature data from each literature;
[0046] The literature classification framework model sorts and classifies each literature data, classifies according to the research field, research theme and research direction, and then structures the literature classification framework graph.
[0047] When the management and classification system is established, the literature interpretation model, the literature management and classification module and the literature search model are first established in the database, but the large amount of literature at the beginning needs to be interpreted, warehoused and screened, and then the literature classification framework model is used to structure the literature classification framework graph, and then the literature correlation processing is performed to form the literature correlation table of each literature and is concentrated in the literature correlation directory table.
[0048] The theme, field, research direction and research method of a literature are structured on the literature classification framework graph, and other literature data of the literature is structured on the literature classification framework graph in a hyperlink manner, and only the method category is displayed on the research method of the literature classification framework graph, and the research method specific operation details and research results are also set in the research method position.
[0049] Some of the content of the literature data is not convenient to display on the literature classification framework diagram, which can be opened through a hyperlink from the database literature data. When you see the location of the literature distribution, you can see the more detailed content of the literature.
[0050] After the literature classification framework model sorts and classifies each piece of literature data, it performs an associated search on each piece of literature. Each piece of literature is associated with a search, forming an associated literature table for each piece of literature, and then each piece of literature is associated with a literature table. The literature association directory table is compiled into the literature association directory table.
[0051] Literature association is mainly for cross-discipline, same field or similar field, literature in literature classification framework diagram location close, can be seen directly, cross-discipline need to find relevant literature from literature association directory table, through literature classification framework model to find.
[0052] Each time a new literature is input, the literature classification framework model classifies the new literature data, classifies it according to the research field, research topic and research direction, and restructures the similar literature architecture position to the literature classification framework diagram;
[0053] If the new literature input has no similar literature on the literature classification framework diagram, a new section will be established on the literature classification framework diagram to structure the new literature information.
[0054] For new literature, it needs to be interpreted, warehoused and structured every time. For new fields or no related literature (in the database), a new section needs to be established as the new literature on the literature classification framework diagram.
[0055] The literature search model also has the function of searching related literature from various literature libraries. When more related literature information is needed for comparison after a piece of literature is interpreted, literature warehousing and literature framework processing, the literature search model is used to search related literature from various literature libraries. The related literature searched from the literature library is interpreted, and the literature data interpreted from the literature library is structured by the literature classification framework model.
[0056] Sometimes, the capacity of the database is not large for some literature comparison or related literature analysis, so it is necessary to search and download from the literature library (on the network), and then interpret, warehouse (or not) and analyze (literature classification framework model analysis) to evaluate the relevance of the new literature (with existing literature, including database and network literature library).
[0057] The associated document table of each document in the document association directory table is provided with a hyperlink on the document classification framework graph established by the classification framework model, and the document association directory table can be opened and the position of the associated document table is displayed by clicking the hyperlink of the associated document table of the corresponding document on the document classification framework graph.
[0058] In the classification framework model, the relevant information is input, the relevant documents are searched, and the position of the better document is found. The document information related to the document can be viewed by clicking the hyperlink, that is, the position on the document association directory table is opened and viewed.
[0059] In the description of the present specification, the description of the terms "one embodiment", "example", "specific example" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are contained in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0060] The preferred embodiments of the application disclosed above are only used to help explain the application. The preferred embodiments do not describe all the details and limit the application to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the present specification. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the application, so that those skilled in the art can well understand and utilize the application. The application is limited only by the claims and their full scope and equivalents.
Claims
1. A big data-based intelligent document management and classification system, characterized by: It includes literature interpretation model, literature management and classification module, literature search model, database and literature association catalog; The document interpretation model is used to interpret the documents input into the system and to interpret the document data including the subject, author information, research direction, research method and research conclusion; The document management and classification module includes a document information storage management unit and a document classification framework model. The document information storage management unit is used to analyze the document data interpreted by the document interpretation model and save the document data. The document information storage management unit is used to save documents and document data to the database or delete documents and only save document data to the database. The document classification framework model is used to construct a document classification framework diagram for documents associated with document data in the database and to search for document data in the database and establish associations between document data. The document search model is used to search and download documents with document data from major network document repositories; The database is used to store documents and document data; The document association catalog table is used to store association information of cross-domain associated documents.
2. The intelligent document management and classification system based on big data according to claim 1 is characterized in that: The document information storage management unit is also used to analyze the research value of documents. The document information storage management unit uses similar documents stored in the database for comparative analysis and sets a score for the research value. If the score of the analyzed document research value is lower than the set value, it will not be stored in the database. If it is higher than or equal to the set value, it will be stored in the database.
3. The intelligent document management and classification system based on big data according to claim 2 is characterized in that: The document classification framework model screens whether the original text of the document is stored in the database according to the document research value score. Only the original text of the document with high research value will be selected to save the document and the document data derived from the document. For documents with low research value but higher than the set value, only the document data derived from the document will be retained. In the process of searching for documents, when the documents found only have document data, they are downloaded from the Internet through the document search model.
4. The intelligent document management and classification system based on big data according to claim 1 is characterized in that: When the document management database is just established, the document interpretation model interprets and analyzes each document to obtain document data; The document classification framework model organizes and classifies the data of each document, classifies them according to research fields, research topics and research directions, and then constructs a document classification framework diagram.
5. The intelligent document management and classification system based on big data according to claim 4 is characterized in that: The subject, field, research direction and research method of a document are structured on the document classification framework diagram, and other document data of the document are structured on the document classification framework diagram in a hyperlinked manner. The research method on the document classification framework diagram only displays the method category, and the specific operation details and research results of the research method are also hyperlinked at the research method position.
6. The intelligent document management and classification system based on big data according to claim 4 is characterized in that: After the document classification framework model finishes sorting and classifying the data of each document, it then performs an associated search on each document, forms an associated document table for each document, and then compiles the associated document table of each document into the document associated catalog table.
7. The intelligent document management and classification system based on big data according to claim 4 is characterized in that: Each time a new document is input, the document classification framework model classifies the new document data, classifies it into the framework position of similar documents according to research fields, research topics and research directions, and restructures and typesets it on the document classification framework diagram; If the new document input has no similar documents on the document classification framework diagram, a new section will be created on the document classification framework diagram to structure the new document information.
8. The intelligent document management and classification system based on big data according to claim 1 is characterized in that: The document search model also has the function of searching for relevant documents from various document libraries. When more relevant document information is needed for comparison after interpreting, storing and framing a document, the document search model is used to search for relevant documents from various document libraries, interpret the relevant documents searched from the document libraries, and structure the document data interpreted from searching for relevant documents from the document libraries through the document classification framework model.
9. The intelligent document management and classification system based on big data according to claim 1 is characterized in that: The associated document table of each document in the document association catalog table is provided with a hyperlink on the document classification framework diagram established by the classification framework model. By clicking the hyperlink of the associated document table of the corresponding document on the document classification framework diagram, the document association catalog table can be opened and the location of the associated document table can be displayed.