A file classification and management system based on big data
By designing an archive classification management system based on big data, the problem of large amount of archive data and inaccurate search results in traditional archive management systems is solved, efficient archive classification and intelligent recommendation are achieved, and the system usage effect and intelligence level are improved.
Patent Information
- Application Number
- CN202510315301.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-18
AI Technical Summary
In the process of digitizing archives, traditional archive management systems have problems such as large amount of data, complex information, and low accuracy of search results, resulting in poor system use.
An archive classification management system based on big data is designed, including archive data collection module, data storage module, archive classification module, search recommendation module and visualization module. Classified storage is performed according to the access frequency and importance of archival data through hierarchical storage units, and intelligent recommendation of archives is performed using metadata and similarity analysis.
It improves the accuracy of archive retrieval and the quality of the system use, reduces the construction cost of the archive classification management system, and improves the intelligence level and user experience of the system.
Smart Images

Figure CN119847998B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of file management, and particularly to a file classification management system based on big data. Background Art
[0002] With the development of modern civilized society, "files" have long become something that people are familiar with. It appears in people's life, study, and work, and runs through various aspects such as scientific research, medical treatment, and litigation. In order to ensure the integrity, originality, etc. of files, the work of file classification management has emerged. With the development of big data technology, in various technical fields, information has gradually been integrated through Internet technology to share resources. File information has the characteristics of large data volume, complex information, and various types. Traditional file information mainly consists of paper files. The collection, classification, and storage of a large amount of file information bring huge workload and pressure to staff.
[0003] With the development of technology, file digitization has become a trend. However, in the current process of file digitization, usually after files are scanned and recognized, they are all stored together. Due to the huge file data and the high cost of high-performance storage devices, the system construction cost is increased. And when retrieving file data, there is a lack of comprehensive analysis of file data and the search results cannot be adjusted, resulting in low accuracy of retrieval results and reducing the use effect of the file classification management system. Summary of the Invention
[0004] The purpose of the present invention is to provide a file classification management system based on big data, which solves the problems raised in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A file classification management system based on big data, including a file data acquisition module, a data storage module, a file classification module, a retrieval recommendation module, and a visualization module;
[0006] The file data acquisition module is used to obtain file data from a data source;
[0007] The data storage module is used to store the data collected by the file data acquisition module. The data storage module includes a hierarchical storage unit, and the hierarchical storage unit is used to divide file data into cold files and hot files and store them using different storage media;
[0008] The file classification module uses a metadata collection tool to obtain the metadata of file data and generates a label for each file data. The label includes time, keywords, data source, and category, and classifies the file data according to the label;
[0009] The retrieval and recommendation module includes a result recommendation unit and a relevance analysis unit. The result recommendation unit is used to provide specific files according to the content input by the user, and the relevance analysis unit compares the files provided by the result recommendation unit with other files and provides relevant files for the user according to the comparison results.
[0010] Optionally, the file data includes paper files and electronic files. The electronic files include audio-visual data. The file data acquisition module includes an optical recognition unit and an audio conversion unit. The optical recognition unit is used to convert paper files into image format and extract the text content in the image format. The audio conversion unit is used to convert audio-visual data, extract text data, and perform digital processing in combination with video transcoding technology.
[0011] Optionally, the process of dividing the file data by the hierarchical storage unit is as follows:
[0012] ;
[0013] where A is the heat score;
[0014] F is the file access frequency, indicating the number of accesses in the past unit time;
[0015] α is the access frequency influence coefficient, and its value range is from 0 to 1;
[0016] T is the time interval since the last access;
[0017] β is the access time interval influence coefficient, and its value range is from 0 to 1;
[0018] Z is the file importance weight, and its value range is from 0 to 1;
[0019] The higher the heat score A, the more frequently the file data is used. On the contrary, it means that the file data is used less. The classification threshold of the heat score A is set as Y1. When A is greater than Y1, it is a hot file. When A is less than Y1, it is a cold file. For hot files, they are moved to a high-performance storage medium to improve access speed and response time. For cold files, they are moved to a low-cost storage medium to reduce storage costs.
[0020] Optionally, the process of the result recommendation unit is as follows:
[0021] ;
[0022] where G represents the user's query content and S represents the candidate score;
[0023] S(D 1 ,G) represents the candidate score of file one according to the user's query content;
[0024] R(D 1 , G) represents the similarity score between the user's query content and Archive 1;
[0025] W1 represents the influence coefficient of R(D 1 , G), and its value range is from 0 to 1;
[0026] H(D 1 ) represents the historical access frequency of Archive 1;
[0027] W2 represents the influence coefficient of H(D 1 ), and its value range is from 0 to 1;
[0028] C(D 1 ) represents the timeliness of Archive 1;
[0029] W3 represents the influence coefficient of C(D 1 ), and its value range is from 0 to 1;
[0030] The process of obtaining the historical access frequency H(D 1 ) of Archive 1 is as follows:
[0031] ;
[0032] where U represents the set of users;
[0033] M p (D 1 ) represents the number of times user p accesses Archive 1;
[0034] M p represents the influence coefficient of user p, and its value range is from 0 to 1, representing the importance of different users;
[0035] S(D 1 , G) the larger it is, the stronger the relevance between Archive 1 and the user's query content. When the user inputs the user query content G during the use of the system, the search result column will automatically recommend the archive with the highest score to ensure the accuracy of archive query. The user can adjust W1, W2, and W3 according to needs to improve the flexibility of archive query.
[0036] Optionally, the analysis process of the correlation analysis unit is as follows:
[0037] ;
[0038] where R(D 1 , D 2 ) represents the similarity score between Archive 1 and Archive 2;
[0039] F i (D 1) represents the frequency of occurrence of the i-th keyword in Archive 1;
[0040] F i (D 2 ) represents the frequency of occurrence of the i-th keyword in Archive 2;
[0041] n represents the number of keywords;
[0042] represents the sum of the minimum frequencies of shared features between Archive 1 and Archive 2, used to calculate how much overlap there is between the two archives for each keyword and to find the minimum value for each keyword;
[0043] represents the maximum value of the total number of features in Archive 1 and Archive 2;
[0044] By calculating the degree of overlap of shared features between the two archives and normalizing it with the maximum value of their total number of features, the correlation between the two archives is measured to obtain a similarity score. The correlation analysis unit can determine the similarity of the two archives in terms of keywords, helping to find similar archives and compare the contents of archives. After the result recommendation unit recommends the archive with the highest candidate score S to the user, the correlation analysis unit will calculate the similarity score between the archive with the highest candidate score S and other archives, and set the 1 ,D 2 ) scoring threshold as Y2. When R(D 1 ,D 2 ) is greater than Y2, it indicates that the similarity between the two archives is high, and the archive compared with the archive with the highest score will be added to the search result column. When R(D 1 ,D 2 ) is less than Y2, it will not be added.
[0045] Optionally, the visualization module provides monitoring data on the system operation for the administrator. The monitoring data includes the archive storage volume, access frequency, and system hardware operation parameters.
[0046] Optionally, the high-performance storage medium includes solid-state drives, and the low-cost storage medium includes mechanical hard drives.
[0047] Optionally, the metadata collection tool is ApacheTika.
[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0049] 1. After classifying the archival data, when the user conducts a search, the result recommendation unit in the retrieval recommendation module recommends the most suitable archives for the user based on the input content. The result recommendation unit includes influencing factors such as historical access frequency, timeliness, and similarity to the user input content, realizing a comprehensive analysis of the archival data. In actual operation, the user can adjust according to needs, thereby improving the accuracy of archival retrieval and the quality of use of the archival classification management system. Then, through the correlation analysis unit, the archives recommended by the result recommendation unit are compared with other archives, and relevant archives are provided for the user for reference, enabling the user to view relevant archives without repeated searches, improving the intelligence level and user experience of the archival classification management system.
[0050] 2. When storing the archival data, the hierarchical storage unit in the data storage module classifies the archival data according to factors such as access frequency and importance, dividing the archival data into hot archives and cold archives. For hot archives, they are moved to high-performance storage media to improve access speed and response time. For cold archives, they are moved to low-cost storage media to reduce storage costs, so that the storage device can be determined according to the actual situation of the archival data, reducing the cost of storage devices while ensuring the system usage efficiency and lowering the construction cost of the archival classification management system. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a block diagram of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0053] Embodiment: Please refer to Figure 1 , this embodiment provides an archival classification management system based on big data, including an archival data acquisition module, a data storage module, an archival classification module, a retrieval recommendation module, and a visualization module;
[0054] The archival data acquisition module is used to obtain archival data from the data source;
[0055] The data storage module is used to store the data collected by the archival data acquisition module. The data storage module includes a hierarchical storage unit, and the hierarchical storage unit is used to divide the archival data into cold archives and hot archives and store them using different storage media;
[0056] The file classification module uses the metadata collection tool to obtain the metadata of the file data and generates labels for each file data. The labels include time, keywords, data source, and category. The file data is classified according to the labels.
[0057] The retrieval and recommendation module includes a result recommendation unit and a relevance analysis unit. The result recommendation unit is used to provide specific files according to the content input by the user. The relevance analysis unit compares the files provided by the result recommendation unit with other files and provides relevant files for the user according to the comparison results.
[0058] More specifically, in this embodiment: The file data acquisition module is used to obtain file data from the data source. When storing through the data storage module later, the hierarchical storage unit classifies it according to factors such as the access frequency and importance of the file data, divides the file data into hot files and cold files, and uses different storage media for storage, so as to be able to determine the use of storage devices according to the actual situation of the file data, reduce the cost of storage devices while ensuring the system usage efficiency, and reduce the construction cost of the file classification management system.
[0059] Then, the file classification module uses the metadata collection tool to obtain the metadata of the file data and generates labels for each file data. The file data is classified according to the labels. When the user conducts a retrieval later, the user can directly enter the content in the search bar for searching, and can directly enter the file label information. At the same time, the result recommendation unit in the retrieval and recommendation module recommends the most suitable file for the user according to the content input by the user. The result recommendation unit includes influencing factors such as historical access frequency, timeliness, and similarity to the content input by the user, realizing the comprehensive analysis of the file data, and the user can adjust according to needs in actual operation, so as to improve the file retrieval accuracy, ensure meeting the user's needs, improve the usage quality of the file classification management system. Then, the relevance analysis unit compares the files recommended by the result recommendation unit with other files and provides relevant files for the user for reference, so that the user can view relevant files without repeated searching, improving the intelligent level and usage experience of the file classification management system.
[0060] Furthermore, the file data includes paper files and electronic files. The electronic files include audio and video data. The file data acquisition module includes an optical recognition unit and an audio conversion unit. The optical recognition unit is used to convert paper files into image format and extract the text content in the image format. The audio conversion unit is used to convert audio and video data, extract text data, and perform digital processing in combination with video transcoding technology.
[0061] Specifically, electronic archives also include PDF, Word, Excel and other files. The audio in audio and video files includes lectures, meeting minutes, interviews, etc., and the video files include meeting videos, lecture videos, surveillance videos, etc. Specifically, ASR speech recognition technology can be used to convert the speech in audio files into text, and video transcoding can be used to convert video files into a data format that can be analyzed. The audio in the video is combined for speech recognition processing, and all audio and video data are stored in text, which can facilitate subsequent archive classification management.
[0062] Furthermore, the hierarchical storage unit archive data division process is as follows:
[0063] ;
[0064] Where A is the heat score;
[0065] F is the file access frequency, which indicates the number of accesses per unit time in the past;
[0066] α is the access frequency influence coefficient, ranging from 0 to 1;
[0067] T is the time interval since the last visit;
[0068] β is the access time interval influence coefficient, ranging from 0 to 1;
[0069] Z is the archive importance weight, ranging from 0 to 1;
[0070] Specifically, the higher the heat score A, the more frequently the archive data is used, and vice versa, the less frequently the archive data is used. The classification threshold of the heat score A is set to Y1. When A is greater than Y1, it is a hot archive, and when A is less than Y1, it is a cold archive. For hot archives, they are moved to high-performance storage media to improve access speed and response time. For cold archives, they are moved to low-cost storage media to reduce storage costs, but the access speed and response time will be reduced accordingly. High-performance storage media include solid-state drives, and low-cost storage media include mechanical hard drives. In actual operations, managers can adjust the archive importance weight, access frequency influence coefficient, and access time interval influence coefficient according to the actual situation of the archive to prevent misclassification. For example, long-term important archives are classified as cold archives because they are accessed too little in the short term, thereby improving the flexibility of this archive classification management system.
[0071] Furthermore, the result recommendation unit process is as follows:
[0072] ;
[0073] Where G represents the user query content, and S represents the candidate score;
[0074] S(D 1 , G) represents the candidate score of Archive 1 according to the user's query content;
[0075] R(D 1 , G) represents the similarity score between the user's query content and Archive 1;
[0076] W1 represents the influence coefficient of R(D 1 , G), and its value range is from 0 to 1;
[0077] H(D 1 ) represents the historical access frequency of Archive 1;
[0078] W2 represents the influence coefficient of H(D 1 ), and its value range is from 0 to 1;
[0079] C(D 1 ) represents the timeliness of Archive 1;
[0080] W3 represents the influence coefficient of C(D 1 ), and its value range is from 0 to 1;
[0081] The process of obtaining the historical access frequency H(D 1 ) of Archive 1 is as follows:
[0082] ;
[0083] where U represents the set of users;
[0084] M p (D 1 ) represents the number of times user p accesses Archive 1;
[0085] M p represents the influence coefficient of user p, and its value range is from 0 to 1, representing the importance of different users;
[0086] Specifically, the larger S(D 1 , G) is, the stronger the relevance between Archive 1 and the user's query content. When the user inputs the user's query content G during the system usage process, the search result column will automatically recommend the archive with the highest score to ensure the accuracy of archive query. In actual operation, the system adds adjustment options with different influence coefficients in the search bar settings, and users can adjust W1, W2, and W3 according to their needs to ensure that the search results meet their own requirements and improve the flexibility and result accuracy of archive query.
[0087] Furthermore, the analysis process of the correlation analysis unit is as follows:
[0088] ;
[0089] where R(D1 , D 2 ) represents the similarity score between Archive One and Archive Two;
[0090] F i (D 1 ) represents the frequency of occurrence of the i-th keyword in Archive One;
[0091] F i (D 2 ) represents the frequency of occurrence of the i-th keyword in Archive Two;
[0092] n represents the number of keywords;
[0093] represents the sum of the minimum frequencies of shared features between Archive One and Archive Two, which is used to calculate how much overlap there is between the two archives for each keyword and find the minimum value for each keyword;
[0094] represents the maximum value of the total number of features in Archive One and Archive Two;
[0095] Specifically, by calculating the degree of overlap of shared features between the two archives and normalizing it with the maximum value of their total number of features, the correlation between the two archives is measured, and the similarity score R(D 1 , D 2 ) of Archive One and Archive Two is obtained. R(D 1 , D 2 ) can determine the similarity between the two archives in terms of keywords or other features, helping to find similar archives and compare the contents of the archives. After the result recommendation unit recommends the archive with the highest score to the user, the correlation analysis unit will calculate the similarity score between the archive with the highest score and other archives, and set the scoring threshold of R(D 1 , D 2 ) to Y2. When R(D 1 , D 2 ) is greater than Y2, it means that the similarity between the two archives is high, and the archive compared with the archive with the highest score will be added to the search result column. When R(D 1 , D 2 ) is less than Y2, it will not be added. In actual application, in order to avoid too many archives in the search result column, the display quantity of the search result column can be set to improve the system usage experience.
[0096] Furthermore, the visualization module provides the administrator with monitoring data on the system operation. The monitoring data includes the archive storage volume, access frequency, and system hardware operation parameters.
[0097] Specifically, during the operation of the system, the management personnel can view the system usage in real time, including the system storage status, the occupancy of the graphics card and CPU, avoid system overload, and display the data in the form of charts through the visualization module, enabling the management personnel to quickly and clearly understand the system usage and improve the system management quality.
[0098] Furthermore, the metadata collection tool is Apache Tika.
[0099] Specifically, Apache Tika is an open-source content analysis tool that can automatically detect and extract the metadata and content of files, and supports multiple file formats, including documents, PDFs, images, audio, video, etc. It can generate structured metadata for archival data, support batch processing of a large number of files, and is suitable for large-scale archival data classification scenarios.
[0100] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A file classification management system based on big data, characterized in that: It includes archive data collection module, data storage module, archive classification module, retrieval recommendation module and visualization module; The archive data acquisition module is used to obtain archive data from a data source; The data storage module is used to store the data collected by the archive data collection module, and the data storage module includes a hierarchical storage unit, and the hierarchical storage unit is used to divide the archive data into cold archives and hot archives, and use different storage media for storage; The archive classification module uses a metadata collection tool to obtain metadata of archive data and generates a label for each archive data. The label includes time, keyword, data source and category, and classifies the archive data according to the label; The search recommendation module includes a result recommendation unit and a correlation analysis unit. The result recommendation unit is used to provide a specific profile according to the content input by the user. The correlation analysis unit compares the profile provided by the result recommendation unit with other profiles and provides the user with relevant profiles according to the comparison results. The hierarchical storage unit archive data division process is as follows: ; Where A is the heat score; F is the file access frequency, which indicates the number of accesses per unit time in the past; α is the access frequency influence coefficient, ranging from 0 to 1; T is the time interval since the last visit; β is the access time interval influence coefficient, ranging from 0 to 1; Z is the archive importance weight, ranging from 0 to 1; The higher the heat score A is, the more frequently the archive data is used. Conversely, the lower the heat score A is, the less frequently the archive data is used. The classification threshold of the heat score A is set to Y1. When A is greater than Y1, it is a hot archive. When A is less than Y1, it is a cold archive. For hot archives, move them to high-performance storage media to improve access speed and response time. For cold archives, move them to low-cost storage media to reduce storage costs. The result recommendation unit process is as follows: ; Where G represents the user query content, and S represents the candidate score; S(D1,G) represents the candidate score of content profile 1 according to the user query; R(D1,G) represents the similarity score between the user query content and profile 1; W1 represents the influence coefficient of R(D1,G), ranging from 0 to 1; H(D1) represents the historical access frequency of file 1; W2 represents the influence coefficient of H(D1), ranging from 0 to 1; C(D1) indicates the timeliness of file 1; W3 represents the influence coefficient of C(D1), ranging from 0 to 1; The historical access frequency H(D1) of file 1 is obtained as follows: ; Where U represents the user set; M p (D1) represents the number of times user p accesses file 1; M p It represents the influence coefficient of user p, ranging from 0 to 1, representing the importance of different users; The larger the S(D1,G), the stronger the relevance between the file 1 and the user's query content. When the user uses the system, after entering the user query content G, the search result bar will automatically recommend a file with the highest score to ensure the accuracy of the file query. The user can adjust W1, W2 and W3 according to needs to improve the flexibility of file query; The analysis process of the correlation analysis unit is as follows: ; Where R(D1,D2) represents the similarity score between file 1 and file 2; f i (D1) represents the frequency of occurrence of the i-th keyword in file 1; f i (D2) represents the frequency of occurrence of the i-th keyword in file 2; n represents the number of keywords; It represents the sum of the minimum values of the shared feature frequencies between profile 1 and profile 2, which is used to calculate how much overlap the two profiles have on each keyword and find the minimum value for each keyword; It represents the maximum value of the total number of features in file 1 and file 2; By calculating the degree of overlap of shared features between two archives and normalizing them with the maximum value of their total number of features, the correlation between the two archives is measured to obtain a similarity score. The correlation analysis unit can determine the similarity of the two archives in terms of keywords, help find similar archives and compare archive contents. After the result recommendation unit recommends the archive with the highest candidate score S to the user, the correlation analysis unit will calculate the similarity score between the archive with the highest candidate score S and other archives, and set the score threshold of R(D1, D2) to Y2. When R(D1, D2) is greater than Y2, it means that the similarity between the two archives is high, and the archive compared with the archive with the highest score is added to the search result column. When R(D1, D2) is less than Y2, it is not added.
2. The file classification management system based on big data according to claim 1 is characterized in that: The archival data includes paper archives and electronic archives, the electronic archives include audio and video data, and the archival data acquisition module includes an optical recognition unit and an audio conversion unit. The optical recognition unit is used to convert paper archives into image format and extract text content in the image format. The audio conversion unit is used to convert audio and video data, extract text data, and perform digital processing in combination with video transcoding technology.
3. The file classification management system based on big data according to claim 1 is characterized by: The visualization module provides the administrator with monitoring data of system operation, including file storage capacity, access frequency and system hardware operation parameters.
4. The file classification management system based on big data according to claim 3 is characterized by: The high-performance storage medium includes a solid-state hard disk, and the low-cost storage medium includes a mechanical hard disk.
5. The file classification management system based on big data according to claim 1 is characterized by: The metadata collection tool is Apache Tika.
Citation Information
Patent Citations
Workshop production scheduling and analysis method based on scheduling rule
CN113570118A
Archive data storage system based on big data
CN117725283A