File management method and system based on data analysis and storage medium
By building a knowledge graph and supervising user accounts, the problems of the archive management system in data relevance and security management are solved, and efficient archive data management and security improvement are achieved.
Patent Information
- Application Number
- CN202510104283.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
The existing archive management system has problems of inefficiency and security risks in the classification storage and security management of archive data, especially when dealing with paper and electronic files, it lacks efficient correlation management and abnormal access detection.
Using a data analysis-based archive management method, a knowledge graph is built through natural language processing, machine learning and reasoning engines to realize the intelligent identification of archive data and the establishment of association relationships. At the same time, by supervising user accounts, identifying abnormal access behaviors, blocking access rights of abnormal access files and related files, and improving security.
Improve the relevance management of archive data, enhance the efficiency of archive resources, and enable managers to quickly track archives related to a certain topic. At the same time, by identifying and blocking abnormal access behaviors, the security of file storage is improved and security risks are reduced.
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of file management, and specifically provides a file management method, system and storage medium based on data analysis. Background Art
[0002] With the development and application of information technology, a large amount of file data has been accumulated in various industries. As an important asset of organizations and institutions, file data is of great significance for decision-making, management, and protection.
[0003] With the continuous update and development of modern technologies, the management of file data has gradually become digital and information-based. The current management forms of file data mainly include electronic files and paper files. Paper files will increase the workload of managers significantly. In the existing method of classifying and storing file data by computer, the degree of association between files is not high during management. Managers cannot find files related to a certain theme in a timely manner based on that theme, making the management system inefficient to use.
[0004] Most existing file management systems block the accounts that access abnormal files the most times. If there are multiple abnormal access accounts, the file management system cannot block the hidden abnormal access accounts, posing certain security risks.
[0005] Therefore, a file management method, system and storage medium based on data analysis are proposed to solve the above problems. Summary of the Invention
[0006] The purpose of the present invention is to provide a file management method, system and storage medium based on data analysis to solve the problems raised in the above background art.
[0007] To achieve the above purpose, the present invention provides the following technical solutions: A file management method, system and storage medium based on data analysis, including the following steps:
[0008] Step S1: The file management system uses a natural language processing model to process electronic files and paper files, and extracts entities from the file content;
[0009] Step S2: The file management system uses a machine learning model to intelligently identify the extracted entities, so as to obtain the association relationships between the entities;
[0010] Step S3: Based on the extracted entities and the association relationships between the entities, the file management system uses an inference engine to perform inferences to obtain the associations between the entities and the association relationships, and completes the establishment of the knowledge graph;
[0011] Step S4: After the file management system completes the establishment of the knowledge graph, it processes the electronic files and paper files to obtain unstructured data;
[0012] Step S5: The file management system processes the unstructured data and converts the unstructured data into structured data;
[0013] Step S6: The file management system stores the converted structured data, the original files of the electronic files and paper files into the knowledge graph; the file management system analyzes and processes the structured data, establishes connections between the structured data, the original files and the existing files in the knowledge graph; when the new files are completed and entered, the file management system updates the knowledge graph.
[0014] Preferably, it includes the following steps:
[0015] Step S7: The file management system monitors the user accounts and identifies abnormal access behaviors;
[0016] Step S8: The abnormally accessed files of the abnormal access behaviors, and taking the abnormally accessed files as the origin, find the directly associated and indirectly associated files with the abnormally accessed files based on the knowledge graph. After the search is completed, the file management system blocks the access permissions of the abnormally accessed files and the associated files;
[0017] Step S9: The file management system obtains the common visitors of the abnormally accessed files and the associated files, performs normalization and weighted calculation on the access times and access durations of each visitor, and combines the calculation results with the marking criteria to mark the risky visitors;
[0018] Step S10: The file management system pushes the obtained abnormally accessed files, associated files and risky visitors to the administrator, and the administrator processes according to the information pushed by the file management system.
[0019] Preferably, the method for processing the electronic files and paper files in step S4 includes:
[0020] Step S41: Convert the paper files into electronic files through a scanner;
[0021] Step S42: The file management system uses optical character recognition technology to extract the text content of the original electronic files and the electronic files converted from the paper files, and at the same time uses a vision algorithm based on machine learning to identify and process the image files to generate image data and text data;
[0022] The generated image data and text data are the unstructured data.
[0023] Preferably, the method for converting unstructured data into structured data in step S5 includes:
[0024] Data cleaning: Remove redundant information in the unstructured data, correct the incorrect data in the unstructured data, and ensure the accuracy of the content of the unstructured data;
[0025] Data encapsulation: Encapsulate the cleaned unstructured data into a standardized format;
[0026] Data integration: Integrate the encapsulated unstructured data into a data format that can be recognized by the knowledge graph.
[0027] Preferably, the analysis and processing of the structured data by the file management system in step S6 includes extracting the metadata of the structured data, and the file management system matches and establishes a relationship with the existing files in the knowledge graph through the metadata of the structured data.
[0028] Preferably, the supervision content of the user account by the file management system in step S7 includes the user name, user permission level, accessed file, access time of the file, and access times of the file for each access.
[0029] Preferably, the evaluation criteria for abnormal access behavior by the file management system in step S7 include:
[0030] Step S71: Normalize the recorded access information to obtain an access status evaluation value;
[0031] Step S72: Set an evaluation threshold A. If the access status evaluation value of the file is greater than the set evaluation threshold, the access situation of the file is considered an abnormal access behavior.
[0032] Preferably, the marking criteria in step S9 include:
[0033] Step S91: Obtain a risk evaluation value through normalization and weighted calculation, and mark the one with the largest value as a risk visitor;
[0034] Step S92: Set a risk threshold B. If the difference between the risk evaluation values of other visitors and the maximum value of the risk evaluation value exceeds the risk threshold B, they are also listed as risk visitors.
[0035] Compared with the prior art, the beneficial effects of the present invention are:
[0036] 1. The present invention constructs a knowledge graph through electronic files and paper files, and processes the electronic files and paper files stored in the knowledge graph into structured data. The file management system analyzes the processed structured data and establishes a connection with the established knowledge graph to increase the correlation between files, make full use of file resources, and enable the stored files to form a complete system. Managers can quickly track files related to a certain theme through the knowledge graph.
[0037] 2. Based on the recorded user access situations, the present invention identifies abnormal access behaviors among them through abnormal access criteria evaluation, and finds abnormal access files through the abnormal access behaviors. The file management system takes the abnormal access files as the origin, and at the same time, based on the knowledge graph, finds the associated files of the abnormal access files. The file management system blocks the abnormal access files and the associated files. The file management system then finds risk visitors through the abnormal access files and the associated files, and pushes the obtained abnormal access files, associated files, and risk visitors to the administrator. The administrator monitors the stored file information through the information pushed by the file management system, increasing the security of file storage. Specific implementation manner
[0038] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0039] A file management method, system and storage medium based on data analysis, including the following steps:
[0040] Step S1: The file management system uses a natural language processing model to process electronic files and paper files, and extracts entities from the file content;
[0041] Step S2: The file management system uses a machine learning model to intelligently identify the extracted entities, so as to obtain the association relationships between the entities;
[0042] Step S3: Based on the extracted entities and the association relationships between the entities, the file management system uses an inference engine for inference to obtain the association between the entities and the association relationships, and completes the establishment of the knowledge graph;
[0043] Step S4: After the file management system completes the establishment of the knowledge graph, it processes the electronic files and paper files to obtain unstructured data;
[0044] Unstructured data specifically refers to content without a predefined data model or rules, which is difficult to be directly used for computer processing, such as document content, pictures, text described in natural language, etc.
[0045] Step S5: The archive management system processes the unstructured data and converts the unstructured data into structured data.
[0046] Structured data specifically refers to data with a fixed format and good organization.
[0047] Step S6: The archive management system stores the converted structured data, the original files of electronic archives and paper archives into the knowledge graph; the archive management system analyzes and processes the structured data, establishes connections between the structured data, the original files and the existing archives in the knowledge graph; when a new archive is completed for entry, the archive management system updates the knowledge graph.
[0048] Specifically, it includes the following steps:
[0049] Step S7: The archive management system monitors the user accounts and identifies abnormal access behaviors.
[0050] Each user should have a corresponding access account when accessing, and each account has its corresponding access permissions. The range of files that can be accessed by accounts with different access permissions is also limited.
[0051] Abnormal access specifically refers to archive access behaviors that exceed the expected range or violate the permission management policy. Possible types of abnormal access include abnormal access frequency, abnormal access time, abnormal access habits, account sharing possibilities, etc.
[0052] Step S8: The abnormal access archives accessed by the abnormal access behaviors, and taking the abnormal access archives as the origin, find the directly associated and indirectly associated associated archives based on the knowledge graph. After the search is completed, the archive management system blocks the access permissions of the abnormal access archives and the associated archives.
[0053] Step S9: The archive management system obtains the common visitors of the abnormal access archives and the associated archives, performs normalization and weighted calculation on the access times and access durations of each visitor, and combines the calculation results with the marking criteria to mark the risky visitors.
[0054] Step S10: The archive management system pushes the obtained abnormal access archives, associated archives and risky visitors to the administrator, and the administrator processes according to the information pushed by the archive management system.
[0055] Files determined to have abnormal access are marked as primary anomalies, and the relevant data is regarded as secondary anomalies. The access permissions for the relevant files in the access chain will be restricted, and the anomalies will be reported to the administrator. To lift the access to the files on this chain, the administrator needs to manually do so.
[0056] The administrator will restrict the access permissions of the marked risky visitors to the files. Only after the administrator's investigation will the access restrictions on the risky visitors be lifted.
[0057] Specifically, the method for processing electronic files and paper files in step S4 includes:
[0058] Step S41: Convert paper files into electronic files through a scanner;
[0059] Step S42: Use optical character recognition technology to extract the text content of the original electronic files and the electronic files converted from paper files. At the same time, use a vision algorithm based on machine learning to identify and process the image files to generate image data and text data;
[0060] The generated image data and text data are unstructured data.
[0061] Specifically, the processing method for converting unstructured data into structured data in step S5 includes:
[0062] Data cleaning: Remove redundant information in the unstructured data, correct the incorrect data in the unstructured data, and ensure the accuracy of the content of the unstructured data;
[0063] Data encapsulation: Encapsulate the cleaned unstructured data into a standardized format;
[0064] Data integration: Integrate the encapsulated unstructured data into a data format that can be recognized by the knowledge graph.
[0065] Specifically, the analysis and processing of structured data by the file management system in step S6 includes extracting the metadata of the structured data. The file management system matches the metadata of the structured data with the existing files in the knowledge graph and establishes relationships.
[0066] The metadata of the structured data specifically includes basic information such as file title, author, date, file type, project number, etc.
[0067] Specifically, the supervision content of the file management system for user accounts in step S7 includes the user name, user permission level, accessed file, access time of the file, and number of accesses to the file for each access.
[0068] Specifically, the evaluation criteria for abnormal access behaviors by the file management system in step S7 include:
[0069] Step S71: Normalize the recorded access information to obtain an access status evaluation value;
[0070] Step S72: Set an evaluation threshold A. If the access status evaluation value of a file is greater than the set evaluation threshold, the access situation of the file is considered an abnormal access behavior.
[0071] Specifically, the marking criteria in Step S9 include:
[0072] Step S91: Obtain a risk assessment value through normalization and weighted calculation, and mark the one with the largest value as a risk visitor;
[0073] Step S92: Set a risk threshold B. If the difference between the risk assessment value of other visitors and the maximum value of the risk assessment value exceeds the risk threshold B, they are also listed as risk visitors.
[0074] Example 1:
[0075] The administrator processes the electronic files and paper files to be stored in the database. First, the administrator uses a scanner to convert the paper files into electronic files. After the administrator completes the conversion of the paper files, the original electronic files and the electronic files converted from the paper files are input into the file management system. The file management system uses optical character recognition technology to extract the text content from the original electronic files and the electronic files converted from the paper files, and at the same time uses a vision algorithm based on machine learning to identify and process the image files to generate image data and text data.
[0076] At this time, the image data and text data generated in the file management system are unstructured data.
[0077] Next, the file management system processes the unstructured data, converts the unstructured data into structured data, and integrates the unstructured data into a data format that can be recognized by the knowledge graph.
[0078] After the file management system completes the processing of the unstructured data, it stores the completed converted structured data, the original files of the electronic files and paper files into the established knowledge graph; the file management system extracts the metadata of the structured data, and establishes a connection between the structured data, the original files and the files already existing in the knowledge graph; when a new file is completed for entry, the file management system updates the knowledge graph to continuously expand the knowledge graph.
[0079] Example 2:
[0080] The file management system monitors the usage of user accounts and records the actual access situation of user accounts in the form of logs.
[0081] When a user uses an account to perform a risk browsing behavior on a file, the file management system uses an abnormal access standard assessment to identify the actual access behavior recorded in the log and marks the abnormal access behavior among them.
[0082] Based on the marked abnormal access behavior, the file management system searches for the abnormal access files accessed by the abnormal access behavior. After the file management system finds the abnormal access files, taking the abnormal access files as the origin and based on the knowledge graph, it searches for the associated files that are directly and indirectly associated with the abnormal access files. After the search is completed, the file management system blocks the access permissions of the abnormal access files and the associated files.
[0083] After the file management system completes blocking the access permissions of the abnormal access files and the associated files, it obtains the common visitors of the abnormal access files and the associated files. The file management system then calculates based on the number of accesses and access duration of each visitor, and combines the calculation results with the marking standard to mark the risk visitors. The file management system restricts the functions of the accounts of the marked risk visitors.
[0084] The file management system pushes the obtained abnormal access files, associated files, and risk visitors to the administrator, and the administrator processes them according to the information pushed by the file management system. The administrator determines whether the file management system has made a mis-block. If it is a mis-block, the administrator will lift the function restriction on the mis-blocked account and the access permission restriction on the abnormal access files and the associated files.
[0085] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to embrace all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention.
Claims
1. A file management method based on data analysis, characterized in that: The steps include: Step S1: The archive management system uses a natural language processing model to process electronic archives and paper archives and extract entities from the archive content; Step S2: The archive management system uses a machine learning model to intelligently identify the extracted entities, thereby obtaining the association relationship between the entities; Step S3: The archive management system uses the inference engine to perform reasoning based on the extracted entities and the association relationships between the entities, obtains the association between the entities and the association relationships, and completes the establishment of the knowledge graph; Step S4: After the archive management system completes the establishment of the knowledge graph, the electronic archives and paper archives are processed to obtain unstructured data; Step S5: The archive management system processes the unstructured data and converts the unstructured data into structured data; Step S6: The archive management system stores the converted structured data and the original files of the electronic archives and paper archives into the knowledge graph; the archive management system analyzes and processes the structured data, and establishes a connection between the structured data, the original files and the existing archives in the knowledge graph; When the new file is entered, the file management system updates the knowledge graph.
2. The file management method based on data analysis according to claim 1, characterized in that: The steps include: Step S7: The archive management system monitors the user account and identifies abnormal access behavior; Step S8: The abnormal access files accessed by the abnormal access behavior are used as the origin, and the associated files directly and indirectly associated with the abnormal access files are found based on the knowledge graph. After the search is completed, the file management system blocks the access rights of the abnormal access files and the associated files; Step S9: The archive management system obtains the abnormal access archives and the common visitors of the associated archives, normalizes and weights the number of visits and the duration of visits of each visitor, and combines the calculation results with the marking standard to mark the risky visitors; Step S10: The archive management system pushes the obtained abnormal access archives, related archives and risky visitors to the administrator, and the administrator processes according to the information pushed by the archive management system.
3. The archive management method based on data analysis according to claim 1, characterized in that: The method for processing electronic archives and paper archives in step S4 includes: Step S41: Converting paper files into electronic files through a scanner; Step S42: The archive management system uses optical character recognition technology to extract text content from the original electronic archives and the electronic archives converted from paper archives, and uses a machine learning-based visual algorithm to recognize and process image files to generate image data and text data; The generated image data and text data are unstructured data.
4. The archive management method based on data analysis according to claim 1, characterized in that: The processing method for converting unstructured data into structured data in step S5 includes: Data cleaning: remove redundant information from unstructured data, correct erroneous data in unstructured data, and ensure the accuracy of the content of unstructured data; Data packaging: packaging cleaned unstructured data into a standardized format; Data integration: Integrate the encapsulated unstructured data into a data format that can be recognized by the knowledge graph.
5. The file management method based on data analysis according to claim 1, characterized in that: In step S6, the archive management system analyzes and processes the structured data, including extracting metadata of the structured data. The archive management system matches and establishes relationships with existing archives in the knowledge graph through the metadata of the structured data.
6. The archive management method, system and storage medium based on data analysis according to claim 1, characterized in that: The monitoring content of the user account by the archive management system in step S7 includes the user name of each access, the user authority level, the archive accessed, the time when the archive was accessed, and the number of times the archive was accessed.
7. The archive management method based on data analysis according to claim 1, characterized in that: The evaluation criteria of the archive management system for abnormal access behavior in step S7 include: Step S71: normalizing the recorded access information to obtain an access status evaluation value; Step S72: setting an evaluation threshold A. If the access status evaluation value of the file is greater than the set evaluation threshold, the access status of the file is considered to be abnormal access behavior.
8. The archive management method based on data analysis according to claim 1, characterized in that: The marking standard in step S9 includes: Step S91: obtaining a risk assessment value through normalization and weighted calculation, wherein the one with the largest value is marked as a risk visitor; Step S92: A risk threshold B is set. If the difference between the risk assessment value of other visitors and the maximum risk assessment value exceeds the risk threshold B, they are also listed as risky visitors.
9. A file management system based on data analysis, comprising: one or more processors; and a storage device for storing one or more programs, wherein the one or more processors implement the method according to any one of claims 1 to 8.
10. An archive management storage medium based on data analysis, on which executable instructions are stored, and when the instructions are executed by a processor, the processor implements the method according to any one of claims 1 to 8.
Citation Information
Cited By
Statistical analysis method and system for archive information
CN121255886A
Cultural industry digital monitoring system and method based on machine learning
CN121543114A