A computer data analysis and management system based on big data

By designing a computer data analysis and management system based on big data, using data analysis and user behavior data to calculate user attention, intelligent classification and storage of files is realized, and the problem of difficulty in identifying and cleaning temporary files is solved by existing tools, and the efficiency and accuracy of file management are improved.

CN119166595BActive Publication Date: 2025-06-06SHANGHAI HENGGE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311518472.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2025-06-06
Estimated Expiration
2043-11-15

AI Technical Summary

Technical Problem

Existing junk file cleaning tools are difficult to adapt to new junk file types or changed usage patterns, and tools based on rules alone cannot identify and delete temporary files that are not of use value in a timely manner, affecting user cleaning efficiency.

Method used

A computer data analysis and management system based on big data is designed, which stores file information and user behavior data through the database. The data collection module extracts file size and deletes records. The data analysis module calculates feature vectors and user attention. The real-time file classification storage module classifies and stores it according to user attention, and provides users with cleaning suggestions through the file cleaning suggestions module.

Benefits of technology

It realizes more comprehensive, accurate and personalized file management, can intelligently classify and store files, accurately identify and clean up junk files, provide dynamically adjusted classification rules, and improve the timeliness and accuracy of file management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119166595B_ABST
    Figure CN119166595B_ABST
Patent Text Reader

Abstract

The present invention discloses a computer data analysis management system based on big data, belonging to the field of data analysis technology. The system of the present invention comprises a database, a data acquisition module, a data analysis module, a real-time file classification storage module, a file cleaning suggestion module and a user management module; the database stores all file information and user behavior data; the data acquisition module extracts the file saving information and deletion records from the database, and performs corresponding processing; the data analysis module analyzes the file information and deletion records, extracts the target file, and calculates the feature vector and user attention; the real-time file classification storage module classifies and stores the real-time files according to the user attention, and dynamically calculates the user attention; the file cleaning suggestion module analyzes the file list and suggests cleaning operations; the user management module receives notification information, performs file management and feeds back the results to the file cleaning suggestion module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and in particular to a computer data analysis management system based on big data. Background Art

[0002] Computer systems have become an important part of modern life and work. People use computers to store, manage and process various types of data and files. However, as time goes by, the amount of data in computer systems gradually increases, which leads to some problems, including the accumulation of junk files. Junk files usually include temporary files that are no longer used, invalid application residual files, duplicate files, etc., which occupy valuable storage space, reduce computer performance, and increase the complexity of data management.

[0003] Existing junk file cleaning tools are usually based on rules and pattern matching. They rely on fixed rules to identify junk files. This approach has some limitations: because the rules are fixed, these tools are difficult to adapt to new junk file types or changing usage patterns; tools based only on rules tend to only identify obvious junk files, such as invalid application residual files and duplicate files, etc., and cannot delete temporary files that have no use value in a timely manner, affecting the efficiency of users in cleaning junk files. Summary of the invention

[0004] The purpose of the present invention is to provide a computer data analysis and management system based on big data to solve the problems raised in the above background technology.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0006] A computer data analysis and management system based on big data, the system includes a database, a data acquisition module, a data analysis module, a real-time file classification storage module, a file cleaning suggestion module and a user management module;

[0007] The database stores all file information and all user behavior data of the files. The file information includes save information and delete records. The save information records the basic properties of the file, such as the file name, size, format, and creation time, while the delete record records the deletion operation of the file, including the deletion time and deletion operator. User behavior data includes various operations of users on files, including users' browsing history, modification history, and sharing history.

[0008] The data acquisition module obtains the file saving information and deletion records from the database, extracts the file size and performs sorting analysis;

[0009] The data analysis module analyzes the file saving information and deletion records, extracts the target files, and calculates the feature vector and user attention based on the user behavior data; through data analysis, the threshold range of user attention is obtained;

[0010] The real-time file classification storage module classifies and stores the files received in real time according to the user attention, and recalculates the user attention within a certain time interval for dynamic adjustment;

[0011] The file cleaning suggestion module analyzes and judges the file lists in the recycle bin, temporary storage space and permanent storage space, and sends corresponding notification information to the user management module to suggest file cleaning operations;

[0012] The user management module receives notification information from the file cleaning suggestion module, the user performs corresponding file management on the received notification information, and sends feedback information to the file cleaning suggestion module to inform the user of the selection and execution result.

[0013] Further, the data collection module includes a file saving information extraction unit, a file deletion record extraction unit, a file size extraction unit and an empty file identification unit;

[0014] The file saving information extraction unit and the file deletion record extraction unit obtain the saving information and deletion records of all files in the database, and construct a file saving information set A according to the timestamp sequence of file generation in the saving information, and A={ai}, where i is a positive integer from 1 to n, and ai represents the i-th file saving information; construct a file deletion record set B according to the timestamp sequence of deleted files in the deletion record, and B={bj}, where j is a positive integer from 1 to m, and bj represents the j-th deletion record;

[0015] The file size extraction unit extracts the sizes of all files in the file storage information set A, records them in the format of "file name-file size", and forms a file size data table; arranges the data in the file size data table in descending order, obtains the file name with a file size of 0 bytes, extracts the file name of the file deletion record set B and searches for it; the empty file identification unit marks the files in set A that do not correspond to the file names in set B, and marks them as empty files.

[0016] Since set A records all file storage information, the files in set A include junk files, temporary files and long-lived files, where temporary files refer to files that are no longer of use value; the purpose of obtaining the size of the files in set A is to eliminate the impact of empty files on subsequent analysis; set B records file deletion information, so there are some junk information deletion records in set B, including the deletion of empty files; the purpose of searching for deleted file names in set B is to find some undeleted empty files in set A, so it is necessary to use an empty file identification unit to mark the undeleted empty files in set A.

[0017] Further, the data analysis module includes a file name comparison unit, a target file feature extraction unit, a feature vector correlation calculation unit, and a user attention calculation unit;

[0018] The file name comparison unit extracts the file names of all files from the file saving information set A and the file deletion record set B, and compares the file names of the set A with those of the set B;

[0019] Obtain the file names that belong only to set A, filter out empty files, and mark the remaining file names that belong only to set A as target files 1; the target file feature extraction unit obtains the user behavior data of the target file 1, obtains the user browsing frequency Lw, user modification frequency Xw, and user sharing frequency Fw of each target file 1 in the selected time period T, and forms a feature vector, which is represented by V1k=(Lwk, Xwk, Fwk), and V1k represents the feature vector of the kth target file 1, where k represents the number of target files 1 in the selected time period, which is a positive integer;

[0020] Get the file name that belongs to both set A and set B, and mark it as target file 2; get the save information and deletion record of the file name corresponding to target file 2, extract the timestamp TA generated by each target file 2 and the timestamp TB deleted by each target file 2, and calculate the time difference ΔT between TA and TB, ΔT = TB-TA;

[0021] For each target file 2, the user browsing frequency Lw, user modification frequency Xw, and user sharing frequency Fw within the ΔT time period are obtained and formed into a feature vector, which is represented by V2 d =(Lw d ,Xw d ,Fw d ), and V2 d The feature vector representing the dth target file 2, where d represents the number of target files 2 in the ΔT time period, and is a positive integer;

[0022] For each feature vector of target file 2, a feature vector correlation calculation unit is used to perform vector linear correlation operations on the feature vectors in target file 1 to obtain the file names of the files corresponding to the feature vectors in target file 2 that are linearly correlated with target file 1, forming a set C.

[0023] The feature vector V2 in target file 2 that is linearly related to target file 1 d , indicating that such feature vectors can be represented by the feature vector of target file 1, representing feature vector V2 d Correlation with the feature vector of target file 1; if the feature vector of target file 2 is linearly correlated with the feature vector of target file 1, it indicates that the corresponding target file 2 is a temporary file, and temporary files mean that such files are useful to users, but have time limits, which means that such files are not always of use value to users. When they are no longer of use value, such files need to be deleted; by analyzing the linear correlation between the feature vector of target file 2 and the feature vector in target file 1, temporary files can be found.

[0024] Furthermore, the user attention calculation unit includes:

[0025] Obtain the user behavior data of each target file 2, and record the number of views Ln and the browsing duration Lt, the number of modifications Xn, and the number of shares Fn within the ΔT time period for each target file 2; calculate the user attention Y of each target file 2, and the specific calculation formula is:

[0026] Y = (ΔT / T)*[(Ln / SLn)*(Lt / SLt)+(Xn / SXn)+(Fn / SFn)], where T represents the time period from the generation time of the first target file 2 to the deletion time of the nth target file 2, SLn represents the sum of the number of views of all target files 2 in the time period T, SLt represents the sum of the viewing duration of all target files 2 in the time period T, SXn represents the sum of the number of modifications of all target files 2 in the time period T, and SFn represents the sum of the number of shares of all target files 2 in the time period T;

[0027] The calculation of user attention Y is based on the product of ΔT / T and (Ln / SLn)*(Lt / SLt)+(Xn / SXn)+(Fn / SFn), where ΔT / T represents the ratio of the time difference ΔT between the timestamp TA of target file 2 generation and the timestamp TB of target file 2 deletion to the time period T of all target file 2 generation and deletion. For the difference ΔT, the smaller ΔT is, the earlier the target file 2 is deleted. No matter how large the value of (Ln / SLn)*(Lt / SLt)+(Xn / SXn)+(Fn / SFn), the smaller ΔT is, it indicates that the target file 2 is no longer of use value. Therefore, the user attention is calculated from two aspects: ΔT and user behavior data, which can make the user attention more comprehensively reflect the importance of the file.

[0028] Get the user attention of target file 2 corresponding to all file names in set C, and get the user attention threshold interval Q, and Q = [q_min, q_max], where q_min is the minimum user attention of target file 2 in set C, and q_max is the maximum user attention of target file 2 in set C.

[0029] Further, the real-time file classification storage module includes a real-time file user attention calculation unit, a user attention threshold comparison unit and a file classification processing unit;

[0030] The real-time file user attention calculation unit includes:

[0031] The user attention degree of the files received in real time is calculated every selected time period T0, and the calculation formula of the user attention degree Ys of the real-time files is:

[0032] Ys=(ΔT0 / T0)*[(Lsn / SLsn)*(Lst / SLst)+(Xsn / SXsn)+(Fsn / SFsn)],

[0033] Wherein, ΔT0 represents the difference between the generation time of each real-time received file and the current time, Lsn represents the number of times each real-time received file is browsed in the T0 time period, SLsn represents the sum of the number of times all real-time received files are browsed in the T0 time period, Xsn represents the number of times each real-time received file is modified in the T0 time period, SXsn represents the sum of the number of times all real-time received files are modified in the T0 time period, Fsn represents the number of times each real-time received file is shared in the T0 time period, and SFsn represents the sum of the number of times all real-time received files are shared in the T0 time period.

[0034] Furthermore, the user attention threshold comparison unit includes:

[0035] The obtained real-time file user attention Ys is compared with the user attention threshold interval Q. If Ys < q_min, it means that the user attention of the real-time file is low, which also reflects that the use value of this file to the user is low. In this case, the real-time received file is classified as a junk file and stored in the recycle bin. The recycle bin is set with an automatic cleaning time T1, and the recycle bin will be automatically cleaned every T1 period.

[0036] If q_min≤Ys≤q_max, it means that the user attention of the real-time file meets the attention threshold range, which also reflects that the file has a certain use value to the user. In this case, the real-time received file is classified as a temporary file and stored in the temporary storage space. The user attention Ys' is recalculated every T2 time period. When Ys'<q_min, the file is classified as a junk file and marked as "temporary file converted to junk file", and the storage address is also moved from the temporary storage space to the recycle bin; when Ys'>q_max, the file is classified as a permanent file and marked as "temporary file converted to permanent file", and the storage address is also moved from the temporary storage space to the permanent storage space;

[0037] If Ys>q_max, the file received in real time will be classified as a permanent file and stored in the permanent storage space, and the user attention Ys' will be recalculated every T3 time period. When q_min≤Ys'≤q_max, the file will be classified as a temporary file and marked as "permanent file converted to temporary file", and the storage address will be moved from the permanent storage space to the temporary space; when Ys'<q_min, the file will be classified as a junk file and marked as "permanent file converted to junk file", and the storage address will be moved from the permanent storage space to the recycle bin.

[0038] Furthermore, the file classification processing unit includes:

[0039] According to the comparison result of the user attention threshold comparison unit, the file is classified; if the user attention of the real-time file is lower than the minimum value of the threshold interval Q, it is classified as a junk file and stored in the recycle bin, and the recycle bin will clean up the junk file according to the set automatic cleaning time period; if the user attention of the real-time file is between the minimum value and the maximum value of the threshold interval Q, including the minimum value and the maximum value, it is stored as a temporary file in the temporary storage space, the user attention is recalculated every T2 time period, the classification type of the file is determined again according to the new user attention, and the labeling and storage position are adjusted; if the user attention of the real-time file is higher than the maximum value of the threshold interval Q, it is classified as a permanent file and stored in the permanent storage space, the user attention is recalculated every T3 time period, the classification type of the file is determined again according to the new user attention, and the labeling and storage position are adjusted.

[0040] The user attention of temporary files is calculated every T2 time period, which provides a dynamic adjustment function to prevent the user attention of temporary files from changing later, causing the temporary files to become junk files and occupy space; similarly, the user attention of long-term files is calculated every T3 time period, and T1>T2, T1>T3.

[0041] Further, the file cleanup suggestion module includes a recycle bin file processing unit, a temporary storage space file processing unit, and a permanent storage space file processing unit;

[0042] The recycle bin file processing unit obtains the file list R in the recycle bin and determines whether the automatic cleaning time T1 has been reached; if the automatic cleaning time T1 has been reached, the files are included in the cleaning suggestion list, compressed and stored in the cloud space, and completely deleted from the recycle bin; if the automatic cleaning time T1 has not been reached, no processing is performed;

[0043] The role of cloud space is to prevent users from accidentally deleting files. When users need to retrieve deleted files, they can log in to the cloud space to search. Cloud space only serves as a storage function, and the management of cloud space requires manual management by users.

[0044] The temporary storage space file processing unit obtains the file list S in the temporary storage space, associates the user attention Ys' recalculated every T2 time period with the file list S; sends the notification information marked "temporary files are converted to junk files" to the user management module, and the user determines whether to delete them immediately. If the user gives an immediate deletion instruction, they are deleted immediately; if the user does not give a deletion instruction, the same subsequent operation is performed as for the files in the recycle bin; sends the notification information marked "temporary files are converted to permanent files" to the user management module, and the user confirms the final storage location of the file;

[0045] The long-term storage space file processing unit obtains the file list P in the long-term storage space, and associates the user attention Ys' recalculated every T3 time period with the file list P; sends the notification information marked "long-term storage files are converted to temporary files" to the user management module, and the user confirms the final storage location of the file; sends the notification information marked "long-term storage files are converted to junk files" to the user management module, and the user determines whether to delete them immediately. If the user gives an immediate deletion instruction, they are deleted immediately; if the user does not give a deletion instruction, the same subsequent operations are performed as for files in the recycle bin.

[0046] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: the present invention provides a more comprehensive, accurate and personalized computer data analysis system, which can not only store and manage the basic attributes of files and user behavior data, but also intelligently classify and store file data through big data analysis, thereby accurately cleaning up junk files, and also provides cloud space storage for deleted junk files, which provides a simple and fast method for users to restore files; the system obtains the file saving information and deletion records in real time through the data acquisition module, and uses the data analysis module to intelligently classify and analyze the attention of files, thereby realizing the automatic management and processing of files; the system classifies files according to the frequency of users' browsing, modification and sharing behaviors, and The system automatically classifies files into junk files, temporary files and long-term files, and stores them in the corresponding space, providing a specific and dynamically adjusted classification rule; the system calculates the user's attention to the file by calculating indicators such as the number of times the user browses the file, the browsing time, the number of modifications and the number of shares, combined with the overall trend in the time period, to determine the importance and storage location of the file; the system calculates the user's attention when receiving files in real time, and classifies and stores the files according to the threshold range, ensuring the timeliness and accuracy of file management; the system judges and processes the files in the recycle bin, temporary space and long-term space through the file cleanup suggestion module, sends the cleanup suggestion list to the user, and receives the user's feedback information, providing a convenient file management method. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0048] Figure 1 It is a module schematic diagram of a computer data analysis and management system based on big data of the present invention. DETAILED DESCRIPTION

[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0050] See also Figure 1 , the present invention provides a technical solution:

[0051] A computer data analysis and management system based on big data, the system includes a database, a data acquisition module, a data analysis module, a real-time file classification storage module, a file cleaning suggestion module and a user management module;

[0052] The database stores all file information and all user behavior data of the files. The file information includes save information and delete records. The save information records the basic properties of the file, such as the file name, size, format, and creation time, while the delete record records the deletion operation of the file, including the deletion time and deletion operator. User behavior data includes various operations of users on files, including users' browsing history, modification history, and sharing history.

[0053] The data acquisition module obtains the file saving information and deletion records from the database, extracts the file size and performs sorting analysis;

[0054] The data analysis module analyzes the file saving information and deletion records, extracts the target files, and calculates the feature vector and user attention based on the user behavior data; through data analysis, the threshold range of user attention is obtained;

[0055] The real-time file classification storage module classifies and stores the files received in real time according to the user attention, and recalculates the user attention within a certain time interval for dynamic adjustment;

[0056] The file cleaning suggestion module analyzes and judges the file lists in the recycle bin, temporary storage space and permanent storage space, and sends corresponding notification information to the user management module to suggest file cleaning operations;

[0057] The user management module receives notification information from the file cleaning suggestion module, the user performs corresponding file management on the received notification information, and sends feedback information to the file cleaning suggestion module to inform the user of the selection and execution result.

[0058] The data collection module includes a file saving information extraction unit, a file deletion record extraction unit, a file size extraction unit and an empty file identification unit;

[0059] The file saving information extraction unit and the file deletion record extraction unit obtain the saving information and deletion records of all files in the database, and construct a file saving information set A according to the timestamp sequence of file generation in the saving information, and A={ai}, where i is a positive integer from 1 to n, and ai represents the i-th file saving information; construct a file deletion record set B according to the timestamp sequence of deleted files in the deletion record, and B={bj}, where j is a positive integer from 1 to m, and bj represents the j-th deletion record;

[0060] The file size extraction unit extracts the sizes of all files in the file storage information set A, records them in the format of "file name-file size", and forms a file size data table; arranges the data in the file size data table in descending order, obtains the file name with a file size of 0 bytes, extracts the file name of the file deletion record set B and searches for it; the empty file identification unit marks the files in set A that do not correspond to the file names in set B, and marks them as empty files.

[0061] Since set A records all file storage information, the files in set A include junk files, temporary files and long-lived files, where temporary files refer to files that are no longer of use value; the purpose of obtaining the size of the files in set A is to eliminate the impact of empty files on subsequent analysis; set B records file deletion information, so there are some junk information deletion records in set B, including the deletion of empty files; the purpose of searching for deleted file names in set B is to find some undeleted empty files in set A, so it is necessary to use an empty file identification unit to mark the undeleted empty files in set A.

[0062] Data analysis module file name comparison unit, target file feature extraction unit, feature vector correlation calculation unit and user attention calculation unit;

[0063] The file name comparison unit extracts the file names of all files from the file saving information set A and the file deletion record set B, and compares the file names of the set A with those of the set B;

[0064] Obtain the file names that belong only to set A, filter out empty files, and mark the remaining file names that belong only to set A as target files 1; the target file feature extraction unit obtains the user behavior data of the target file 1, obtains the user browsing frequency Lw, user modification frequency Xw, and user sharing frequency Fw of each target file 1 in the selected time period T, and forms a feature vector, which is represented by V1k=(Lwk, Xwk, Fwk), and V1k represents the feature vector of the kth target file 1, where k represents the number of target files 1 in the selected time period, which is a positive integer;

[0065] Get the file name that belongs to both set A and set B, and mark it as target file 2; get the save information and deletion record of the file name corresponding to target file 2, extract the timestamp TA generated by each target file 2 and the timestamp TB deleted by each target file 2, and calculate the time difference ΔT between TA and TB, ΔT = TB-TA;

[0066] For each target file 2, the user browsing frequency Lw, user modification frequency Xw, and user sharing frequency Fw within the ΔT time period are obtained and formed into a feature vector, which is represented by V2d =(Lw d ,Xw d ,Fw d ), and V2 d The feature vector representing the dth target file 2, where d represents the number of target files 2 in the ΔT time period, and is a positive integer;

[0067] For each feature vector of target file 2, a feature vector correlation calculation unit is used to perform vector linear correlation operations on the feature vectors in target file 1 to obtain the file names of the files corresponding to the feature vectors in target file 2 that are linearly correlated with target file 1, forming a set C.

[0068] The feature vector V2 in target file 2 that is linearly related to target file 1 d , indicating that such feature vectors can be represented by the feature vector of target file 1, representing feature vector V2 d Correlation with the feature vector of target file 1; if the feature vector of target file 2 is linearly correlated with the feature vector of target file 1, it indicates that the corresponding target file 2 is a temporary file, and temporary files mean that such files are useful to users, but have time limits, which means that such files are not always of use value to users. When they are no longer of use value, such files need to be deleted; by analyzing the linear correlation between the feature vector of target file 2 and the feature vector in target file 1, temporary files can be found.

[0069] In this embodiment,

[0070] Assume that the matrix composed of the feature vectors of target file 1 is A1=[V1 k1 ,V1 k2 ,...,V1 kW ], and V1 k1 =(Lw k1 ,Xw k1 ,Fw k1 );

[0071] The matrix composed of the feature vectors of target file 2 is B1=[V2 d1 ,V2 d2 ,...,V2 dv ], and V2 d1 =(Lw d1 ,Xw d1 ,Fw d1 );

[0072] Solve A1X1=V2 respectively d1 ,...,AvXv=V2 dvBy solving these equations that are not all zero, we can get the linear relationship between target file 1 and target file 2, so as to find the temporary file in target file 2, that is, the file corresponding to the file name in set C.

[0073] The user attention calculation unit includes:

[0074] Obtain the user behavior data of each target file 2, and record the number of views Ln and the browsing duration Lt, the number of modifications Xn, and the number of shares Fn within the ΔT time period for each target file 2; calculate the user attention Y of each target file 2, and the specific calculation formula is:

[0075] Y = (ΔT / T)*[(Ln / SLn)*(Lt / SLt)+(Xn / SXn)+(Fn / SFn)], where T represents the time period from the generation time of the first target file 2 to the deletion time of the nth target file 2, SLn represents the sum of the number of views of all target files 2 in the time period T, SLt represents the sum of the viewing duration of all target files 2 in the time period T, SXn represents the sum of the number of modifications of all target files 2 in the time period T, and SFn represents the sum of the number of shares of all target files 2 in the time period T;

[0076] The calculation of user attention Y is based on the product of ΔT / T and (Ln / SLn)*(Lt / SLt)+(Xn / SXn)+(Fn / SFn), where ΔT / T represents the ratio of the time difference ΔT between the timestamp TA of target file 2 generation and the timestamp TB of target file 2 deletion to the time period T of all target file 2 generation and deletion. For the difference ΔT, the smaller ΔT is, the earlier the target file 2 is deleted. No matter how large the value of (Ln / SLn)*(Lt / SLt)+(Xn / SXn)+(Fn / SFn), the smaller ΔT is, it indicates that the target file 2 is no longer of use value. Therefore, the user attention is calculated from two aspects: ΔT and user behavior data, which can make the user attention more comprehensively reflect the importance of the file.

[0077] Get the user attention of target file 2 corresponding to all file names in set C, and get the user attention threshold interval Q, and Q = [q_min, q_max], where q_min is the minimum user attention of target file 2 in set C, and q_max is the maximum user attention of target file 2 in set C.

[0078] In this embodiment, it is assumed that there are three target files 2 in set C, namely File1, File2, and File3;

[0079] For File1, user behavior data in time period T (30 minutes): ΔT1 = 10 minutes; number of views Ln1 = 10; browsing time Lt1 = 500 seconds; number of modifications Xn1 = 1; number of shares Fn = 0;

[0080] For File2, user behavior data within time period T (30 minutes):

[0081] ΔT2=5min; Number of views Ln2=50; Viewing duration Lt2=200 seconds; Number of modifications Xn2=10; Number of shares Fn2=5;

[0082] For File3, user behavior data in time period T (30 minutes): ΔT3 = 8 minutes; number of views Ln3 = 100; browsing time Lt3 = 400 seconds; number of modifications Xn3 = 8; number of shares Fn3 = 4;

[0083] According to the formula Y = (ΔT / T)*[(Ln / SLn)*(Lt / SLt)+(Xn / SXn)+(Fn / SFn)], calculate the user attention Y:

[0084] Y_File1=(10 / 30)*[10 / (10+50+100)+500 / (500+200+400)+1 / (1+10+8)+0 / 0+5+4)=0.19;

[0085] Y_File2=(5 / 30)*[50 / (10+50+100)+200 / (500+200+400)+10 / (1+10+8)+5 / 0+5+4)=0.26;

[0086] Y_File3=(8 / 30)*[100 / (10+50+100)+400 / (500+200+400)+8 / (1+10+8)+4 / 0+5+4)=0.49;

[0087] In summary, the user attention threshold interval Q is [0.19, 0.49].

[0088] The real-time file classification storage module includes a real-time file user attention calculation unit, a user attention threshold comparison unit and a file classification processing unit;

[0089] The real-time file user attention calculation unit includes:

[0090] The user attention degree of the files received in real time is calculated every selected time period T0, and the calculation formula of the user attention degree Ys of the real-time files is:

[0091] Ys=(ΔT0 / T0)*[(Lsn / SLsn)*(Lst / SLst)+(Xsn / SXsn)+(Fsn / SFsn)],

[0092] Wherein, ΔT0 represents the difference between the generation time of each real-time received file and the current time, Lsn represents the number of times each real-time received file is browsed in the T0 time period, SLsn represents the sum of the number of times all real-time received files are browsed in the T0 time period, Xsn represents the number of times each real-time received file is modified in the T0 time period, SXsn represents the sum of the number of times all real-time received files are modified in the T0 time period, Fsn represents the number of times each real-time received file is shared in the T0 time period, and SFsn represents the sum of the number of times all real-time received files are shared in the T0 time period.

[0093] The user attention threshold comparison unit includes:

[0094] The obtained real-time file user attention Ys is compared with the user attention threshold interval Q. If Ys < q_min, it means that the user attention of the real-time file is low, which also reflects that the use value of this file to the user is low. In this case, the real-time received file is classified as a junk file and stored in the recycle bin. The recycle bin is set with an automatic cleaning time T1, and the recycle bin will be automatically cleaned every T1 period.

[0095] If q_min≤Ys≤q_max, it means that the user attention of the real-time file meets the attention threshold range, which also reflects that the file has a certain use value to the user. In this case, the real-time received file is classified as a temporary file and stored in the temporary storage space. The user attention Ys' is recalculated every T2 time period. When Ys'<q_min, the file is classified as a junk file and marked as "temporary file converted to junk file", and the storage address is also moved from the temporary storage space to the recycle bin; when Ys'>q_max, the file is classified as a permanent file and marked as "temporary file converted to permanent file", and the storage address is also moved from the temporary storage space to the permanent storage space;

[0096] If Ys>q_max, the file received in real time will be classified as a permanent file and stored in the permanent storage space, and the user attention Ys' will be recalculated every T3 time period. When q_min≤Ys'≤q_max, the file will be classified as a temporary file and marked as "permanent file converted to temporary file", and the storage address will be moved from the permanent storage space to the temporary space; when Ys'<q_min, the file will be classified as a junk file and marked as "permanent file converted to junk file", and the storage address will be moved from the permanent storage space to the recycle bin.

[0097] The file classification processing unit includes:

[0098] According to the comparison result of the user attention threshold comparison unit, the file is classified; if the user attention of the real-time file is lower than the minimum value of the threshold interval Q, it is classified as a junk file and stored in the recycle bin, and the recycle bin will clean up the junk file according to the set automatic cleaning time period; if the user attention of the real-time file is between the minimum value and the maximum value of the threshold interval Q, including the minimum value and the maximum value, it is stored as a temporary file in the temporary storage space, the user attention is recalculated every T2 time period, the classification type of the file is determined again according to the new user attention, and the labeling and storage position are adjusted; if the user attention of the real-time file is higher than the maximum value of the threshold interval Q, it is classified as a permanent file and stored in the permanent storage space, the user attention is recalculated every T3 time period, the classification type of the file is determined again according to the new user attention, and the labeling and storage position are adjusted.

[0099] The user attention of temporary files is calculated every T2 time period, which provides a dynamic adjustment function to prevent the user attention of temporary files from changing later, causing the temporary files to become junk files and occupy space; similarly, the user attention of long-term files is calculated every T3 time period, and T1>T2, T1>T3.

[0100] The file cleaning suggestion module includes a recycle bin file processing unit, a temporary storage space file processing unit, and a permanent storage space file processing unit;

[0101] The recycle bin file processing unit obtains the file list R in the recycle bin and determines whether the automatic cleaning time T1 has been reached; if the automatic cleaning time T1 has been reached, the files are included in the cleaning suggestion list, compressed and stored in the cloud space, and completely deleted from the recycle bin; if the automatic cleaning time T1 has not been reached, no processing is performed;

[0102] The role of cloud space is to prevent users from accidentally deleting files. When users need to retrieve deleted files, they can log in to the cloud space to search. Cloud space only serves as a storage function, and the management of cloud space requires manual management by users.

[0103] The temporary storage space file processing unit obtains the file list S in the temporary storage space, associates the user attention Ys' recalculated every T2 time period with the file list S; sends the notification information marked "temporary files are converted to junk files" to the user management module, and the user determines whether to delete them immediately. If the user gives an immediate deletion instruction, they are deleted immediately; if the user does not give a deletion instruction, the same subsequent operation is performed as for the files in the recycle bin; sends the notification information marked "temporary files are converted to permanent files" to the user management module, and the user confirms the final storage location of the file;

[0104] The long-term storage space file processing unit obtains the file list P in the long-term storage space, and associates the user attention Ys' recalculated every T3 time period with the file list P; sends the notification information marked "long-term storage files are converted to temporary files" to the user management module, and the user confirms the final storage location of the file; sends the notification information marked "long-term storage files are converted to junk files" to the user management module, and the user determines whether to delete them immediately. If the user gives an immediate deletion instruction, they are deleted immediately; if the user does not give a deletion instruction, the same subsequent operations are performed as for files in the recycle bin.

[0105] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0106] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A computer data analysis and management system based on big data, Features: The system includes a database, a data collection module, a data analysis module, a real-time file classification storage module, a file cleaning suggestion module and a user management module; The database stores all file information and all user behavior data of the files, wherein the file information includes save information and delete records, the save information records the basic attributes of the file, and the delete record records the deletion operation of the file; the user behavior data includes various operations of the user on the file, including the user's browsing history, modification history and sharing history; The data acquisition module obtains the file saving information and deletion records from the database, extracts the file size and performs sorting analysis; The data analysis module analyzes the file saving information and deletion records, extracts the target file, and calculates the feature vector and user attention according to the user behavior data; and obtains the threshold range of user attention through data analysis; The real-time file classification storage module classifies and stores the files received in real time according to the user attention, and recalculates the user attention within a certain time interval to perform dynamic adjustment; The file cleaning suggestion module analyzes and judges the file lists in the recycle bin, temporary storage space and permanent storage space, and sends corresponding notification information to the user management module to suggest file cleaning operations; The user management module receives notification information from the file cleaning suggestion module, the user performs corresponding file management on the received notification information, and sends feedback information to the file cleaning suggestion module to inform the user of the selection and execution result.

2. A computer data analysis and management system based on big data according to claim 1, Features: The data acquisition module includes a file storage information extraction unit, a file deletion record extraction unit, a file size extraction unit and an empty file identification unit; The file saving information extraction unit and the file deletion record extraction unit obtain the saving information and deletion records of all files in the database, and construct a file saving information set A according to the timestamp sequence of file generation in the saving information, and A={ai}, where i is a positive integer from 1 to n, and ai represents the i-th file saving information; construct a file deletion record set B according to the timestamp sequence of deleted files in the deletion record, and B={bj}, where j is a positive integer from 1 to m, and bj represents the j-th deletion record; The file size extraction unit extracts the sizes of all files in the file storage information set A and forms a file size data table; arranges the data in the file size data table in descending order, obtains the file name with a file size of 0 bytes, extracts the file name of the file deletion record set B and searches for it; the empty file identification unit marks the file in set A that corresponds to the file name in set B and marks it as an empty file.

3. A computer data analysis and management system based on big data according to claim 1, Features: The data analysis module includes a file name comparison unit, a target file feature extraction unit, a feature vector correlation calculation unit and a user attention calculation unit; The file name comparison unit extracts the file names of all files from the file saving information set A and the file deletion record set B, and compares the file names of the set A with those of the set B; Obtain the file names that belong only to set A, filter out empty files, and mark the remaining file names that belong only to set A as target files 1; the target file feature extraction unit obtains the user behavior data of the target file 1, obtains the user browsing frequency Lw, user modification frequency Xw, and user sharing frequency Fw of each target file 1 in the selected time period T, and forms a feature vector, which is represented by V1k=(Lwk, Xwk, Fwk), and V1k represents the feature vector of the kth target file 1, where k represents the number of target files 1 in the selected time period, which is a positive integer; Get the file name that belongs to both set A and set B, and mark it as target file 2; Obtain the saving information and deletion record of the file name corresponding to the target file 2, extract the timestamp TA generated by each target file 2 and the timestamp TB deleted by each target file 2, and calculate the time difference ΔT between TA and TB, ΔT = TB-TA; For each target file 2, the user browsing frequency Lw, user modification frequency Xw, and user sharing frequency Fw within the ΔT time period are obtained and formed into a feature vector, which is represented by V2 d =(Lw d ,Xw d ,Fw d ), and V2 d The feature vector representing the dth target file 2, where d represents the number of target files 2 in the ΔT time period, and is a positive integer; For each feature vector of target file 2, a feature vector correlation calculation unit is used to perform vector linear correlation operations on the feature vectors in target file 1 to obtain the file names of the files corresponding to the feature vectors in target file 2 that are linearly correlated with target file 1, forming a set C.

4. A computer data analysis and management system based on big data according to claim 3, Features: The user attention calculation unit comprises: Obtain the user behavior data of each target file 2, and record the number of views Ln and the browsing duration Lt, the number of modifications Xn, and the number of shares Fn within the ΔT time period for each target file 2; calculate the user attention Y of each target file 2, and the specific calculation formula is: Y = (ΔT / T)*[(Ln / SLn)*(Lt / SLt)+(Xn / SXn)+(Fn / SFn)], where T represents the time period from the generation time of the first target file 2 to the deletion time of the nth target file 2, SLn represents the sum of the number of views of all target files 2 in the time period T, SLt represents the sum of the viewing duration of all target files 2 in the time period T, SXn represents the sum of the number of modifications of all target files 2 in the time period T, and SFn represents the sum of the number of shares of all target files 2 in the time period T; Get the user attention of target file 2 corresponding to all file names in set C, and get the user attention threshold interval Q, and Q = [q_min, q_max], where q_min is the minimum user attention of target file 2 in set C, and q_max is the maximum user attention of target file 2 in set C.

5. A computer data analysis and management system based on big data according to claim 1, Features: The real-time file classification storage module includes a real-time file user attention calculation unit, a user attention threshold comparison unit and a file classification processing unit; The real-time file user attention calculation unit comprises: The user attention degree of the files received in real time is calculated every selected time period T0, and the calculation formula of the user attention degree Ys of the real-time files is: Ys=(ΔT0 / T0)*[(Lsn / SLsn)*(Lst / SLst)+(Xsn / SXsn)+(Fsn / SFsn)], Wherein, ΔT0 represents the difference between the generation time of each real-time received file and the current time, Lsn represents the number of times each real-time received file is browsed in the T0 time period, SLsn represents the sum of the number of times all real-time received files are browsed in the T0 time period, Xsn represents the number of times each real-time received file is modified in the T0 time period, SXsn represents the sum of the number of times all real-time received files are modified in the T0 time period, Fsn represents the number of times each real-time received file is shared in the T0 time period, and SFsn represents the sum of the number of times all real-time received files are shared in the T0 time period.

6. A computer data analysis and management system based on big data according to claim 5, Features: The user attention threshold comparison unit comprises: The obtained real-time file user attention Ys is compared with the user attention threshold interval Q. If Ys < q_min, the real-time received file is classified as a junk file and stored in the recycle bin. The recycle bin is set with an automatic cleaning time T1. The recycle bin will automatically clean up every T1 period. If q_min≤Ys≤q_max, the file received in real time is classified as a temporary file and stored in the temporary storage space, and the user attention Ys' is recalculated every T2 time period. When Ys'<q_min, the file is classified as a junk file and marked as "temporary file converted to junk file", and the storage address is also moved from the temporary storage space to the recycle bin; when Ys'>q_max, the file is classified as a permanent file and marked as "temporary file converted to permanent file", and the storage address is also moved from the temporary storage space to the permanent storage space; If Ys>q_max, the file received in real time will be classified as a permanent file and stored in the permanent storage space, and the user attention Ys' will be recalculated every T3 time period. When q_min≤Ys'≤q_max, the file will be classified as a temporary file and marked as "permanent file converted to temporary file", and the storage address will be moved from the permanent storage space to the temporary space; when Ys'<q_min, the file will be classified as a junk file and marked as "permanent file converted to junk file", and the storage address will be moved from the permanent storage space to the recycle bin.

7. A computer data analysis and management system based on big data according to claim 6, Features: The file classification processing unit comprises: According to the comparison result of the user attention threshold comparison unit, the file is classified; if the user attention of the real-time file is lower than the minimum value of the threshold interval Q, it is classified as a junk file and stored in the recycle bin, and the recycle bin will clean up the junk file according to the set automatic cleaning time period; if the user attention of the real-time file is between the minimum value and the maximum value of the threshold interval Q, including the minimum value and the maximum value, it is stored as a temporary file in the temporary storage space, the user attention is recalculated every T2 time period, the classification type of the file is determined again according to the new user attention, and the labeling and storage position are adjusted; if the user attention of the real-time file is higher than the maximum value of the threshold interval Q, it is classified as a permanent file and stored in the permanent storage space, the user attention is recalculated every T3 time period, the classification type of the file is determined again according to the new user attention, and the labeling and storage position are adjusted.

8. A computer data analysis and management system based on big data according to claim 1, Features: The file cleaning suggestion module includes a recycle bin file processing unit, a temporary storage space file processing unit and a permanent storage space file processing unit; The recycle bin file processing unit obtains the file list R in the recycle bin and determines whether the automatic cleaning time T1 has been reached; if the automatic cleaning time T1 has been reached, the files are included in the cleaning suggestion list, compressed and stored in the cloud space, and completely deleted from the recycle bin; if the automatic cleaning time T1 has not been reached, no processing is performed; The temporary storage space file processing unit obtains the file list S in the temporary storage space, associates the user attention Ys' recalculated every T2 time period with the file list S; sends the notification information marked "temporary files are converted to junk files" to the user management module, and the user determines whether to delete them immediately. If the user gives an immediate deletion instruction, they are deleted immediately; if the user does not give a deletion instruction, the same subsequent operation is performed as for the files in the recycle bin; sends the notification information marked "temporary files are converted to permanent files" to the user management module, and the user confirms the final storage location of the file; The long-term storage space file processing unit obtains a file list P in the long-term storage space, and associates the user attention Ys' recalculated every T3 time period with the file list P; Send the notification message marked "Converting permanent files to temporary files" to the user management module, and the user confirms the final storage location of the file; A notification message marked "long-lived files converted to junk files" is sent to the user management module, and the user decides whether to delete it immediately. If the user gives an immediate deletion instruction, it is deleted immediately; if the user does not give a deletion instruction, the same subsequent operation is performed as for files in the recycle bin.

Citation Information

Patent Citations

  • Big data-based analysis algorithm

    CN112579671A

  • Computer resource management system and method based on big data

    CN115952140A