Archive management method, system and program
By extracting features, determining the cutoff distance and calculating density indicators in archive management, the limitations of existing clustering algorithms in archive data processing are solved, and more accurate archive classification and management are achieved.
Patent Information
- Application Number
- CN202510213220.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing clustering algorithms have limitations in archival data processing, such as sensitivity to initial clustering center selection, high computational complexity, and sensitivity to noise and outliers.
By extracting the features of the classified archives, determining the target features and initial truncation distance, adjusting the truncation distance to maximize the weighted local density of the archives, calculating standardized local density and high density minimum distances, determining the noise and cluster center, and achieving more accurate archive classification.
It improves clustering accuracy in archive management, reduces data dimensions, automatically determines the cutoff distance, adapts to complex density distribution, and reduces misclassification of noise archives.
Smart Images

Figure CN120104854A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of archive management, and in particular to an archive management method, system and program. Background Art
[0002] With the increasing number of archives, traditional manual management and simple classification methods can no longer meet the needs of modern archive management. Traditional management methods often rely on manual classification, manual labeling or classification systems based on simple rules, resulting in low management efficiency, large errors, and difficulty in adapting to the variability of archive content and the management needs of massive data. Cluster analysis, as an unsupervised learning method, can automatically divide archives into different categories based on the similarity between archive contents, thereby realizing automatic organization and management of archives, helping managers to achieve more efficient archive classification and query.
[0003] However, existing clustering algorithms still have some limitations when processing archival data. For example, the traditional K-means algorithm is very sensitive to the selection of the initial cluster center and is prone to fall into the local optimal solution. Although hierarchical clustering can construct a tree structure, it has a large amount of calculation when processing large-scale data and is highly sensitive to noise and outliers. The Density Peak Clustering (DPC) algorithm has the advantages of being able to automatically discover clusters of any shape and not requiring the number of clusters to be set in advance. However, the choice of the cutoff distance has a great impact on the clustering results, and the computational complexity is high when processing high-dimensional data. These are all problems that need to be solved urgently in archival management. Summary of the invention
[0004] The present invention provides an archive management method, system and program, which can not only reduce data dimension and automatically determine the truncation distance, but also improve the local density in DPC, thereby improving the classification accuracy in the clustering of archive management.
[0005] Specifically, in a first aspect, the present invention provides a file management method, the method comprising: Extract the features of the classified files, obtain the number of categories of the classified files, cluster each feature, and determine the target feature from the features based on the clustering results and the classification of the classified files; determine the initial cutoff distance based on the classification of the classified files and the target feature; Extract the target feature from all the archives including the classified archives and the archives to be classified to obtain the feature vector of the archives, adjust the initial cutoff distance within a preset range so that the weighted local density of the archives is maximized and obtain the cutoff distance, and calculate the standardized local density and the high-density minimum distance of each archive according to the cutoff distance; Based on the standardized local density of the archives and the minimum distance of high density, the noise and several cluster centers of the categories are determined, and the non-noise archives are clustered according to the cluster centers to complete the archive classification.
[0006] In some implementations of the first aspect, determining the initial cutoff distance based on the classification status of the classified archives and the target feature is specifically: The target feature is used to calculate the average distance between all archives in each archive category; Calculate the minimum distance between the centers of archives of different categories; An initial cutoff distance is obtained according to the average distance and the minimum distance.
[0007] In some implementations of the first aspect, the initial cutoff distance is adjusted within a preset range so that the weighted local density of the archive is maximized and the cutoff distance is obtained, specifically: Determine a preset range according to the initial cutoff distance, and increase the cutoff distance according to a preset step length starting from the lower limit of the preset range; After obtaining a new cutoff distance each time, the distribution balance and local density of each file within the cutoff distance are calculated by using the feature vector, and a weighted local density is obtained according to the distribution balance and local density; When the cutoff distance reaches the upper limit of the preset range, the maximum value of the weighted local density and the cutoff distance corresponding to the maximum value of the weighted local density are counted.
[0008] In some implementations of the first aspect, the calculating of the normalized local density and the high-density minimum distance of each archive according to the cutoff distance is specifically: The standardized local density of the archives was calculated using the maximum weighted local density of the archives and the cutoff distance; The distance from each file to the nearest neighbor file with higher normalized local density is calculated, and the distance is taken as the high-density minimum distance of the file.
[0009] In some implementations of the first aspect, the determining of the noise and the number of cluster centers of the category based on the archive-based standardized local density and high-density minimum distance is specifically: Selecting several files of the category with the largest standardized distance density and the largest high density from the decision graph composed of the standardized local density and the minimum high density distance as the cluster center; The cluster center is deleted from the decision diagram, and the standardized local density average and high-density minimum distance average of the remaining files are calculated. If the standardized local density of the file is less than the standardized local density average, and the high-density minimum distance is greater than the high-density minimum distance average, the file is regarded as noise, and the file corresponding to the noise is manually classified.
[0010] In some implementations of the first aspect, the determining of the noise and the number of cluster centers of the category based on the archive-based standardized local density and high-density minimum distance is specifically: A decision graph of classified archives is constructed based on the standardized local density and high-density minimum distance of classified archives, and a global decision graph is constructed based on the standardized local density and high-density minimum distance of classified archives and archives to be classified; Obtain the maximum value of the high-density minimum distance and the minimum value of the standardized local density in the classified archive decision graph except for the cluster center, and determine the adjusted maximum value of the high-density minimum distance and the adjusted minimum value of the standardized local density according to the classified archive and all the archives; If in the global decision graph, the high-density minimum distance of the file is greater than the adjusted high-density minimum distance maximum value, and the standardized local density is less than the adjusted standardized local density minimum value, the file is regarded as noise, and the file corresponding to the noise is manually classified.
[0011] In a second aspect, the present invention provides a file management system, the system comprising: The feature extraction module is used to extract the features of the classified files, obtain the number of categories of the classified files, cluster each feature, determine the target feature from the features according to the clustering results and the classification of the classified files; determine the initial cutoff distance based on the classification of the classified files and the target feature; A distance calculation module is used to extract the target feature from all archives including classified archives and archives to be classified to obtain a feature vector of the archives, adjust the initial cutoff distance within a preset range so that the weighted local density of the archives is maximized and obtain the cutoff distance, and calculate the standardized local density and high-density minimum distance of each archive according to the cutoff distance; The classification module is used to determine the noise and several cluster centers of the category based on the standardized local density and high-density minimum distance of the archives, and cluster the non-noise archives according to the cluster centers to complete the archive classification.
[0012] In a third aspect, the present invention provides a computer executable program, the program comprising: The feature extraction module is used to extract the features of the classified files, obtain the number of categories of the classified files, cluster each feature, determine the target feature from the features according to the clustering results and the classification of the classified files; determine the initial cutoff distance based on the classification of the classified files and the target feature; A distance calculation module is used to extract the target feature from all archives including classified archives and archives to be classified to obtain a feature vector of the archives, adjust the initial cutoff distance within a preset range so that the weighted local density of the archives is maximized and obtain the cutoff distance, and calculate the standardized local density and high-density minimum distance of each archive according to the cutoff distance; The classification module is used to determine the noise and several cluster centers of the category based on the standardized local density and high-density minimum distance of the archives, and cluster the non-noise archives according to the cluster centers to complete the archive classification.
[0013] In a fourth aspect, the present invention provides a computer-readable medium having instructions stored thereon, which, when executed on a computer, causes the computer to execute the archive management method provided in the first aspect and various possible implementations of the first aspect.
[0014] The present invention allocates an independent cutoff distance to each file, which can better handle files with complex density distribution; and by introducing distribution balance, it can distinguish files with uniform distribution of surrounding files, and avoid taking files with surrounding files distributed on one side as cluster centers. In the identification of noisy files, the noise range is expanded to avoid files being misclassified. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a flow chart of Embodiment 1; Figure 2 A schematic diagram of the cluster center in the decision graph; Figure 3 A schematic diagram of noise in a decision graph; Figure 4 is a flowchart of step 103; Figure 5 is a structural diagram of Embodiment 2; Figure 6 A block diagram of a computer device. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described in detail below in conjunction with the drawings and specific implementation methods of the specification. It should be understood that in the various embodiments of the present application, the size of the sequence number of each process does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. It should be understood that determining B based on A does not mean determining B only based on A, and B can also be determined based on A and / or other information.
[0017] Figure 1 A flow chart of a first embodiment of the present invention is shown. Figure 1 As shown, the archive management method includes: 101, extracting features of the classified files, obtaining the number of categories of the classified files, clustering each feature, and determining a target feature from the features according to the clustering result and the classification of the classified files; determining an initial cutoff distance based on the classification of the classified files and the target feature; For the classified archives, the features of the archives are extracted by means of keywords, etc. Keywords include but are not limited to name, date of joining the work, work unit, graduation school, major, etc. Keywords can also include other structured and unstructured content in the archives, such as pictures, etc. In one embodiment, keywords are used as features, and the content corresponding to the keywords in the archives is used as feature values. After the feature values are obtained, the non-digital content is encoded to obtain the feature values of each feature after encoding, and the feature values of all features constitute the feature vectors of the features.
[0018] Since the classified files have been classified, each classified file has a category. Since there are many features of the files, in order to reduce the dimension, the classified files are re-clustered, and then the target features are determined based on the re-clustering results and the classification situation, thereby achieving dimensionality reduction of the features. Specifically, each feature is clustered separately, and the degree of conformity between the clustering results and the classification situation of the classified files is judged. If the degree of conformity is high, the feature is used as the target feature; for example, for the feature "professional", the classified files are re-clustered according to "professional". If the clustered clusters have a high degree of overlap with the classified file categories, "professional" is used as the target feature. In a specific embodiment, the target feature is determined from the features according to the clustering results and the classification of the classified files, specifically: the classified files are clustered according to each feature, the number of clusters in the cluster is the number of categories, the label of the file with the largest number of labels with the same category in the cluster is used as the label of the cluster, the ratio of the number of files in the cluster with the same label as the cluster to the number of all files in the cluster is calculated, the average value of the ratios of all clusters is calculated, the average value is used as the degree of conformity, the features are sorted in order of conformity from large to small, and the first n features are selected as target features, wherein n is a positive integer; or a feature with a conformity greater than the average conformity is selected as the target feature.
[0019] After obtaining the target feature, the cutoff distance of the DPC is further determined according to the classification status of the classified archives and the target feature. In one embodiment, the initial cutoff distance is determined by continuous attempts. Specifically, a preset initial cutoff distance set is obtained, and the initial cutoff distance in the initial cutoff distance set is used as the cutoff distance of the DPC. The classified archives are clustered using the target feature, and the degree of conformity between each cutoff distance cluster and the classified archives is calculated, and the truncation distance with the greatest degree of conformity is used as the initial truncation distance.
[0020] In yet another embodiment, the initial cutoff distance is determined based on the classification of the classified archives and the target feature, specifically: The target feature is used to calculate the average distance between all archives in each archive category; Calculate the minimum distance between the centers of archives of different categories; An initial cutoff distance is obtained according to the average distance and the minimum distance.
[0021] Obtain a feature vector composed of target features of classified archives, wherein each value in the feature vector is a value corresponding to a target feature. For example, if there are 3 target features, the feature vector is 1×3. Use the feature vector to calculate the average distance of all archives in the same category; at the same time, calculate the minimum distance between the centers of archives in different categories. First, calculate the center vector of each category of archives. Preferably, the center vector is the average value or centroid of the feature vector of a category of archives. Then calculate the distance between the center vectors of different categories, and take the shortest distance of the center vectors of all categories as the minimum distance. After obtaining the average distance and the minimum distance, obtain the initial truncation distance according to the average distance and the minimum distance. For example, take the average value of the average distance and the minimum distance as the initial truncation distance, or weight the average value as the initial truncation distance.
[0022] 102, extracting the target feature from all the archives including the classified archives and the archives to be classified to obtain the feature vector of the archives, adjusting the initial cutoff distance within a preset range so that the weighted local density of the archives is maximized and obtaining the cutoff distance, and calculating the standardized local density and the high-density minimum distance of each archive according to the cutoff distance; For all archives including those that have been classified and those to be classified, information about the target features is extracted and converted into a numerical form that can be processed by a computer. For example, if the target feature is a keyword, the value corresponding to the keyword is converted into a vector using a bag-of-words model or a word vector model, and the vectors of all target features constitute the feature vector of the archive. In order to find the best cutoff distance, a preset range is set near the initial cutoff distance, such as ±20% of the initial cutoff distance. Within the preset range, the cutoff distance is gradually adjusted, and the weighted local density of each archive is calculated. By adjusting the cutoff distance, the cutoff distance that makes the weighted local density of the archive reach the maximum is found, and the cutoff distance with the maximum weighted local density is used as the cutoff distance of the archive. The cutoff distances of different archives are different. According to the determined cutoff distance, the standardized local density of each archive is calculated, and the standardized local density is the ratio of the weighted local density at the cutoff distance of the archive to the cutoff distance. The distance from each archive to the archive with a higher standardized local density and the closest distance is calculated to obtain the high-density minimum distance.
[0023] The traditional local density calculation method does not fully consider the uniformity of file distribution. For example, if the local density of a file is very large, but it is distributed on one side or in one direction of the file, the file is not suitable as a cluster center even though the local density is very large. In a specific embodiment, the initial cutoff distance is adjusted within a preset range so that the weighted local density of the file is maximized and the cutoff distance is obtained, specifically: Determine a preset range according to the initial cutoff distance, and increase the cutoff distance according to a preset step length starting from the lower limit of the preset range; After obtaining a new cutoff distance each time, the distribution balance and local density of each file within the cutoff distance are calculated by using the feature vector, and a weighted local density is obtained according to the distribution balance and local density; When the cutoff distance reaches the upper limit of the preset range, the maximum value of the weighted local density and the cutoff distance corresponding to the maximum value of the weighted local density are counted.
[0024] Based on the determined initial truncation distance, a preset range is defined, and starting from the minimum value of the range, the truncation distance is gradually increased with a fixed increment, i.e., the preset step size. For example, if the initial truncation distance is 10, the preset range can be [9, 11], and the step size is 0.5, then the truncation distance starts from 9 and increases by 0.5 each time until it reaches 11.
[0025] Using the characteristic vector of the archive, the local density of each archive within its current cutoff distance is calculated, that is, the density of the surrounding archives, and the uniformity of the distribution of these archives is evaluated. Then, the local density and distribution balance are combined to calculate a comprehensive density index, namely the weighted local density, in a weighted manner. Among them, the calculation method of distribution balance includes but is not limited to calculating the density of different areas around the archive. If the density distribution is uniform, the distribution balance is high, or calculating the angle between the archive and the vector of other archives within its cutoff distance. If the angle distribution is relatively uniform, it means that the archive distribution is relatively uniform and the value of distribution balance is large; if the angle distribution is relatively concentrated, it means that the archive distribution is relatively uneven and the value of distribution balance is small. In one embodiment, the weighted local density is equal to the product or weighted sum of the distribution balance and the local density.
[0026] When the cutoff distance increases to the maximum value of the preset range, the search ends. During this process, the weighted local density corresponding to each cutoff distance is recorded. By comparing these values, the maximum weighted local density is found and the corresponding cutoff distance is determined.
[0027] In a specific embodiment, the normalized local density and high-density minimum distance of each file are calculated according to the cutoff distance, specifically: The standardized local density of the archives was calculated using the maximum weighted local density of the archives and the cutoff distance; The distance from each file to the nearest neighbor file with higher normalized local density is calculated, and the distance is taken as the high-density minimum distance of the file.
[0028] After obtaining the maximum weighted local density and cutoff distance of the file, the maximum weighted local density is divided by the cutoff distance as the normalized local density. For example, if the maximum weighted local density of file A is 8 and the cutoff distance is 10, then the normalized local density of file A is 0.8.
[0029] For each archive, find all archives with higher normalized local density than it, calculate the distance from this archive to these archives, and take the minimum value of the distance as the high-density minimum distance of this archive.
[0030] 103, based on the standardized local density of the archives and the minimum distance of high density, determine the noise and several cluster centers of the category, cluster the non-noise archives according to the cluster centers to complete the archive classification.
[0031] After calculating the standardized local density and the minimum distance of high density, a decision diagram with the standardized local density as the horizontal axis and the minimum distance of high density as the vertical axis is constructed, and each file is represented as a point. From the decision diagram composed of the standardized local density and the minimum distance of high density, several files of the category with the largest standardized distance density and the largest high density are selected as cluster centers; the cluster centers are deleted from the decision diagram, and the average standardized local density and the average high density minimum distance of the remaining files are calculated. If the standardized local density of the file is less than the average standardized local density, and the minimum distance of high density is greater than the average high density minimum distance, the file is regarded as noise, and the files corresponding to the noise are manually classified.
[0032] The point in the upper right corner of the decision graph, i.e., the file with the highest normalized local density and the highest minimum distance to high density, is surrounded by many similar files and is far away from other high-density areas. It is the core representative of each category. Several such points of the category are selected as cluster centers, such as Figure 2 As shown. The selected cluster centers are deleted from the decision diagram to facilitate subsequent analysis of the remaining archives. The standardized local density average and the high-density minimum distance average of the remaining archives are calculated. For each remaining archive, if its standardized local density is less than the average value and the high-density minimum distance is greater than the average value, the archive is considered to be noise, as shown in Figure 3 As shown in the figure, noise files are usually quite different from other files, making them difficult to classify automatically or easily misclassified, so the noise files are manually classified. The remaining non-noise files are assigned to the corresponding categories according to their similarity to the centers of each cluster, completing the automatic classification of the files.
[0033] In order to more accurately identify the noise profile, Figure 4 In another embodiment, the noise and the number of cluster centers of the category are determined based on the standardized local density and high density minimum distance of the archive, specifically: 301, constructing a classified archive decision graph according to the standardized local density and high-density minimum distance of the classified archives, and constructing a global decision graph according to the standardized local density and high-density minimum distance of the classified archives and the archives to be classified; 302, obtaining the maximum value of the high-density minimum distance and the minimum value of the standardized local density in the classified archive decision graph except for the cluster center, and determining the adjusted maximum value of the high-density minimum distance and the adjusted minimum value of the standardized local density according to the classified archive and all the archives; 303. If in the global decision graph, the high-density minimum distance of the file is greater than the adjusted high-density minimum distance maximum value, and the standardized local density is less than the adjusted standardized local density minimum value, the file is regarded as noise, and the file corresponding to the noise is manually classified.
[0034] Construct two decision graphs, one only contains classified archives, and the other contains all archives. The classified archive decision graph can provide the decision boundary of the classified archives, and the global decision graph provides the decision boundary of all archives. By comparing the two decision graphs, the noise points can be found more accurately. Obtain the points in the classified archive decision graph, excluding the cluster center, and find the maximum value of the high-density minimum distance and the minimum value of the standardized local density. Determine the adjusted high-density minimum distance maximum value and the adjusted standardized local density minimum value according to the classified archives and all the archives. Specifically, calculate the quantity ratio of all archives to the classified archives, and use the quantity ratio to adjust the high-density minimum distance maximum value and the standardized local density minimum value. Preferably, the product of the high-density minimum distance maximum value and the inverse of the quantity ratio is used as the adjusted high-density minimum distance maximum value, and the product of the standardized local density minimum value and the quantity ratio is used as the adjusted standardized local density minimum value. In the global decision graph, if the high-density minimum distance of a point is greater than the adjusted high-density minimum distance maximum value, and the standardized local density is less than the adjusted standardized local density minimum value, then this point is classified as noise. By using the information of classified archives to build a special decision diagram and determine the noise threshold, it is possible to identify noise more accurately and reduce misjudgments.
[0035] Example 2 Figure 5 The file management system shown in the figure comprises: The feature extraction module is used to extract the features of the classified files, obtain the number of categories of the classified files, cluster each feature, determine the target feature from the features according to the clustering results and the classification of the classified files; determine the initial cutoff distance based on the classification of the classified files and the target feature; A distance calculation module is used to extract the target feature from all archives including classified archives and archives to be classified to obtain a feature vector of the archives, adjust the initial cutoff distance within a preset range so that the weighted local density of the archives is maximized and obtain the cutoff distance, and calculate the standardized local density and high-density minimum distance of each archive according to the cutoff distance; The classification module is used to determine the noise and several cluster centers of the category based on the standardized local density and high-density minimum distance of the archives, and cluster the non-noise archives according to the cluster centers to complete the archive classification.
[0036] In a third embodiment, the present invention provides a computer executable program, the program comprising: The feature extraction module is used to extract the features of the classified files, obtain the number of categories of the classified files, cluster each feature, determine the target feature from the features according to the clustering results and the classification of the classified files; determine the initial cutoff distance based on the classification of the classified files and the target feature; A distance calculation module is used to extract the target feature from all archives including classified archives and archives to be classified to obtain a feature vector of the archives, adjust the initial cutoff distance within a preset range so that the weighted local density of the archives is maximized and obtain the cutoff distance, and calculate the standardized local density and high-density minimum distance of each archive according to the cutoff distance; The classification module is used to determine the noise and several cluster centers of the category based on the standardized local density and high-density minimum distance of the archives, and cluster the non-noise archives according to the cluster centers to complete the archive classification.
[0037] In a fourth embodiment, the present invention provides a computer-readable medium having instructions stored thereon. When the instructions are executed on a computer, the computer is enabled to execute the archive management method provided in the first embodiment and various possible implementations of the first embodiment.
[0038] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, instructions may be distributed over a network or through other computer-readable media. Therefore, a machine-readable medium may include any mechanism for storing or transmitting information in a machine (e.g., computer) readable form, including, but not limited to, a floppy disk, an optical disk, an optical disk, a magneto-optical disk, a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic card or an optical card, a flash memory, or a tangible machine-readable memory for transmitting information (e.g., a carrier wave, an infrared signal, a digital signal, etc.) using the Internet in an electrical, optical, acoustic, or other form of propagation signal. Accordingly, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (eg, a computer).
[0039] References to "one embodiment" or "an embodiment" in the specification mean that the specific features, structures, or characteristics described in conjunction with the embodiment are included in at least one exemplary implementation or technology disclosed according to the embodiment of the present application. The appearance of the phrase "in one embodiment" in various places in the specification does not necessarily all refer to the same embodiment.
[0040] The embodiment of the present invention can also be Figure 6 The computer device shown in the specification is executed, and the computer device includes at least a processor and a memory. The device can be constructed specifically for the required purpose or it can include a general-purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program can be stored in a computer-readable medium, such as, but not limited to, any type of disk, including a floppy disk, an optical disk, a CD-ROM, a magneto-optical disk, a read-only memory (ROM), a random access memory (RAM), an EPROM, an EEPROM, a magnetic or optical card, an application-specific integrated circuit (ASIC) or any type of medium suitable for storing electronic instructions, and can be coupled to a computer system bus. In addition, the computer mentioned in the specification may include a single processor or may be an architecture involving multiple processors for increased computing power.
[0041] In addition, the language used in this specification has been primarily selected for readability and instructional purposes and may not be selected to describe or limit the disclosed subject matter. Therefore, the present application embodiment disclosure is intended to illustrate rather than limit the scope of the concepts discussed herein.
Claims
1. A file management method, characterized in that: The method comprises: Extract the features of the classified files, obtain the number of categories of the classified files, cluster each feature, and determine the target feature from the features based on the clustering results and the classification of the classified files; determine the initial cutoff distance based on the classification of the classified files and the target feature; Extract the target feature from all the archives including the classified archives and the archives to be classified to obtain the feature vector of the archives, adjust the initial cutoff distance within a preset range so that the weighted local density of the archives is maximized and obtain the cutoff distance, and calculate the standardized local density and the high-density minimum distance of each archive according to the cutoff distance; Based on the standardized local density of the archives and the minimum distance of high density, the noise and several cluster centers of the categories are determined, and the non-noise archives are clustered according to the cluster centers to complete the archive classification.
2. The method according to claim 1, characterized in that The initial cutoff distance is determined based on the classification of the classified archives and the target features, specifically: The target feature is used to calculate the average distance between all archives in each archive category; Calculate the minimum distance between the centers of archives of different categories; An initial cutoff distance is obtained according to the average distance and the minimum distance.
3. The method according to claim 1, characterized in that The initial cutoff distance is adjusted within a preset range so that the weighted local density of the archive is maximized and the cutoff distance is obtained, specifically: Determine a preset range according to the initial cutoff distance, and increase the cutoff distance according to a preset step length starting from the lower limit of the preset range; After obtaining a new cutoff distance each time, the distribution balance and local density of each file within the cutoff distance are calculated by using the feature vector, and a weighted local density is obtained according to the distribution balance and local density; When the cutoff distance reaches the upper limit of the preset range, the maximum value of the weighted local density and the cutoff distance corresponding to the maximum value of the weighted local density are counted.
4. The method according to claim 1, characterized in that The standardized local density and high-density minimum distance of each file are calculated according to the cutoff distance, specifically: The standardized local density of the archives was calculated using the maximum weighted local density of the archives and the cutoff distance; The distance from each file to the nearest neighbor file with higher normalized local density is calculated, and the distance is taken as the high-density minimum distance of the file.
5. The method according to claim 1, characterized in that The standardized local density and high density minimum distance based on the archive determine the noise and the cluster centers of the category, specifically: Selecting several files of the category with the largest standardized distance density and the largest high density from the decision graph composed of the standardized local density and the minimum high density distance as the cluster center; The cluster center is deleted from the decision diagram, and the standardized local density average and high-density minimum distance average of the remaining files are calculated. If the standardized local density of the file is less than the standardized local density average, and the high-density minimum distance is greater than the high-density minimum distance average, the file is regarded as noise, and the file corresponding to the noise is manually classified.
6. The method according to claim 1, characterized in that The standardized local density and high density minimum distance based on the archive determine the noise and the cluster centers of the category, specifically: A decision graph of classified archives is constructed based on the standardized local density and high-density minimum distance of classified archives, and a global decision graph is constructed based on the standardized local density and high-density minimum distance of classified archives and archives to be classified; Obtain the maximum value of the high-density minimum distance and the minimum value of the standardized local density in the classified archive decision graph except for the cluster center, and determine the adjusted maximum value of the high-density minimum distance and the adjusted minimum value of the standardized local density according to the classified archive and all the archives; If in the global decision graph, the high-density minimum distance of the file is greater than the adjusted high-density minimum distance maximum value, and the standardized local density is less than the adjusted standardized local density minimum value, the file is regarded as noise, and the file corresponding to the noise is manually classified.
7. A file management system, characterized in that: The system comprises: The feature extraction module is used to extract the features of the classified files, obtain the number of categories of the classified files, cluster each feature, determine the target feature from the features according to the clustering results and the classification of the classified files; determine the initial cutoff distance based on the classification of the classified files and the target feature; A distance calculation module is used to extract the target feature from all archives including classified archives and archives to be classified to obtain a feature vector of the archives, adjust the initial cutoff distance within a preset range so that the weighted local density of the archives is maximized and obtain the cutoff distance, and calculate the standardized local density and high-density minimum distance of each archive according to the cutoff distance; The classification module is used to determine the noise and several cluster centers of the category based on the standardized local density and high-density minimum distance of the archives, and cluster the non-noise archives according to the cluster centers to complete the archive classification.
8. The system according to claim 7, characterized in that The initial cutoff distance is determined based on the classification of the classified archives and the target features, specifically: The target feature is used to calculate the average distance between all archives in each archive category; Calculate the minimum distance between the centers of archives of different categories; An initial cutoff distance is obtained according to the average distance and the minimum distance.
9. The system according to claim 7, characterized in that The initial cutoff distance is adjusted within a preset range so that the weighted local density of the archive is maximized and the cutoff distance is obtained, specifically: Determine a preset range according to the initial cutoff distance, and increase the cutoff distance according to a preset step length starting from the lower limit of the preset range; After obtaining a new cutoff distance each time, the distribution balance and local density of each file within the cutoff distance are calculated by using the feature vector, and a weighted local density is obtained according to the distribution balance and local density; When the cutoff distance reaches the upper limit of the preset range, the maximum value of the weighted local density and the cutoff distance corresponding to the maximum value of the weighted local density are counted.
10. A computer executable program, characterized in that: The procedures include: The feature extraction module is used to extract the features of the classified files, obtain the number of categories of the classified files, cluster each feature, determine the target feature from the features according to the clustering results and the classification of the classified files; determine the initial cutoff distance based on the classification of the classified files and the target feature; A distance calculation module is used to extract the target feature from all archives including classified archives and archives to be classified to obtain a feature vector of the archives, adjust the initial cutoff distance within a preset range so that the weighted local density of the archives is maximized and obtain the cutoff distance, and calculate the standardized local density and high-density minimum distance of each archive according to the cutoff distance; The classification module is used to determine the noise and several cluster centers of the category based on the standardized local density and high-density minimum distance of the archives, and cluster the non-noise archives according to the cluster centers to complete the archive classification.