A density clustering method and system based on SOM network

Through the density clustering method based on SOM network, the density hierarchical diffusion and topological relationship merge similar clusters are used to solve the problems of mismatch and leakage clustering in traditional clustering algorithms, and high-precision data clustering is achieved.

CN120086624BActive Publication Date: 2025-07-22JIANGXI CLOUD EYE VISION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510578640.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-22
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

Traditional clustering algorithms often face the problems of mis-aggregation and misaggregation when processing complex data, especially in scenarios where high-precision clustering is required, resulting in incorrect decision-making and analysis results.

Method used

The density clustering method based on SOM network is adopted, and by training feature data, acquiring topology map tables, performing density hierarchical clustering and cluster merging, the density hierarchical diffusion mechanism and topological relationship are used to merge similar clustering algorithms, making up for the shortcomings of traditional clustering algorithms that only utilize feature similarity.

Benefits of technology

High-precision clustering is achieved, the error rate and leakage rate are reduced, and the accuracy of clustering is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086624B_ABST
    Figure CN120086624B_ABST
Patent Text Reader

Abstract

The present invention provides a density clustering method and system based on a SOM network. The method includes obtaining training feature data, inputting the training feature data into a preset SOM network for training to obtain a trained SOM network; obtaining a feature set, inputting the feature set into the trained SOM network to obtain a topological mapping table; performing density hierarchical clustering processing on the feature set to obtain a number of clustering clusters; and performing cluster merging on the number of clustering clusters based on the topological mapping table to obtain final clustering clusters. The present invention can achieve high-precision clustering. The mechanism of density hierarchical diffusion makes the diffusion of clusters more strict, reduces the probability of features that seem similar but are not being included in the clusters, thereby reducing the misclustering rate. At the same time, it makes up for the deficiency of traditional clustering algorithms that only utilize a single piece of information, namely feature similarity, thereby effectively improving the missing clustering rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data clustering, and specifically relates to a density clustering method and system based on a SOM network. Background Art

[0002] In recent years, deep learning technology has experienced rapid development, and its engineering applications in various fields have become increasingly widespread, from image recognition, medical diagnosis to financial risk assessment, etc. Clustering plays an important role in these applications. For example, in image recognition, through clustering of image features, automatic classification and similarity retrieval of images can be achieved; in bioinformatics, clustering of gene expression data can help scientists discover functional modules of genes; in the financial field, clustering of transaction data can be used for risk assessment and fraud detection.

[0003] However, traditional clustering algorithms often face the problems of misclustering and missed clustering when dealing with complex data. Misclustering refers to wrongly classifying data points that do not belong to the same category into the same cluster, while missed clustering refers to wrongly classifying data points of the same category into multiple clusters. These problems may lead to incorrect decisions and analysis results in practical applications, especially in scenarios that require high-precision clustering, such as medical image analysis, financial risk assessment, etc. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention provides a density clustering method and system based on a SOM network to solve the technical problems in the prior art.

[0005] On the one hand, the present invention provides the following technical solution. A density clustering method based on a SOM network includes:

[0006] Obtain training feature data, and input the training feature data into a preset SOM network for training to obtain a trained SOM network;

[0007] Obtain a feature set, and input the feature set into the trained SOM network to obtain a topological mapping table;

[0008] Perform density hierarchical clustering processing on the feature set to obtain a number of clustering clusters;

[0009] Based on the topological mapping table, perform cluster merging on a number of the clustering clusters to obtain final clustering clusters.

[0010] Compared with the prior art, the beneficial effects of the present invention are as follows: First, the present invention obtains training feature data, inputs the training feature data into a preset SOM network for training to obtain a trained SOM network; then obtains a feature set, inputs the feature set into the trained SOM network to obtain a topological mapping table; then performs density hierarchical clustering processing on the feature set to obtain a number of clustering clusters; finally, based on the topological mapping table, performs cluster merging on the number of clustering clusters to obtain final clustering clusters. The present invention can achieve high-precision clustering through density hierarchical clustering. The mechanism of density hierarchical diffusion makes the diffusion of clusters more strict, reduces the probability of features that are similar but not identical from entering the cluster, thereby reducing the misclustering rate. Then, the topological relationship mapped by the SOM network is used to find and merge similar clusters. The introduction of topological information makes up for the deficiency of traditional clustering algorithms that only use a single piece of information, namely feature similarity, thus effectively improving the missed clustering rate.

[0011] Preferably, the step of inputting the feature set into the trained SOM network to obtain a topological mapping table is specifically as follows:

[0012] Input the feature set into the trained SOM network, and through the trained SOM network, map the feature set to a two-dimensional space to obtain the row index number and column index number of each data in the feature set in the SOM output layer. Based on the row index number and the column index number, determine the topological mapping table , where the topological mapping table is:

[0013] ;

[0014] In the formula, represents the row index number and column index number of the th feature data in the feature set.

[0015] Preferably, the steps of performing density hierarchical clustering processing on the feature set to obtain a number of clustering clusters include:

[0016] Calculate the similarity between each feature data in the feature set, determine the neighborhood density of each feature data according to the similarity, and use the feature data with the largest neighborhood density as the basic data point for diffusion;

[0017] If there is feature data in the feature set whose similarity with the basic data point is not less than the similarity threshold and the number of feature data whose similarity is not less than the similarity threshold is not less than the quantity threshold, then create a new clustering cluster and deposit the basic data point into the clustering cluster;

[0018] Using the basic data point Centered on a preset similarity Determine basic data points with a radius to define the neighborhood range, and store all the feature data within the neighborhood range of the basic data points in the neighborhood feature set;

[0019] Based on the basic data points and the neighborhood feature set, output several clustering clusters.

[0020] Preferably, the step of outputting several clustering clusters based on the basic data points and the neighborhood feature set includes:

[0021] Arbitrarily select a feature data from the neighborhood feature set , and calculate the matching degree between the basic data points and the feature data : :

[0022] ;

[0023] In the formula, , respectively represent the neighborhood ranges of the basic data points and the feature data ;

[0024] If the matching degree is greater than the preset matching degree, store the feature data in the clustering cluster where the basic data points are located;

[0025] If there are feature data within the neighborhood range of the feature data whose similarity to the feature data is not less than the similarity threshold and the number of feature data with a similarity not less than the similarity threshold is not less than the quantity threshold, store the feature data within the neighborhood range of the feature data in the neighborhood feature set;

[0026] Repeat the density hierarchical clustering process to output several clustering clusters.

[0027] Preferably, the step of performing cluster merging on several of the clustering clusters based on the topological mapping table to obtain the final clustering clusters includes:

[0028] Calculate the mean topological center of each clustering cluster based on the topological mapping table :

[0029] ; ;

[0030] In the formula, represents the number of feature data in the th clustering cluster, , respectively represent the row index number and column index number of the th feature data in the th clustering cluster;

[0031] Calculate the topological distance between any two of the clustering clusters :

[0032] ;

[0033] In the formula, , respectively represent the mean topological centers of the th clustering cluster and the th clustering cluster;

[0034] If the topological distance is less than the preset topological distance, then merge the corresponding two clustering clusters. If the topological distance is not less than the preset topological distance, then do not merge, so as to obtain several final clustering clusters.

[0035] In a second aspect, the present invention provides the following technical solution, a density clustering system based on a SOM network, the system includes:

[0036] A training module, configured to obtain training feature data, input the training feature data into a preset SOM network for training to obtain a trained SOM network;

[0037] A mapping module, configured to obtain a feature set, input the feature set into the trained SOM network to obtain a topological mapping table;

[0038] A clustering module, configured to perform density hierarchical clustering processing on the feature set to obtain several clustering clusters;

[0039] A merging module, configured to perform cluster merging on several of the clustering clusters based on the topological mapping table to obtain final clustering clusters.

[0040] Preferably, the mapping module is specifically configured to:

[0041] Input the feature set into the trained SOM network, map the feature set to a two-dimensional space through the trained SOM network to obtain the row index number and column index number of each data in the feature set in the SOM output layer, and determine the topological mapping table based on the row index number and the column index number , where the topological mapping table is:

[0042] ;

[0043] In the formula, represents the row index number and column index number of the th feature data in the feature set.

[0044] Preferably, the merging module includes:

[0045] A center determination sub-module for calculating the mean topological center of each of the clustering clusters based on the topological mapping table :

[0046] ; ;

[0047] In the formula, represents the number of feature data in the th clustering cluster, , respectively represent the row index number and column index number of the th feature data in the th clustering cluster;

[0048] A distance sub-module for calculating the topological distance between any two of the clustering clusters :

[0049] ;

[0050] In the formula, , respectively represent the mean topological centers of the th clustering cluster and the th clustering cluster;

[0051] A merging sub-module for merging the corresponding two clustering clusters if the topological distance is less than a preset topological distance, and not merging if the topological distance is not less than the preset topological distance, so as to obtain a number of final clustering clusters.

[0052] In a third aspect, the present invention provides the following technical solution. A computer includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the density clustering method based on the SOM network as described above is implemented.

[0053] Fourth aspect, the present invention provides the following technical solution: a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the density clustering method based on the SOM network as described above. Description of the Drawings

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0055] Figure 1 It is a flowchart of the density clustering method based on the SOM network provided in the first embodiment of the present invention;

[0056] Figure 2 It is a structural block diagram of the density clustering system based on the SOM network provided in the second embodiment of the present invention;

[0057] Figure 3 It is a schematic diagram of the hardware structure of a computer provided in another embodiment of the present invention.

[0058] The following will further illustrate the embodiments of the present invention with reference to the drawings. Detailed Embodiments

[0059] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are intended to explain the embodiments of the present invention and should not be construed as a limitation to the present invention.

[0060] In the description of the embodiments of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the embodiments of the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.

[0061] In addition, the terms "first" and "second" are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, the meaning of "a plurality" is two or more, unless otherwise specifically defined.

[0062] In the embodiments of the present invention, unless otherwise clearly stipulated and defined, terms such as "installed", "connected", "linked", "fixed", etc. shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present invention can be understood according to specific circumstances.

[0063] Embodiment 1

[0064] In Embodiment 1 of the present invention, as Figure 1 shown, a density clustering method based on a SOM network includes:

[0065] S1. Obtain training feature data, input the training feature data into a preset SOM network for training to obtain a trained SOM network;

[0066] Specifically, the input layer in the preset SOM network here can be determined according to the dimension of the feature data, and the output layer can be determined according to the category scale of the feature data. By randomly initializing the parameters of the preset SOM network, inputting the training feature data and adopting a competitive learning mechanism for training, the SOM network is trained.

[0067] S2. Obtain a feature set, input the feature set into the trained SOM network to obtain a topological mapping table;

[0068] Specifically, step S2 is specifically:

[0069] Input the feature set into the trained SOM network, map the feature set to a two-dimensional space through the trained SOM network to obtain the row index number and column index number of each data in the feature set in the SOM output layer, and determine the topological mapping table based on the row index number and the column index number , where the topological mapping table is:

[0070] ;

[0071] In the formula, Indicates the row index number and column index number of the th feature data in the feature set;

[0072] Specifically, the SOM network maps the input high-dimensional features to a two-dimensional space while maintaining the topological relationship between the input features, that is, features with close topological relationships in the input are also close in the output two-dimensional mapping Map. For the topological mapping table, it records the row and column index numbers of the BMU of each feature data in the SOM output layer.

[0073] S3. Perform density hierarchical clustering on the feature set to obtain several clustering clusters;

[0074] Among them, the step S3 includes:

[0075] S31. Calculate the similarity between each feature data in the feature set, determine the neighborhood density of each feature data according to the similarity, and use the feature data with the largest neighborhood density as the basic data point for diffusion;

[0076] Specifically, the similarity here is determined by the distance between two feature data, and the distance is inversely proportional to the similarity. The neighborhood density specifically refers to the number of feature data in the feature set whose similarity with this data point is not less than the similarity threshold.

[0077] S32. If there is feature data in the feature set whose similarity with the basic data point is not less than the similarity threshold and the number of feature data whose similarity is not less than the similarity threshold is not less than the quantity threshold, then create a new clustering cluster and store the basic data point in the clustering cluster;

[0078] The similarity threshold here is specifically the preset similarity , and the quantity threshold can be determined according to the data type and data scale.

[0079] S33. Determine the neighborhood range of the basic data point with the preset similarity as the radius, and store all the feature data within the neighborhood range of the basic data point in the neighborhood feature set; Specifically, the basic data point is not included in the neighborhood feature set here.

[0080] S34. Output several clustering clusters based on the basic data point

[0081] and the neighborhood feature set; ​

[0082] Among them, the step S34 includes:

[0083] S341. Arbitrarily select a feature data from the neighborhood feature set , and calculate the matching degree between the basic data point and the feature data : :

[0084] ;

[0085] In the formula, , respectively represent the neighborhood ranges of the basic data point , the feature data ;

[0086] S342. If the matching degree is greater than the preset matching degree, then store the feature data into the clustering cluster where the basic data point is located;

[0087] Specifically, the range of the preset matching degree here is between 0 and 1. When the preset matching degree is 0, it cannot effectively improve the mis-clustering problem. When the preset matching degree is 1, serious under-clustering problems will occur. Therefore, considering the improvement effects of mis-clustering and under-clustering comprehensively, the range of the preset matching degree is set to 0.38 - 0.62.

[0088] S343. If there are feature data with a similarity not less than the similarity threshold between the feature data in the neighborhood range of the feature data and the number of feature data with a similarity not less than the similarity threshold is not less than the number threshold, then store the feature data in the neighborhood range of the feature data into the neighborhood feature set;

[0089] Specifically, in the subsequent steps, taking the feature data as the basic data point and repeating the above process, a complete clustering cluster can be obtained. And for the neighborhood feature set, for each feature data in it, it needs to be marked with the corresponding identifier to indicate its source. If a neighborhood feature data is determined by the basic data point , then it needs to be marked with . If a neighborhood feature data is determined by the feature data , then it needs to be marked with .

[0090] S344. Repeatedly execute the density hierarchical clustering process to output several clustering clusters;

[0091] Specifically, first in the present application, the neighborhood density of each feature data in the feature set is identified, and the feature data with the largest neighborhood density is used as the basic data point. Diffusion is carried out. After a complete clustering cluster is determined, all the feature data except this clustering cluster in the feature set are formed into a set. Then, the point with the largest neighborhood density in this set is used as the basic data point, and the above steps are repeatedly executed to obtain several clustering clusters.

[0092] S4. Cluster merging is performed on several of the clustering clusters based on the topological mapping table to obtain the final clustering cluster.

[0093] Among them, the step S4 includes:

[0094] S41. Calculate the mean topological center of each of the clustering clusters based on the topological mapping table :

[0095] ; ;

[0096] In the formula, represents the number of feature data in the th clustering cluster, , respectively represent the row index number and column index number of the th feature data in the th clustering cluster.

[0097] S42. Calculate the topological distance between any two of the clustering clusters :

[0098] ;

[0099] In the formula, , respectively represent the mean topological centers of the th clustering cluster and the th clustering cluster;

[0100] S43. If the topological distance is less than the preset topological distance, the corresponding two clustering clusters are merged. If the topological distance is not less than the preset topological distance, they are not merged to obtain several final clustering clusters.

[0101] The density clustering method based on SOM network provided in the first embodiment of the present invention first obtains training feature data, inputs the training feature data into a preset SOM network for training to obtain a trained SOM network; then obtains a feature set, inputs the feature set into the trained SOM network to obtain a topological mapping table; then performs density hierarchical clustering processing on the feature set to obtain a number of clustering clusters; finally, performs cluster merging on the number of clustering clusters based on the topological mapping table to obtain the final clustering clusters. The present invention can achieve high-precision clustering through density hierarchical clustering. The mechanism of density hierarchical diffusion makes the diffusion of clusters more strict, reduces the probability of features that are similar but not exactly the same from entering the cluster, thereby reducing the misclustering rate. Then, the topological relationship mapped by the SOM network is used to find and merge similar clusters. The introduction of topological information makes up for the deficiency of traditional clustering algorithms that only use a single piece of information, namely feature similarity, thus effectively improving the missed clustering rate.

[0102] Embodiment 2

[0103] As Figure 2 shown, the second embodiment of the present invention provides a density clustering system based on SOM network. The system includes:

[0104] A training module 1, configured to obtain training feature data, input the training feature data into a preset SOM network for training to obtain a trained SOM network;

[0105] A mapping module 2, configured to obtain a feature set, input the feature set into the trained SOM network to obtain a topological mapping table;

[0106] A clustering module 3, configured to perform density hierarchical clustering processing on the feature set to obtain a number of clustering clusters;

[0107] A merging module 4, configured to perform cluster merging on a number of the clustering clusters based on the topological mapping table to obtain the final clustering clusters.

[0108] The mapping module 2 is specifically configured to:

[0109] Input the feature set into the trained SOM network, map the feature set to a two-dimensional space through the trained SOM network to obtain the row index number and column index number of each data in the feature set in the SOM output layer, and determine the topological mapping table based on the row index number and the column index number , where the topological mapping table is:

[0110] ;

[0111] In the formula, represents the The row index number and column index number of the feature data.

[0112] The clustering module 3 includes:

[0113] An identification sub-module, configured to calculate the similarity between each feature data in the feature set, determine the neighborhood density of each feature data according to the similarity, and use the feature data with the largest neighborhood density as the basic data point Perform diffusion;

[0114] A similarity sub-module, configured to if there is feature data in the feature set whose similarity with the basic data point is not less than the similarity threshold and the number of feature data whose similarity is not less than the similarity threshold is not less than the quantity threshold, then create a new clustering cluster and store the basic data point into the clustering cluster;

[0115] A neighborhood sub-module, configured to determine the neighborhood range of the basic data point centered on the basic data point and with a preset similarity as the radius, and store all the feature data within the neighborhood range of the basic data point into the neighborhood feature set;

[0116] An output sub-module, configured to output several clustering clusters based on the basic data point and the neighborhood feature set.

[0117] The output sub-module includes:

[0118] A matching unit, configured to arbitrarily select a feature data in the neighborhood feature set and calculate the matching degree between the basic data point and the feature data : :

[0119] ;

[0120] In the formula, , respectively represent the neighborhood ranges of the basic data point and the feature data ;

[0121] A storing unit, configured to if the matching degree is greater than the preset matching degree, then store the feature data into the clustering cluster where the basic data point is located;

[0122] A neighborhood feature expansion unit, configured to if the feature data Within the neighborhood range of, there is characteristic data that has a similarity with the said characteristic data and the number of characteristic data with a similarity not less than the similarity threshold is not less than the quantity threshold, then the said characteristic data within the neighborhood range of will be stored into the said neighborhood characteristic set;

[0123] An output unit, configured to repeatedly execute the density hierarchical clustering process to output several clustering clusters.

[0124] The said merging module 4 includes:

[0125] A center determination sub-module, configured to calculate the mean topological center of each of the said clustering clusters based on the said topological mapping table :

[0126] ; ;

[0127] In the formula, represents the number of characteristic data in the th clustering cluster, , respectively represent the row index number and column index number of the th characteristic data in the th clustering cluster;

[0128] A distance sub-module, configured to calculate the topological distance between any two of the said clustering clusters :

[0129] ;

[0130] In the formula, , respectively represent the mean topological centers of the th clustering cluster and the th clustering cluster;

[0131] A merging sub-module, configured to, if the topological distance is less than the preset topological distance, merge the corresponding two clustering clusters, and if the topological distance is not less than the preset topological distance, do not merge, so as to obtain several final clustering clusters.

[0132] In some other embodiments of the present invention, the present invention embodiment provides the following technical solution, a computer, including a memory 102, a processor 101, and a computer program stored on the said memory 102 and executable on the said processor 101, and when the processor 101 executes the said computer program, it implements the density clustering method based on the SOM network as described above.

[0133] Specifically, the above-mentioned processor 101 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present invention.

[0134] Among them, the memory 102 may include a mass storage for data or instructions. By way of example and not limitation, the memory 102 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In appropriate cases, the memory 102 may include removable or non-removable (or fixed) media. In appropriate cases, the memory 102 may be internal or external to the data processing device. In a particular embodiment, the memory 102 is non-volatile memory. In a particular embodiment, the memory 102 includes a read-only memory (ROM) and a random access memory (RAM). In appropriate cases, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. In appropriate cases, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended data out dynamic random access memory (EDODRAM), a synchronous dynamic random-access memory (SDRAM), etc.

[0135] The memory 102 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 101.

[0136] The processor 101 reads and executes the computer program instructions stored in the memory 102 to implement the above-mentioned density clustering method based on the SOM network.

[0137] In some embodiments, the computer may further include a communication interface 103 and a bus 100. Among them, as Figure 3 shown, the processor 101, the memory 102, and the communication interface 103 are connected through the bus 100 to complete communication with each other.

[0138] The communication interface 103 is used to implement communication between various modules, devices, units, and / or devices in the embodiments of the present invention. The communication interface 103 can also implement data communication with other components such as external devices, image / data acquisition devices, databases, external storage, and image / data processing workstations.

[0139] Bus 100 includes hardware, software, or both, and couples components of a computer device to each other. Bus 100 includes, but is not limited to, at least one of the following: Data Bus, Address Bus, Control Bus, Expansion Bus, Local Bus. By way of example and not limitation, Bus 100 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable bus or a combination of two or more of these. In suitable cases, Bus 100 may include one or more buses. Although embodiments of the present invention describe and illustrate specific buses, the present invention contemplates any suitable bus or interconnect.

[0140] The computer can execute the density clustering method based on the SOM network of the present invention based on the obtained density clustering system based on the SOM network, so as to realize density clustering based on the SOM network.

[0141] In still some other embodiments of the present invention, in combination with the above-mentioned density clustering method based on the SOM network, embodiments of the present invention provide the following technical solution: a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned density clustering method based on the SOM network is realized.

[0142] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0143] More specific examples (a non-exhaustive list) of the readable medium include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0144] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0145] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity in description, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0146] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. A density clustering method based on SOM network, characterized in that, Including: Obtain training feature data, and input the training feature data into a preset SOM network for training to obtain a trained SOM network; Obtain a feature set, and input the feature set into the trained SOM network to obtain a topological mapping table; Perform density hierarchical clustering processing on the feature set to obtain a number of clustering clusters; Perform cluster merging on a number of the clustering clusters based on the topological mapping table to obtain final clustering clusters; The step of inputting the feature set into the trained SOM network to obtain a topological mapping table is specifically: Input the feature set into the trained SOM network, and map the feature set to a two-dimensional space through the trained SOM network to obtain the row index number and column index number of each data in the feature set in the SOM output layer, and determine the topological mapping table based on the row index number and the column index number , where the topological mapping table is as follows: ; In the formula, represents the row index number and column index number of the th feature data in the feature set; The step of performing density hierarchical clustering processing on the feature set to obtain a number of clustering clusters includes: Calculate the similarity between each piece of feature data in the feature set, determine the neighborhood density of each piece of feature data according to the similarity, and use the feature data with the maximum neighborhood density as the basic data point Perform diffusion; If there is feature data within the feature set whose similarity to the basic data point is not less than the similarity threshold and the number of feature data with a similarity not less than the similarity threshold is not less than the quantity threshold, then a new cluster is created and the basic data point is stored in the cluster; Taking the said basic data point as the center and a preset similarity as the radius to determine the neighborhood range of the basic data point , and storing all the feature data within the neighborhood range of the said basic data point into the neighborhood feature set; Based on the basic data points Output several clustering clusters with the neighborhood feature set; Wherein, the neighborhood density is the number of feature data in the feature set whose similarity with this data point is not less than the similarity threshold; The step of performing cluster merging on a number of the clustering clusters based on the topological mapping table to obtain final clustering clusters includes: Calculate the mean topological center of each of the clustering clusters based on the topological mapping table : ; ; In the formula, represents the number of feature data in the th clustering cluster, , respectively represent the row index number and column index number of the th feature data in the th clustering cluster; Calculate the topological distance between any two of the said clustering clusters : ; In the formula, , respectively represent the mean topological centers of the th clustering cluster and the th clustering cluster; If the topological distance is less than the preset topological distance, then the corresponding two clustering clusters are merged. If the topological distance is not less than the preset topological distance, then they are not merged, so as to obtain several final clustering clusters.

2. The density clustering method based on SOM network according to claim 1, characterized in that, Based on the basic data points The step of outputting a plurality of clustering clusters from the neighborhood feature set includes: Arbitrarily select a piece of feature data from the neighborhood feature set , and calculate the matching degree of the basic data point and the feature data : : ; In the formula, and respectively represent the neighborhood ranges of the basic data points and the characteristic data ; If the matching degree is greater than the preset matching degree, then the feature data is stored in the clustering cluster where the basic data point is located; If there is feature data within the neighborhood range of the said feature data and the similarity between the said feature data is not less than the similarity threshold, and the number of feature data with similarity not less than the similarity threshold is not less than the quantity threshold, then the feature data within the neighborhood range of the said feature data is stored into the said neighborhood feature set; Repeatedly execute the density hierarchical clustering process to output a number of clustering clusters.

3. A density clustering system based on SOM network, the system adopts the density clustering method based on SOM network as described in claim 1, characterized in that, The system includes: A training module, configured to obtain training feature data, and input the training feature data into a preset SOM network for training to obtain a trained SOM network; A mapping module, configured to obtain a feature set, and input the feature set into the trained SOM network to obtain a topological mapping table; A clustering module, configured to perform density hierarchical clustering processing on the feature set to obtain a number of clustering clusters; A merging module, configured to perform cluster merging on a number of the clustering clusters based on the topological mapping table to obtain final clustering clusters.

4. A computer, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the density clustering method based on the SOM network according to any one of claims 1 to 2.

5. A storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by the processor, it implements the density clustering method based on the SOM network according to any one of claims 1 to 2.

Citation Information

Patent Citations

  • Path planning method and system based on scene classification

    CN107609709A

  • Network flow classification method based on SOM and K-means fusion algorithm

    CN111211994A