Information processing system, information processing method, and program product
Through the improved hierarchical clustering method, the data density is calculated using Gini coefficients and the vectors with high density are selected for clustering, which solves the problem of excessive calculation load of high-dimensional data, and reduces the calculation amount and shortens the learning time.
Patent Information
- Application Number
- CN202510123385.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-06
- Filing Date
- 2025-01-26
- Publication Date
- 2025-08-08
AI Technical Summary
When existing information processing systems perform vector data clustering, the calculation load is too high, especially in high-dimensional data, which is prone to dimensional disasters, resulting in an explosion in computational volume.
The improved hierarchical clustering method is used to determine the data density by calculating the Gini coefficient of the vector, select vectors with high density for clustering, and stop the calculation when the target cluster number is reached, avoiding repeated iterations.
It effectively reduces the computational load and shortens the creation time of learning models, especially in the learning process of large language models, which significantly reduces the computational amount.
Smart Images

Figure CN120448843A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing system, an information processing method, and a program. Background Art
[0002] An information processing system is known that includes a clustering unit that converts input data into multidimensional vector data and groups data with close relative distances based on the relative distances of the converted vector data (see, for example, Patent Document 1).
[0003] Patent Document 1: Japanese Patent No. 6562984 Summary of the Invention
[0004] In order to avoid the huge combinatorial explosion when calculating the relative distance between vector data, the above system uses the K-nearest neighbor algorithm to group vectors with similar distances.
[0005] However, when grouping is performed after specifying the number of clusters in advance as in the K-nearest neighbor algorithm, there may be a problem that the computational load becomes excessive due to repeated iterative calculations.
[0006] The present disclosure has been made to solve such problems, and its main object is to provide an information processing system, an information processing method, and a program that can reduce the calculation load.
[0007] One method of the present disclosure for achieving the above-mentioned purpose is an information processing system, which includes a data acquisition unit and a clustering unit, wherein the data acquisition unit acquires multidimensional vector data, and the clustering unit performs the following hierarchical clustering, that is, based on the relative distance of the vector data acquired by the data acquisition unit, the data with similar relative distances are grouped with each other, and the grouping is repeated, wherein the information processing system includes a density calculation unit, and the density calculation unit calculates the density of each cluster of the vector data separately, and when performing the hierarchical clustering, the clustering unit divides or integrates the clusters based on the density of each cluster calculated by the density calculation unit.
[0008] In this method, the following method can also be adopted, that is, the clustering unit extracts at least two vectors by comparing the density of each vector component of the multidimensional vector data, and extracts the dense part where the vector values are dense in each extracted vector, and extracts the focus area for clustering based on the extracted dense part, and performs the hierarchical clustering on the extracted focus area and the surrounding area surrounding the focus area.
[0009] In this method, the following method can also be adopted, that is, the clustering unit divides the clusters when it is judged that the density of the cluster calculated by the density calculation unit is above the threshold, and when it is judged that the density of each cluster is not above the threshold, it is judged whether the current number of clusters is above the target value, and when it is judged that the current number of clusters is above the target value, the hierarchical clustering is stopped.
[0010] In this aspect, when the clustering unit determines that the density of each cluster has reached a predetermined value during the hierarchical clustering, it may temporarily stop the hierarchical clustering and output the current number of clusters and the vectors included in each cluster.
[0011] In this aspect, an aspect may be adopted in which the density is a Gini coefficient.
[0012] One method of the present disclosure for achieving the above-mentioned purpose is an information processing method, comprising: a step of obtaining multidimensional vector data; a step of performing the following hierarchical clustering, that is, based on the relative distance of the obtained vector data, grouping the data with similar relative distances to each other, and repeating the grouping; a step of calculating the density of each cluster of the vector data respectively; and a step of dividing or integrating the clusters based on the calculated density of each cluster when performing the hierarchical clustering.
[0013] One method of the present disclosure for achieving the above-mentioned purpose is a program that enables a computer to perform the following processing, namely: obtaining multidimensional vector data; performing the following hierarchical clustering processing, namely, grouping data with similar relative distances based on the relative distances of the obtained vector data, and repeating the grouping; calculating the density of each cluster of the vector data respectively; and when performing the hierarchical clustering, dividing or integrating the clusters based on the calculated density of each cluster.
[0014] According to the present disclosure, an information processing system, an information processing method, and a program capable of reducing calculation load can be provided.
[0015] The above and other objects, features and advantages of the present disclosure will be more fully understood from the detailed description and accompanying drawings given below. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a block diagram showing a schematic system configuration of a processing system according to this embodiment.
[0017] Figure 2 This is a flowchart showing an example of the processing flow of the processing system according to this embodiment.
[0018] Figure 3 This figure compares the advantages and disadvantages of the top-down method, the bottom-up method, and the method according to this embodiment.
[0019] Figure 4 This is a block diagram showing an example of a schematic system configuration of an information processing system according to the present embodiment.
[0020] Figure 5 This is a flowchart showing an example of the flow of the information processing method of the information processing system according to the present embodiment.
[0021] Figure 6 This figure shows an example of the Gini coefficient of each component of a high-dimensional vector after projecting each morpheme onto the high-dimensional vector.
[0022] Figure 7 is an example of a graph showing the simultaneous distribution of vector α and vector β.
[0023] Figure 8 A diagram showing an example of a densely populated portion of a vector.
[0024] Figure 9 FIG. 1 is a diagram showing an example of a target area and surrounding areas.
[0025] Figure 10 A diagram showing cluster demarcation lines. DETAILED DESCRIPTION
[0026] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings.
[0027] For example, when creating a document retrieval system or a learned model of a large language model used for document generation, so-called clustering is implemented to convert input data into multidimensional vector data and group relatively close data based on the relative distance of the converted vector data.
[0028] In order to avoid the explosive increase in combinations when calculating the relative distance between vector data, it is assumed that vectors with close distances are grouped using the K-nearest neighbor algorithm.
[0029] However, when using so-called unsupervised clustering such as the K-nearest neighbor algorithm, in which the number of clusters (groups) is specified in advance and grouped, repeated iterative calculations may cause a problem of increasing the amount of calculation and excessive computational load.
[0030] The information processing system according to the present embodiment executes a processing method that solves the problems that may occur in the above-mentioned K-means algorithm.
[0031] For example, in unsupervised learning clustering, which performs clustering on large, high-dimensional data without pre-determining cluster centers, the number of conceivable dimensions becomes enormous. Consequently, the so-called curse of dimensionality can occur, making clustering impossible. To address the curse of dimensionality in this clustering, there are two methods:
[0032] (1) If the required number of clusters is reached, the calculation is stopped.
[0033] (2) Use a method that does not rely on initial values to avoid iterative calculations.
[0034] The information processing system according to this embodiment takes into account the methods (1) and (2) above and implements the following improved hierarchical clustering as a method for avoiding the curse of dimensionality. In this embodiment, while stopping when the desired number of clusters is reached, the system considers "which clusters to focus on and how to form the desired cluster aggregation," thus focusing on the density of the data.
[0035] The information processing system according to this embodiment calculates the Gini coefficient representing the density of each high-dimensional vector (for example, 200 dimensions in the case of word2vec and 768 dimensions in the case of BERT).
[0036] When the function of the vector value is set to L(x), the Gini coefficient is expressed as 1-2∫ 1 0L(x)dx. The Gini coefficient is closer to 0 as the data is more uniform, and closer to 1 as the data is more dense.
[0037] For example, count the number of words in group 1, the number of words in group 2, the number of words in group 3, ..., the number of words in group N contained in a certain range, calculate the ratio of each to the total number of words, find the value obtained by multiplying the ratios (= sum of products), and subtract it from 1 to obtain the Gini coefficient.
[0038] While methods for determining data density include methods using the centroid, the method using the Gini coefficient, as used in this embodiment, is more advantageous in terms of ease of calculation. The Gini coefficient method simply counts the number of data items contained in a group and the number of group types, requiring only four simple arithmetic operations (counting the number of items, calculating the ratio, summing the products of the ratios, and summing the products of 1-ratio), making calculations easy.
[0039] The information processing system involved in this embodiment selects multiple vectors with high Gini coefficients and performs hierarchical clustering based on the values of these vectors. Hierarchical clustering is a method that groups data with high similarity and close distances between vectors, and then repeatedly grouping these data into larger groups. This method allows for intuitive understanding of the structure of data and hidden relationships.
[0040] The information processing system of this embodiment can suppress increases in computational load, for example, significantly reducing the time required to create a learned model. This time reduction helps address the issue of reducing the time required to create a learned model for a large language model.
[0041] The fundamental challenge of this embodiment is to obtain the required features for data with hundreds of millions of dimensions, known as big data, without exploding the computational complexity. Furthermore, it is necessary to appropriately compress high-dimensional data while retaining the features.
[0042] The information processing system according to this embodiment uses the K-means algorithm or the center of gravity to avoid the exponential increase in the amount of calculation (o(N 2 )) explosion and by stopping the calculation midway, the amount of calculation is reduced linearly (=o(1 / N)). This makes it possible to obtain the required feature value from large data.
[0043] In particular, in language data, for example, a minimum corpus consists of 100,000 articles and 71,000 words, and the dimension of a vector for a single word can be as large as several hundred. Therefore, obtaining feature quantities while minimizing computational effort may be essential for document retrieval or summarization.
[0044] On the other hand, if big data is densely populated across all parts of the data, it becomes difficult to reduce the amount of computation required to observe trends. However, if there is a bias in the data distribution, such as sparseness or density within the data as a whole, it can be effective to prioritize computations in the denser parts and more easily complete computations in the sparser parts.
[0045] Therefore, as described above, the information processing system according to this embodiment focuses on data density and calculates the Gini coefficient representing the density of vectors, selects multiple vectors with high Gini coefficients, and performs hierarchical clustering based on the values of these vectors.
[0046] Next, an example of the hardware configuration of the processing system according to this embodiment will be described. Figure 11 is a block diagram showing a schematic system configuration of a processing system according to the present embodiment. A processing system 1 according to the present embodiment includes a morphological analysis unit 2, a relative distance calculation unit 3, and the aforementioned information processing system 4.
[0047] In addition, the information processing system 4 has, for example, the hardware structure of a conventional computer, which includes processors 4a such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit), internal memories 4b such as a RAM (Random Access Memory) and a ROM (Read Only Memory), storage devices 4c such as an HDD (Hard Disk Drive) and an SDD (Solid State Drive), input and output I / Fs 4d for connecting to peripheral devices such as a display, and communication I / Fs 4e for communicating with devices outside the device.
[0048] Figure 2 The flowchart is an example of a processing flow of the processing system involved in this embodiment. In addition, the processing system 1 involved in this embodiment is, for example, a system for finding common factors that replace the meaning of text from text, and can be applied to other natural language processing services (retrieval, document summarization, translation, etc.).
[0049] In order to find common factors in texts, it is necessary to determine the similarity of the texts, and therefore the words contained in the texts are digitized. To digitize, the processing system 1 according to this embodiment executes the following steps S101 to S103.
[0050] First, for example, text data is input to the morphological analysis unit 2 (step S101 ). The morphological analysis unit 2 performs morphological analysis to output units of words or morphemes based on the input text data (step S102 ).
[0051] Based on the text decomposed into morpheme units by the morphological analysis unit 2, the relative distance calculation unit 3 maps words and morphemes to relative positions based on usage in each language and converts them into numerical values. Specifically, the relative distance calculation unit 3 generates high-dimensional vector data by projecting the data onto a high-dimensional vector (e.g., several hundred dimensions) (step S103). The relative distance calculation unit 3 outputs this high-dimensional vector data to the information processing system 4.
[0052] The information processing system 4 performs improved hierarchical clustering (step S104 ) described later based on the high-dimensional vector output from the relative distance calculation unit 3 .
[0053] Since the common factor search method based on clustering is well known, a detailed description is omitted here. For example, as clustering methods, there are top-down methods (K-center clustering) that divide the data as a whole, and bottom-up methods (hierarchical clustering) that group the two closest data points and repeat the grouping.
[0054] Figure 3 A diagram comparing the advantages and disadvantages of top-down methods, bottom-up methods, and improved hierarchical clustering.
[0055] The information processing system 4 according to the present embodiment takes the above advantages and disadvantages into consideration and performs improved hierarchical clustering that eliminates the disadvantages of the top-down method.
[0056] Figure 4 This is a block diagram showing an example of a schematic system configuration of an information processing system according to the present embodiment. The information processing system 4 according to the present embodiment includes a data acquisition unit 41 that acquires high-dimensional vector data, a density calculation unit 42 that calculates the density of each cluster of the high-dimensional vectors acquired by the data acquisition unit 41, and a clustering unit 43 that performs hierarchical clustering based on the density of each cluster calculated by the density calculation unit 42.
[0057] Figure 5 This is a flowchart showing an example of the flow of the information processing method of the information processing system according to the present embodiment.
[0058] The data acquisition unit 41 acquires high-dimensional vector data from, for example, the relative distance calculation unit 3 (step S201 ).
[0059] The density calculation unit 42 calculates, for example, the Gini coefficient of each cluster as the density of each cluster of the high-dimensional vector acquired by the data acquisition unit 41 (step S202 ).
[0060] The clustering unit 43 determines whether the Gini coefficient of each cluster calculated by the density calculation unit 42 is equal to or greater than a threshold value (step S203). The threshold value of the Gini coefficient may be set in advance in the clustering unit 43.
[0061] When the clustering unit 43 determines that the Gini coefficient of the cluster calculated by the density calculation unit 42 is equal to or greater than the threshold value (Yes in step S203 ), it divides the cluster (step S204 ) and returns to step S202 .
[0062] On the other hand, if the clustering unit 43 determines that the Gini coefficient of each cluster calculated by the density calculation unit 42 is not greater than the threshold value (No in step S203), it determines whether the current number of clusters is greater than the target value (step S205). The target value of the number of clusters (number of common factors) may also be set in advance in the clustering unit 43.
[0063] If the clustering unit 43 determines that the current number of clusters is greater than the target value (Yes in step S205), the clustering unit 43 stops hierarchical clustering (step S206). On the other hand, if the clustering unit 43 determines that the current number of clusters is not greater than the target value (No in step S205), the process returns to step S204.
[0064] As described above, the information processing system 4 according to the present embodiment calculates the density of each cluster of high-dimensional vectors and repeatedly divides clusters having high density and dense data until the target number of clusters is reached.
[0065] The improved hierarchical clustering method of this embodiment can pre-set the target number of clusters and perform calculations while avoiding the drawbacks of dependency on initial values and increased computational complexity. For example, words or morphemes can be clustered until the desired number of clusters is reached, and then aggregated into common factors.
[0066] Next, a more detailed description will be given of the improved hierarchical clustering of the clustering unit 43. The clustering unit 43 according to this embodiment is characterized in that it can significantly reduce the computational load for large data without sacrificing accuracy.
[0067] As described above, calculating the distance between high-dimensional vectors without requiring any computational techniques requires comparing the vector dimensions x the number of words. However, the "close distances between words" required to create clusters, i.e., data density, does not necessarily occur in all dimensions. Therefore, the clustering unit 43 of this embodiment calculates the Gini coefficient, which represents the density of the data, for each dimension (each vector component) of the high-dimensional vector.
[0068] Vector components with high Gini coefficients have dense data. Therefore, they require further clustering to further refine the common factors. On the other hand, vector components with low Gini coefficients have evenly distributed data. Therefore, the common factors can be divided without further clustering.
[0069] Based on the above, the clustering unit 43 according to this embodiment concentrates clusters implemented by the improved hierarchical clustering described above on areas where data is particularly dense. Next, a method for determining concentrated areas by the improved hierarchical clustering described above will be described in detail.
[0070] The clustering unit 43 first calculates the Gini coefficient of each vector component based on the high-dimensional vector acquired by the data acquisition unit 41 . Figure 6 Graph showing an example of the Gini coefficient of each component of a high-dimensional vector after projecting each morpheme onto the high-dimensional vector.
[0071] For example, Figure 6 As shown, the clustering unit 43 compares the calculated Gini coefficients of the components of the high-dimensional vectors and extracts two vectors: vector α consisting of vector components with the highest Gini coefficient and vector β consisting of vector components with the second highest Gini coefficient.
[0072] Figure 7 is an example of a graph showing the simultaneous distribution of vector α and vector β. Figure 7 In the simultaneous distribution shown, the darker colored portion of the graph is where the vector values are dense, and the lighter colored portion is where the vector values are sparse.
[0073] like Figure 7 As shown, the clustering unit 43 obtains a portion where vector values are densely packed and a portion where vector values are sparsely packed, and then determines a concentrated portion for improved hierarchical clustering based on the obtained results.
[0074] Figure 8 This is a diagram showing an example of a densely populated portion of a vector. Figure 8 As shown, the portion where the vector values are particularly dense is i in the component of vector α and ii in the component of vector β.
[0075] Here, clustering unit 43 extracts i and ii so that the common portion of the vectors i and ii, where the vector values are particularly dense, reaches a predetermined value. Specifically, the common portion of i and ii is defined as a region of interest, and the region surrounding the region of interest is defined as a peripheral region. Clustering unit 43 extracts i and ii so that the peripheral region accounts for a predetermined ratio (e.g., 45%) of the total region. The predetermined ratio can also be set experimentally, for example, taking into account computational complexity. Figure 9 FIG. 1 is a diagram showing an example of the target region and the surrounding region calculated as described above.
[0076] In addition, although the clustering unit 43 extracts two vectors in the above description, it may extract three or more vectors and extract the target region and the surrounding region based on the extracted vectors.
[0077] The clustering unit 43 performs the above-mentioned clustering on the target area and the surrounding area extracted as described above. Figure 5 Improved hierarchical clustering is shown. Hierarchical clustering is performed not only on the region of interest but also on the surrounding areas because it is difficult to properly cluster the boundary portion of the region of interest. Therefore, the clustering unit 43 involved in this embodiment performs improved hierarchical clustering not only on the region of interest but also on the surrounding areas of the region of interest.
[0078] Furthermore, the clustering unit 43 may temporarily stop the improved hierarchical clustering when it determines that the Gini coefficient of each cluster has reached a predetermined value (for example, when it determines that the amount of data has decreased by approximately 20%). For example, when the clustering unit 43 temporarily stops the hierarchical clustering, it may randomly select a representative vector from each cluster.
[0079] The clustering unit 43 assigns a name to each vector included in the cluster. This allows the clustering unit 43 to output the number of clusters and the vectors included in the cluster (e.g., their names, morphemes A, etc.). The clustering unit 43 can output this information using, for example, a display device or printer. The user can adjust the threshold and target value of the Gini coefficient, described later, based on the output from the clustering unit 43.
[0080] Thus, for example, Figure 10 As shown in FIG, by adjusting the cluster division line according to the density of the data, the influence of density fluctuation can be minimized and the dense part can be divided more focused.
[0081] The clustering unit 43 restarts the improved hierarchical clustering after the temporary stop, and finally stops the improved hierarchical clustering when it determines that the number of clusters has reached the target value.
[0082] Furthermore, the target number of clusters can be determined based on, for example, the number of meanings that can be formed by a word. This pseudo-indicates the number of different meanings of a word. Therefore, when categorizing more specific meanings, the target number of clusters can be increased. When categorizing meanings broadly and using a wide range of synonyms, the target number of clusters can be decreased. In this embodiment, for example, 100,000 meanings are set, and the target value is set to 100,000.
[0083] As described above, the information processing system 4 according to this embodiment includes: a data acquisition unit 41 that acquires multidimensional vector data; a clustering unit 43 that performs hierarchical clustering by grouping data with similar relative distances based on the relative distances of the vector data acquired by the data acquisition unit 41, and repeating the grouping; and a density calculation unit 42 that calculates the density of each cluster of the vector data. When performing hierarchical clustering, the clustering unit 43 divides the vector data into clusters whose density exceeds a threshold value based on the density of each cluster calculated by the density calculation unit 42.
[0084] According to the information processing system 4 of this embodiment, the density of each cluster of a high-dimensional vector is calculated, and clusters with high density and dense data are divided. This process is repeated until the target number of clusters is reached. This avoids the disadvantages of dependence on initial values and increased computational complexity, and reduces the computational load while enabling calculations to be performed with the target number of clusters set in advance.
[0085] While several embodiments of the present disclosure have been described, these embodiments are provided as examples and are not intended to limit the scope of the invention. These new embodiments can be implemented in various other ways, and various omissions, substitutions, and modifications can be implemented without departing from the scope of the invention. These embodiments and their variations are included within the scope and spirit of the invention and are included within the scope of the invention described in the technical solution and its equivalents.
[0086] Furthermore, in the above embodiment, the clustering unit 43, when performing hierarchical clustering, divides clusters whose density is higher than a threshold and whose vector data is dense based on the density of each cluster calculated by the density calculation unit 42. However, the present invention is not limited to this. For example, the clustering unit 43 may also, when performing hierarchical clustering, integrate clusters whose density is lower than a threshold and whose vector data is sparse based on the density of each cluster calculated by the density calculation unit 42.
[0087] As described above, the density of each cluster of a high-dimensional vector is calculated, and clusters with low density and sparse data are integrated. This process is repeated until the target number of clusters is reached. This, similar to the case of cluster division, avoids the disadvantages of dependence on initial values and increased computational complexity, and reduces the computational load while setting the target number of clusters in advance and performing the calculation.
[0088] The present disclosure can also be implemented by, for example, causing a processor to execute a computer program. Figure 5 The processing shown.
[0089] The program can be stored and provided to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., floppy disks, magnetic tapes, hard disk drives), optical magnetic storage media (e.g., optical magnetic disks), CD-ROMs (read only memory), CD-Rs, CD-R / Ws, semiconductor memories (e.g., mask ROMs, PROMs (programmable ROMs), EPROMs (erasable PROMs), flash memories, and RAMs (random access memory)).
[0090] In addition, the program can also be provided to the computer via various types of transitory computer-readable media. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. Transitory computer-readable media can provide the program to the computer via a wired communication line such as an electric wire or optical cable, or a wireless communication line.
[0091] Each component constituting the information processing system 4 involved in the above-mentioned embodiment can be implemented not only by a program, but also partially or entirely by dedicated hardware such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field-Programmable Gate Array).
[0092] It is obvious from the present disclosure described above that the embodiments of the present disclosure can be changed in many ways. Such changes should not be considered to depart from the spirit and scope of the present disclosure, and all such modifications obvious to those skilled in the art are included in the scope of the technical solution.
Claims
1. An information processing system comprising a data acquisition unit and a clustering unit, wherein the data acquisition unit acquires multidimensional vector data, and the clustering unit performs hierarchical clustering as follows: based on the relative distance of the vector data acquired by the data acquisition unit, the data having a close relative distance are grouped together, and the grouping is repeated, wherein: The information processing system includes a density calculation unit that calculates the density of each cluster of the vector data. When performing the hierarchical clustering, the clustering unit divides or integrates the clusters based on the density of the clusters calculated by the density calculation unit.
2. The information processing system according to claim 1, wherein: The clustering unit extracts at least two vectors by comparing the densities of each vector component of the multidimensional vector data, and extracts the dense parts where the vector values are dense in each extracted vector, and extracts the focus area for clustering based on the extracted dense parts, and performs the hierarchical clustering on the extracted focus area and the surrounding area surrounding the focus area.
3. The information processing system according to claim 1, wherein: The clustering unit divides the clusters when it is determined that the density of the clusters calculated by the density calculation unit is greater than a threshold value, and If it is determined that the density of each cluster is not equal to or greater than the threshold, it is determined whether the current number of clusters is equal to or greater than the target value. If it is determined that the current number of clusters is equal to or greater than the target value, the hierarchical clustering is stopped.
4. The information processing system according to claim 1, wherein: When the clustering unit determines that the density of each cluster has reached a predetermined value during the hierarchical clustering, the clustering unit temporarily stops the hierarchical clustering and outputs the current number of clusters and vectors included in each cluster.
5. The information processing system according to claim 1, wherein: The density is the Gini coefficient.
6. An information processing method, comprising: Steps to obtain multi-dimensional vector data; Performing the following hierarchical clustering step, that is, grouping data with close relative distances based on the obtained relative distances of the vector data, and repeating the grouping; The step of respectively calculating the density of each cluster of the vector data; When performing the hierarchical clustering, the clusters are divided or integrated based on the calculated densities of the clusters.
7. A program product that causes a computer to execute the following processing: Obtain processing of multi-dimensional vector data; Performing the following hierarchical clustering process, that is, based on the obtained relative distance of the vector data, grouping the data with close relative distances, and repeating the grouping; Calculating the density of each cluster of the vector data respectively; When performing the hierarchical clustering, the clusters are divided or integrated based on the calculated densities of the clusters.