Data analysis device, data analysis method, and program

JPWO2025115218A1Pending Publication Date: 2025-06-05
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025560513
Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2023-12-01
Publication Date
2025-06-05
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art
No text content released.

Claims

1. A clustering means that performs clustering of words represented by tags included in a set of combinations of data names and tags representing the attributes of the data names, A representative word determination means for determining a representative word that represents each of the clusters generated based on the aforementioned clustering, In the set, a replacement means for replacing a tag representing a word belonging to the cluster with a representative word that represents the cluster to which the word belongs, A data analysis device having the following features.

2. The set has a generation means for generating frequency information relating to the frequency of occurrence of each word in the set, The data analysis apparatus according to claim 1, wherein the clustering means filters words to be clustered based on the frequency information.

3. The collection has a generation means for generating classification information based on the part of speech or spelling of each word in the collection, The data analysis apparatus according to claim 1, wherein the clustering means performs clustering for each group of words separated based on the classification information.

4. The data analysis apparatus according to claim 1, further comprising cluster filtering means for excluding or modifying clusters that do not meet predetermined conditions among the clusters generated by the clustering described above.

5. The data analysis apparatus according to claim 4, wherein the cluster filtering means sets a condition relating to the number of words in the cluster, or a condition relating to the variation of word vectors obtained by vectorizing the words within the cluster, as the predetermined conditions.

6. The data analysis apparatus according to claim 1, wherein the representative word determination means generates the representative word for each of the clusters based on a language model that has been machine-trained to generate an answer to a fill-in-the-blank question when a fill-in-the-blank question is given, or a language model that has been machine-trained to generate a word that represents a set of words when a set of words is given.

7. The data analysis apparatus according to claim 1, further comprising word filtering means for filtering out words that are not suitable to belong to the cluster represented by the representative word, based on the representative word.

8. The data analysis apparatus according to claim 1, wherein the replacement means modifies the replacement based on the search results of the data name associated with the replaced tag and the representative word replaced by the replacement.

9. Computers Clustering is performed on the words represented by tags included in the set of combinations of data names and tags representing the attributes of the said data names. Based on the clustering described above, a representative word is determined to represent each of the clusters generated. In the set, the tags representing words belonging to the cluster are replaced with the representative word that represents the cluster to which the word belongs. Data analysis methods.

10. Clustering is performed on the words represented by tags included in the set of combinations of data names and tags representing the attributes of the said data names. Based on the clustering described above, a representative word is determined to represent each of the clusters generated. A program that causes a computer to perform a process of replacing tags representing words belonging to a cluster in the aforementioned set with a representative word that represents the cluster to which the word belongs.