Data analysis device, data analysis method, and storage medium

The data analysis apparatus and method address the inefficiency caused by multiple similar tags by clustering and aggregating them using representative words, thereby enhancing data analysis efficiency.

WO2025115218A1PCT designated stage expired Publication Date: 2025-06-05NEC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/043126
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

In data analysis, the presence of multiple similar tags in a table increases the number of unique tags, which can hinder efficiency in data analysis.

Method used

A data analysis apparatus, method, and storage medium that perform clustering of tags, determine representative words for each cluster, and replace similar tags with their representative words, thereby aggregating tags and reducing the number of unique tags.

Benefits of technology

The solution effectively aggregates tags, reducing the number of unique tags and improving the efficiency of data analysis by grouping similar attributes into representative words.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023043126_05062025_PF_FP_ABST
    Figure JP2023043126_05062025_PF_FP_ABST
Patent Text Reader

Abstract

A data analysis device 1X mainly comprises a clustering means 52X, a representative word determination means 54X, and a replacement means 56X. The clustering means 52X performs clustering of words represented by tags included in a set of combinations of a data name and a tag representing an attribute of the data name. The representative word determination means 54X determines a representative word representing each cluster generated on the basis of the clustering. In the set, the replacement means 56X replaces a tag representing a word belonging to a cluster with a representative word representing the cluster to which the word belongs.
Need to check novelty before this filing date? Find Prior Art

Description

Data analysis device, data analysis method and storage medium

[0001] The present disclosure relates to the technical fields of a data analysis device, a data analysis method, and a storage medium that perform processing related to data analysis.

[0002] There are technologies that assign tags that represent characteristics of an analysis target in order to improve the efficiency of data analysis, etc. Patent Literature 1 discloses a technology that extracts characteristic words from clusters generated by word clustering.

[0003] Japanese Patent Application Laid-Open No. 2002-041544

[0004] In a table with a large number of tags, the existence of multiple similar tags increases the number of unique tags, which may be undesirable for efficient data analysis.

[0005] In view of the above-mentioned problems, one of the objectives of the present disclosure is to provide a data analysis device, a data analysis method, and a storage medium that are capable of suitably aggregating tags included in a set of combinations of data names and tags representing attributes of the data names.

[0006] One aspect of the data analysis device is a data analysis device having: a clustering means for clustering words represented by tags included in a set of combinations of data names and tags representing attributes of the data names; a representative word determination means for determining a representative word representing each of the clusters generated based on the clustering; and a replacement means for replacing, in the set, tags representing words belonging to the clusters with the representative word representing the cluster to which the word belongs.

[0007] One aspect of the data analysis method is a data analysis method in which a computer clusters words represented by tags included in a set of combinations of data names and tags representing attributes of the data names, determines a representative word to represent each of the clusters generated based on the clustering, and replaces tags representing words belonging to the clusters in the set with the representative word that represents the cluster to which the word belongs.

[0008] One aspect of the storage medium is a storage medium that stores a program that causes a computer to perform the following processes: clustering words represented by tags included in a set of combinations of data names and tags representing attributes of the data names; determining a representative word that represents each of the clusters generated based on the clustering; and replacing tags representing words that belong to the clusters in the set with the representative word that represents the cluster to which the word belongs.

[0009] One example of the effect of the present disclosure is that tags included in a set of combinations of data names and tags representing attributes of the data names can be suitably aggregated.

[0010] 1 shows the configuration of a data analysis system; 2 shows the hardware configuration of a data analysis device; 3 shows an example of functional blocks of a processor of the data analysis device; 4 shows an overview of the processing of a data analysis unit; 5 shows a functional block diagram of a data analysis unit; 6 shows an example of a display screen showing the results of tag aggregation processing; 7 is an example of a flowchart showing an overview of processing executed by a data analysis device; 8 shows the configuration of a data analysis system; 9 shows the relationship between a user, a data analysis device, and a terminal device; 10 is a functional block diagram of a data analysis device; 11 is an example of a flowchart showing the processing procedure of a data analysis device

[0011] Hereinafter, embodiments of a data analysis device, a data analysis method, and a storage medium will be described with reference to the drawings. Hereinafter, a "query" refers to a natural language inquiry (including a question or a hypothesis sentence) passed from a user to the data analysis system 100. An "answer" refers to a natural language sentence or text data representing a natural language sentence output by the data analysis system 100 in response to the query.

[0012] <First Embodiment> (1) System Configuration Fig. 1 shows the configuration of a data analysis system 100. The data analysis system 100 is a system that aggregates tags that represent attributes by analyzing specified data, and mainly includes a data analysis device 1, an input device 2, a display device 3, and a storage device 4.

[0013] The data analysis device 1 analyzes a specified table (table data) to aggregate tags contained in the table and controls the display of information related to the tag aggregation results. Hereinafter, the table into which tags are aggregated is also referred to as a "tag table." A tag table associates, for each record, the name of the data to be tagged (also referred to as a "data name") with a set of tags representing the attributes of the data name. Note that the data name may be any name to be tagged, such as a person, store, or product. The data analysis device 1 communicates data with the input device 2, display device 3, and storage device 4 via a communication network or by direct wireless or wired communication. While the following embodiment describes the case in which the data analysis device 1 handles table data, the present invention is also applicable to simple character string data, such as CSV file format. A table is an example of a "set of combinations of data names and tags representing the attributes of the data names."

[0014] The input device 2 is an interface that accepts user input, which is external input, and corresponds to, for example, a touch panel, buttons, a keyboard, a voice input device, etc. The input device 2 supplies input information generated based on the user input to the data analysis apparatus 1.

[0015] The display device 3 is, for example, a display, a projector, or the like, and performs a predetermined display based on the display information supplied from the data analysis device 1 .

[0016] The storage device 4 is a memory that stores various information necessary for the processing executed by the data analysis device 1. The storage device 4 may store, for example, model information (configuration information) for configuring a machine-learned large-scale language model (LLM), model information for configuring any natural language understanding model such as BERT (Bidirectional Encoder Representations from Transformers) used in natural language processing, etc. The model information includes, for example, various parameters of the machine-learned deep learning model, such as the layer structure, the neuron structure of each layer, the number and filter size of filters in each layer, and the weight of each element of each filter.

[0017] The LLM is a natural language processing model trained using a large amount of text data, and receives text data representing sentences as input and outputs text data representing sentences. When text data representing a question is input to the LLM, the LLM outputs text data representing an answer. Hereinafter, the text data input to the LLM will be referred to as a "prompt." The LLM will also be simply referred to as a language model. In addition to the above-mentioned BERT, a specific example of a language model is a Generative Pre-Trained Transformer (GPT), which predicts a character string that is likely to follow an input character string and outputs a sentence containing the input character string. Other examples of language models include T5 (Text-to-Text Transfer Transformer), RoBERTa (Robustly optimized BERT approach), and ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately).

[0018] The storage device 4 may be a storage device such as a hard disk connected to or built into the data analysis device 1, or may be a storage medium such as a flash memory. The storage device 4 may also be a server device that performs data communication with the data analysis device 1. In this case, the storage device 4 may be composed of multiple server devices.

[0019] The configuration of the data analysis system 100 shown in FIG. 1 is an example, and various modifications may be made to the configuration. For example, the input device 2 and the display device 3 may be configured as an integrated device. In this case, the input device 2 and the display device 3 may be configured as a tablet terminal integrated with the data analysis device 1. The data analysis device 1 may be connected to or have a built-in sound output device, such as a speaker, and output information by sound. The data analysis device 1 may also be configured from multiple devices. In this case, the multiple devices that make up the data analysis device 1 exchange information necessary to execute pre-assigned processing between these multiple devices.

[0020] (2) Hardware Configuration of Data Analysis Apparatus Fig. 2 shows the hardware configuration of the data analysis apparatus 1. The data analysis apparatus 1 includes, as hardware components, a processor 11, a memory 12, and an interface 13. The processor 11, the memory 12, and the interface 13 are connected via a data bus 19.

[0021] The processor 11 executes predetermined processes by executing programs stored in the memory 12. The processor 11 is a processor such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or a TPU (Tensor Processing Unit). The processor 11 may be composed of multiple processors. The processor 11 is an example of a computer.

[0022] The memory 12 is composed of various types of volatile and non-volatile memories, such as RAM (Random Access Memory) and ROM (Read Only Memory). The memory 12 also stores programs for the data analysis apparatus 1 to execute various processes. The memory 12 is also used as a working memory, and temporarily stores information obtained from the storage device 4. The memory 12 may also function as the storage device 4. Similarly, the storage device 4 may also function as the memory 12 of the data analysis apparatus 1. The programs executed by the data analysis apparatus 1 may be stored in a storage medium other than the memory 12.

[0023] The interface 13 is an interface for electrically connecting the data analysis device 1 to other devices. These interfaces may be wireless interfaces such as network adapters for wirelessly transmitting and receiving data to and from other devices, or may be hardware interfaces for connecting to other devices via cables or the like.

[0024] The hardware configuration of the data analysis device 1 is not limited to the configuration shown in Fig. 2. For example, the data analysis device 1 may include at least one of the input device 2 and the display device 3. Furthermore, the data analysis device 1 may be connected to or have a built-in sound output device such as a speaker.

[0025] (3) Processing Overview Fig. 3 shows an example of functional blocks of the processor 11. Functionally, the processor 11 has a data analysis unit 15 and a UI (User Interface) control unit 16. Note that in Fig. 3, blocks that exchange data are connected by solid lines, but the combination of blocks that exchange data is not limited to Fig. 3. The same applies to other functional block diagrams described later.

[0026] The data analysis unit 15 refers to various information stored in the storage device 4, memory 12, etc., performs data analysis on the tag table specified by the user, and aggregates tags in the tag table. The data analysis unit 15 then supplies the tag table modified so that the tags are aggregated (also referred to as a "modified tag table") to the UI control unit 16 as a tag aggregation result. Details of the processing by the data analysis unit 15 will be described later.

[0027] The UI control unit 16 controls the reception of user input and the display of information to be viewed by the user. For example, the UI control unit 16 identifies a tag table based on input information (i.e., external input) supplied from the input device 2 in response to a user operation, and supplies the tag table to the data analysis unit 15. In this case, the UI control unit 16 identifies, for example, a table with a tag specified by the user or a table corresponding to some of its records as the tag table.

[0028] Furthermore, the UI control unit 16 generates display information based on the tag aggregation result generated by the data analysis unit 15, and controls the display of the display device 3 by supplying the generated display information to the display device 3. Specific processing by the UI control unit 16 will be described later with reference to display examples.

[0029] The components of the data analysis unit 15 and the UI control unit 16 described in FIG. 3 can be realized, for example, by the processor 11 executing a program. Alternatively, the necessary programs may be recorded on any non-volatile storage medium and installed as needed to realize the components. At least some of these components may not necessarily be realized by software programs, but may be realized by any combination of hardware, firmware, and software. At least some of these components may be realized using a user-programmable integrated circuit, such as an FPGA (Field-Programmable Gate Array) or a microcontroller. In this case, the integrated circuit may be used to realize a program consisting of the above components. Furthermore, at least a portion of each component may be configured by an ASSP (Application Specific Standard Product), an ASIC (Application Specific Integrated Circuit), or a quantum processor (quantum computer control chip). In this way, each component may be realized by various hardware. The same applies to other embodiments described below. Furthermore, each of these components may be realized by the cooperation of multiple computers, for example, using cloud computing technology.

[0030] (4) Tag Aggregation Processing Next, the tag aggregation processing executed by the data analysis unit 15 will be described.

[0031] (4-1) Overview Fig. 4 is a diagram showing an overview of the tag aggregation process by the data analysis unit 15. As shown in Fig. 4, when the data analysis unit 15 receives a tag table from the UI control unit 16, it performs tag aggregation processing and generates a modified tag table in which each data name is associated with the aggregated tag.

[0032] 4, the tag table associates a data name representing a store name with a set of tags representing the attributes of the corresponding data name. For example, the first record in the tag table associates a set of tags such as "lunch" and "cafe au lait" with the data name "Store X," and the second record in the tag table associates a set of tags such as "family" and "restaurant" with the data name "Store Y."

[0033] In this case, the data analysis unit 15 clusters the words represented by the tags based on the frequency of occurrence of words (e.g., lunch, cafe au lait, family, restaurant) from the set of tags included in the tag table, excluding duplicates, and auxiliary information such as the classification (category) of each word, and determines a representative word for each cluster. As will be described later, the data analysis unit 15 also filters (i.e., selectively excludes) the words to be clustered and filters the clusters generated by clustering. In this example, the representative word for the cluster consisting of "lunch," "terrace," and "picnic" is "gourmet," and the representative word for the cluster consisting of "cafe au lait," "cafe mocha," and "Frappuccino" is "coffee." Hereinafter, "words" are considered to be the same word when aggregating the frequency of occurrence, etc., for words with the same character string.

[0034] Next, the data analysis unit 15 replaces tags in the tag table based on the determined representative words. The data analysis unit 15 also determines the validity of the tag replacement and corrects any inappropriate tag replacements. This generates a corrected tag table in which tags are appropriately aggregated. In the example of Figure 4, a set of tags such as "gourmet" (a tag replaced from "lunch") and "coffee" (a tag replaced from "cafe au lait") is associated with the data name "Store X."

[0035] In this way, the data analysis unit 15 determines a representative word that represents a group of attributes represented by tags and replaces similar tags (attributes) with the representative word, thereby preferably reducing the number of unique tags. For example, tags with similar meanings, such as "family," "family," and "family demographic," can be prevented from being treated as representing different attributes. Furthermore, by clustering words based on the frequency of appearance or classification of each word represented by a tag, it is possible to preferably prevent clustering failure due to the presence of unimportant words or the generation of unnatural representative words.

[0036] Although Figure 4 illustrates a tag table related to a store, tag tables related to products, services (including healthcare), celebrities (including medical professionals such as doctors and nurses), diseases, patients, place names, occupations, advertisements, and any other taggable items may also be used.

[0037] 5 is a functional block diagram of the data analysis unit 15. Functionally, the data analysis unit 15 includes an auxiliary information generation unit 51, a clustering unit 52, a cluster filtering unit 53, a representative word determination unit 54, a word filtering unit 55, a tag replacement / correction unit 56, and a re-execution unit 57.

[0038] The auxiliary information generation unit 51 generates auxiliary information indicating auxiliary information for each word represented by a tag included in the tag table (such as "lunch" and "cafe au lait" in the example of FIG. 4), and supplies the generated auxiliary information to the clustering unit 52. In this case, the auxiliary information corresponds to meta-information for each word, and includes, for example, at least one of frequency information regarding the frequency of appearance in the tag table and classification information regarding classification in a linguistic system.

[0039] For example, the auxiliary information generation unit 51 calculates the number of occurrences of each word in the entire tag table as the frequency information. Instead of the number of occurrences, frequency information indicating an index value representing a frequency equivalent to the number of occurrences, such as the appearance rate for each record, may be generated. Furthermore, as the classification information, the auxiliary information generation unit 51 recognizes the part-of-speech classification (such as noun or adjective) of each word based on any natural language processing, and generates information indicating the recognized part-of-speech classification (category) as classification information associated with the word. In another example, the auxiliary information generation unit 51 recognizes the orthographic classification (such as katakana or English) of each word based on any natural language processing, and generates information indicating the recognized orthographic classification as classification information associated with the word.

[0040] The clustering unit 52 performs clustering (word clustering) of words associated with the auxiliary information, based on the auxiliary information generated by the auxiliary information generation unit 51. Details of word clustering will be described later. The clustering unit 52 supplies the clustering result, which indicates the words classified into clusters, to the cluster filtering unit 53.

[0041] The cluster filtering unit 53 filters out clusters that do not satisfy a predetermined condition from the clusters generated by the clustering unit 52 (cluster filtering). In other words, the cluster filtering unit 53 excludes clusters that do not satisfy the predetermined condition, or modifies the clusters by excluding some words so that the condition is satisfied. The predetermined condition may be based on the number of words in each cluster, or may be based on the variation in word vectors for each cluster. Here, cluster filtering generates words (tags) that do not belong to any cluster. A set of words that do not belong to any cluster is also called an "exclusion word set." The cluster filtering unit 53 supplies information about the filtered clusters and the exclusion word set to the representative word determination unit 54.

[0042] The representative word determination unit 54 determines a representative word that represents each cluster after cluster filtering. In this case, the representative word determination unit 54 may extract a representative word from the words in each cluster, or may generate a representative word from the words in each cluster. The representative word determination unit 54 notifies the word filtering unit 55 of information about each cluster and its representative word.

[0043] The word filtering unit 55 performs word filtering to remove words that are not suitable for belonging to the cluster represented by each representative word, based on each representative word. Specifically, the word filtering unit 55 removes words that are inconsistent with the representative word for each cluster and adds them to a set of excluded words. The word filtering unit 55 then supplies information indicating each cluster after word filtering and the representative word for each cluster to the tag replacement / modification unit 56.

[0044] The tag replacement / correction unit 56 replaces each tag in the tag table with the representative word of the cluster to which the tag belongs. The tag replacement / correction unit 56 also performs processing to correct inappropriate tags for each record in the tag table after the replacement. The tag replacement / correction unit 56 supplies the tag table after the above processing to the re-execution unit 57 as a corrected tag table.

[0045] The re-execution unit 57 determines whether or not it is necessary to re-execute (hereinafter simply referred to as "re-execution") part or all of the series of processes for generating a corrected tag table from a tag table, based on information output by the auxiliary information generation unit 51, the clustering unit 52, the cluster filtering unit 53, the representative word determination unit 54, the word filtering unit 55, and / or the tag replacement / correction unit 56. The method for this determination will be described later. If the re-execution unit 57 determines that re-execution is necessary, it instructs the other functional blocks to re-execute the processes.

[0046] On the other hand, if the re-execution unit 57 determines that re-execution is not necessary, it supplies the tag aggregation result including the modified tag table generated immediately before by the tag replacement / modification unit 56 to the UI control unit 16. In this case, the UI control unit 16 may display information about the tag aggregation result on the display device 3. An example of this display will be described later.

[0047] Hereinafter, the processes executed by the clustering unit 52, the cluster filtering unit 53, the representative word determination unit 54, the word filtering unit 55, the tag replacement / correction unit 56, and the re-execution unit 57 will be described in detail.

[0048] (4-2) Word Clustering The word clustering performed by the clustering unit 52 will now be described.

[0049] In a first example of word clustering, the clustering unit 52 performs word clustering based on frequency information of auxiliary information associated with each word. For example, the clustering unit 52 excludes (filters) words whose occurrence frequency indicated by the frequency information is lower than a predetermined threshold (e.g., 5 times), and performs clustering based on an arbitrary clustering method on words whose occurrence frequency is equal to or higher than the predetermined threshold. In this case, the clustering unit 52 adds words whose occurrence frequency is lower than the predetermined threshold (i.e., words excluded by filtering) to a set of excluded words.

[0050] Examples of the above-mentioned clustering method include the k-means method and hierarchical clustering (aggregative clustering). In this case, the clustering unit 52 calculates a distance / similarity matrix between all words. It is generally known that basic clustering can be calculated from a distance matrix. The above-mentioned similarity may be calculated as an index representing notational closeness, such as a Jaccard coefficient, or as an index representing semantic closeness. For example, as an index representing semantic closeness, the clustering unit 52 may calculate cosine similarity or the like after vectorizing words using a language model such as word2vec or BERT.

[0051] In a second example of word clustering, the clustering unit 52 performs word clustering based on the classification information of the auxiliary information associated with each word. The clustering unit 52 groups words into categories based on the parts of speech or spellings indicated by the classification information, and uses all clusters obtained by performing clustering for each group as the clustering result. For example, the clustering unit 52 groups words into two groups based on the categories based on the parts of speech or spellings indicated by the classification information, generates m clusters (m is a positive integer) for one group, and generates n clusters (n is a positive integer) for the other group, and outputs a total of m + n clusters as the clustering result.

[0052] The clustering unit 52 may execute a combination of the first and second examples of word clustering described above. In this case, the clustering unit 52 performs filtering based on frequency information, and then performs clustering for each group divided into categories based on classification information.

[0053] The clustering unit 52 may perform word clustering without using auxiliary information. A specific example of word clustering without using auxiliary information will be described below.

[0054] For example, the clustering unit 52 performs any clustering such as the k-means method or hierarchical clustering described above on the words represented by the tags included in the tag table.

[0055] In another example, the clustering unit 52 performs word clustering using an LLM. In this case, for example, template information indicating a template for generating a prompt to be input to the LLM is stored in the storage device 4 or the like, and the clustering unit 52 generates a prompt by referring to the template information. For example, the clustering unit 52 generates a prompt using the following template, and obtains the answer output by the LLM in response to the prompt as the clustering result.

[0056] "{Word list} Please cluster the above words and output them in CSV format."

[0057] In this case, a character string listing the words represented by the tags included in the tag table is inserted into "{word list}." By generating such a prompt, the clustering unit 52 can preferably obtain clustering results.

[0058] (4-3) Cluster Filtering Next, the cluster filtering performed by the cluster filtering unit 53 will be described.

[0059] In a first example of cluster filtering, the cluster filtering unit 53 determines whether to exclude or correct a cluster based on the number of words in the cluster. Specifically, the cluster filtering unit 53 considers a cluster with a word count greater than a predetermined threshold (e.g., 100) to be a cluster to be excluded or corrected. The threshold is stored, for example, in the storage device 4 or the memory 12. Note that the number of words in this case may be a number that takes into account the number of occurrences indicated by the frequency information of each word. That is, the cluster filtering unit 53 may calculate the number of words in the cluster as the sum of the number of occurrences of each word in the cluster (i.e., the number of tags).

[0060] When a cluster to be excluded or corrected is identified, the cluster filtering unit 53, for example, deletes the identified cluster and adds the words belonging to the deleted cluster to the set of excluded words. In another example, the cluster filtering unit 53 excludes words belonging to the identified cluster from the cluster and adds them to the set of excluded words until the number of words belonging to the identified cluster reaches the above-mentioned threshold. In this case, the cluster filtering unit 53, for example, refers to word frequency information and excludes words from the cluster in order of decreasing frequency indicated by the frequency information.

[0061] In the second example of cluster filtering, the cluster filtering unit 53 determines whether to exclude or correct a cluster based on the variance within the cluster of word vectors obtained by vectorizing words. Specifically, the cluster filtering unit 53 identifies a cluster whose variance within the cluster is greater than a predetermined threshold as a cluster to be excluded or corrected. The threshold is stored in, for example, the storage device 4 or the memory 12. After identifying a cluster to be excluded or corrected, the cluster filtering unit 53 executes the same process as in the first example of cluster filtering to exclude or correct the identified cluster.

[0062] In the calculation of the variance value described above, the cluster filtering unit 53 may calculate a variance value that is sloped according to the number of words in the cluster so that the variance value increases as the number of words in the cluster increases. For example, if the variance value before slope is "V" and the number of words in the cluster is "C", the cluster filtering unit 53 calculates the variance value "V'" after slope as follows: V' = √C V

[0063] In this case, the greater the number of words, the greater the calculated variance value. Note that the cluster filtering unit 53 may gradient the variance value based on any information, not limited to the number of words in the cluster. The cluster filtering unit 53 may also calculate a word vector for each tag and calculate the variance value. In this case, the cluster filtering unit 53 calculates the variance value using the same word vector for each word the number of times it appears.

[0064] In a third example of cluster filtering, the cluster filtering unit 53 performs cluster filtering using an LLM. In this case, for example, template information indicating a template for generating a prompt to be input to the LLM is stored in the storage device 4 or the like, and the clustering unit 52 generates a prompt by referring to the template information. For example, the cluster filtering unit 53 generates a prompt for each cluster using the following template, and obtains the answer output by the LLM in response to the prompt as the result of cluster filtering for each cluster.

[0065] "{List of words in cluster} is the result of word clustering. Are the clusters above cohesive? If not, remove the words that prevent cohesiveness and output a new cluster."

[0066] In this case, a list of words belonging to the target cluster is inserted in "{list of words in cluster}." By generating such a prompt, the cluster filtering unit 53 can preferably obtain the results of cluster filtering for each cluster.

[0067] (4-4) Determination of Representative Words Next, the determination of representative words for each cluster by the representative word determination unit 54 will be described.

[0068] In a first example of a method for determining a representative word, the representative word determination unit 54 extracts a word that best represents a cluster from among the words belonging to the cluster as the representative word. In this case, for example, the representative word determination unit 54 selects a representative word by using, as the target text, a text in which the words belonging to the cluster are listed using appropriate delimiters, etc., and executing any keyphrase extraction (or keyword extraction) algorithm. In this case, the representative word determination unit 54 may select, as the representative word, a word corresponding to a word vector that is closest to the average vector of the word vectors of the words in the cluster.

[0069] In a second example of a representative word determination method, the representative word determination unit 54 performs a process of generating representative words (so-called keyphrase generation) using a language model. For example, the representative word determination unit 54 generates representative words using an MLM (Masked Language Model) such as BERT. In this case, for example, the representative word determination unit 54 inputs the following sentence into the MLM to obtain a representative word that matches "{MASK}".

[0070] "The representative word that best represents {list of words in the cluster} is {MASK}"

[0071] A list of words belonging to the target cluster is inserted into "{list of words in cluster}" in the above sentence. In this case, template information indicating a template of a sentence to be input to the MLM is stored in the storage device 4 or the like, and the representative word determination unit 54 generates the above sentence by referring to the template information and inputs it to the MLM. The MLM is an example of a language model that has undergone machine learning so as to generate an answer to a fill-in-the-blank question when given.

[0072] The representative word determination unit 54 may generate representative words using an LLM. In this case, for example, template information indicating a template for generating a prompt to be input to the LLM is stored in the storage device 4 or the like, and the representative word determination unit 54 generates a prompt by referring to the template information. For example, the representative word determination unit 54 generates a prompt using the following template, and obtains the answer output by the LLM in response to the prompt as the clustering result.

[0073] "What is the most representative word that best describes {list of words in the cluster}? Please answer in one word."

[0074] A list of words belonging to the target cluster is inserted into {List of words in cluster} in the above prompt. The representative word determination unit 54 inputs the above prompt into the LLM and obtains the words output by the LLM as representative words. The LLM is an example of a language model that has undergone machine learning so that, when a set of words is given, it generates a word that represents the set of words.

[0075] (4-5) Word Filtering Next, word filtering by the word filtering unit 55 will be described.

[0076] In a first example of word filtering, the word filtering unit 55 calculates the similarity between the representative word and each word in each cluster using each word vector, and excludes from the cluster any words whose similarity is equal to or less than a predetermined threshold. The threshold is stored in, for example, the storage device 4 or the memory 12. In this case, the word filtering unit 55 adds the words excluded from the cluster to a set of excluded words.

[0077] In a second example of word filtering, the word filtering unit 55 performs word filtering using the LLM. In this case, for example, template information indicating a template for generating a prompt to be input to the LLM is stored in the storage device 4 or the like, and the word filtering unit 55 generates a prompt by referring to the template information. For example, the word filtering unit 55 generates a prompt for each cluster using the following template, and obtains the answer output by the LLM in response to the prompt as the result of word filtering for each cluster.

[0078] "Please remove words that are unrelated to {representative word} from the word group {word list in cluster}."

[0079] In this case, a list of words belonging to the target cluster is inserted into "{list of words in cluster}" and a representative word of the target cluster is inserted into "{representative word}." By generating such a prompt, the word filtering unit 55 can preferably obtain the word filtering results for each cluster.

[0080] (4-6) Tag Replacement / Modification Next, tag replacement and modification of the replacement results by the tag replacement / modification unit 56 will be described.

[0081] The tag replacement / modification unit 56 searches the tag table for tags that match words belonging to each cluster, and replaces the searched tags with the representative words of the clusters to which the words belong. In other words, the tag replacement / modification unit 56 performs a process of replacing tags representing words belonging to one of the clusters with the representative words representing the clusters to which the words belong, for all words belonging to the clusters in order. Here, the tag replacement / modification unit 56 may leave tags representing words belonging to the excluded word set in the tag table as they are, or may remove them from the tag table.

[0082] As a first example of correcting the replacement result, the tag replacement / correction unit 56 determines whether the replacement is appropriate based on the search results of the data name associated with the replaced tag and the representative word replaced with the tag. If it determines that the replacement is inappropriate, it corrects the replaced tag. For example, the tag replacement / correction unit 56 sequentially selects each tag in the tag table, performs a search using an arbitrary search engine with a pair of the selected tag and the corresponding data name as a search query, and removes the selected tag from the tag table if the number of search hits is below a predetermined threshold. The threshold is stored, for example, in the storage device 4 or memory 12. The search database in this case may be web information on the Internet or a database managed by an administrator of the data analysis system 100.

[0083] As a second example of modifying the replacement results, the tag replacement / modification unit 56 modifies the tags in the tag table based on the LLM. In this case, for example, template information indicating a template for generating a prompt to be input into the LLM is stored in the storage device 4 or the like, and the tag replacement / modification unit 56 references the template information and generates a prompt for each tag replaced by the representative word, requesting a determination of whether or not modification is necessary and the presentation of an alternative if modification is necessary. For example, the tag replacement / modification unit 56 generates a prompt for each replaced tag using the following template, and modifies the tag replaced by the representative word based on the answer output by the LLM in response to the prompt.

[0084] "Word group: {All representative words} Is it correct that {data name}, which was originally tagged with {tag list in record}, ​​will be tagged with {replaced tag}? If not, please select up to three appropriate words from the word group above and output them."

[0085] In this case, a string listing the representative words for all clusters is inserted in "{all representative words}," a string listing all tags in the record to which the tag replaced by the representative word belongs is inserted in "{tag list in record}," the data name corresponding to the record is inserted in "{data name}," and the representative word representing the replaced tag is inserted in "{replaced tag}." If the LLM outputs a response indicating that the replaced tag is incorrect, the tag replacement / correction unit 56 replaces the replaced tag with a candidate selected by any method from the candidate corrected tag indicated by the response (up to three candidates in this case). When such a prompt is used, the candidate corrected tag indicated by the LLM's response is selected from the representative words, so the number of unique tags in the corrected tag table does not increase.

[0086] (4-7) Re-execution Unit A specific example of determining whether or not it is necessary to re-execute part or all of the series of processes for generating a modified tag table from a tag table will be described.

[0087] For example, the re-execution unit 57 determines whether re-execution is necessary based on the number of tags for each record in the modified tag table. In this case, for example, the re-execution unit 57 sequentially selects records in the modified tag table, and determines that re-execution is necessary if the number of tags in the selected record is less than a predetermined threshold (for example, three or less), and determines that re-execution is not necessary if the number of tags is more than the predetermined threshold. The above-mentioned threshold is stored, for example, in the storage device 4 or the memory 12.

[0088] The re-execution unit 57 may determine whether re-execution is necessary based on the output result of any processing block. For example, the re-execution unit 57 determines that re-execution is necessary when there is a representative word among the representative words generated by the representative word determination unit 54 whose number of characters is equal to or greater than a predetermined threshold, and determines that re-execution is unnecessary when the number of characters of all the representative words is less than the predetermined threshold. In another example, the re-execution unit 57 determines that re-execution is necessary when the number of clusters indicated by the cluster filtering result generated by the cluster filtering unit 53 is less than a predetermined threshold, and determines that re-execution is unnecessary when the number of clusters is equal to or greater than the predetermined threshold. The above-mentioned threshold is stored, for example, in the storage device 4 or the memory 12.

[0089] Next, a specific embodiment of re-execution will be described. When the re-execution unit 57 determines that re-execution is necessary, it may instruct the start of re-execution from any process selected from the word clustering process by the clustering unit 52, the cluster filtering process by the cluster filtering unit 53, the representative word determination process by the representative word determination unit 54, the word filtering process by the word filtering unit 55, the tag replacement process by the tag replacement / correction unit 56, or the tag replacement correction process by the tag replacement / correction unit 56. In this case, the data analysis device 1 preferably re-executes each process by changing the method and / or parameters (such as thresholds) randomly or according to a predetermined rule. In the re-execution, a series of processes is executed, from the process that starts the re-execution to the tag replacement correction process by the tag replacement / correction unit 56. This enables the data analysis device 1 to generate a new tag table as a corrected tag table that is different from the tag table before the re-execution.

[0090] For example, when re-execution is performed after word clustering, the clustering unit 52 performs new word clustering on the excluded word set and generates a clustering result by adding the generated cluster to the cluster generated before the re-execution. Hereinafter, the cluster generated by the re-execution is also referred to as the "second cluster." The clustering unit 52 then supplies the clustering result indicating the generated second cluster to the cluster filtering unit 53, etc. In this case, the clustering unit 52 may perform word clustering on the second cluster by changing at least one of the method or parameters of the word clustering performed before the re-execution. The clustering unit 52 may generate the second cluster described above using only the words added to the excluded word set by cluster filtering by the cluster filtering unit 53, or may generate the second cluster described above using the words added to the excluded word set by cluster filtering by the cluster filtering unit 53 and word filtering by the word filtering unit 55. Thereafter, the representative word determination unit 54 determines a representative word for the second cluster, the word filtering unit 55 performs word filtering for the second cluster based on the representative word, and the tag replacement / correction unit 56 performs the process of replacing tags representing words belonging to the second cluster with the representative word of the second cluster and the correction process, in the same manner as the process before re-execution.

[0091] (5) Display Example Fig. 6 is an example of a display screen showing the results of the tag aggregation process. When the re-execution unit 57 determines that re-execution is unnecessary, the UI control unit 16 generates a display signal to be supplied to the display device 3 based on each processing result supplied from the data analysis unit 15, and supplies the display signal to the display device 3, thereby causing the display device 3 to display the display screen shown in Fig. 6. The UI control unit 16 provides a tag table display field 61, a tag aggregation result display field 62, and a tag substitution table 63 on the display screen.

[0092] The tag table display field 61 is a field for displaying a tag table, and is provided with a file selection button 611. In this example, when the UI control unit 16 detects that the file selection button 611 has been selected, it displays a GUI for selecting a file to be used as the tag table, and accepts an input specifying a file to be used as the tag table from the files stored in the storage device 4 or the memory 12 (or an external device).

[0093] The tag aggregation result display field 62 is a display field that displays the tag aggregation result for the tag table displayed in the tag table display field 61. In this example, the UI control unit 16 displays the corrected tag table finally generated by the data analysis unit 15 in the tag aggregation result display field 62 as the tag aggregation result.

[0094] The tag replacement table 63 shows a correspondence table of tags before and after replacement when correcting the tag table displayed in the tag table display field 61 to the corrected tag table displayed in the tag aggregation result display field 62. The UI control unit 16 displays, as the tag replacement table 63, a table showing the correspondence between tags before replacement and tags after replacement based on the results of tag replacement and correction executed by the tag replacement / correction unit 56.

[0095] According to the display example shown in FIG. 6, the tag aggregation results for the tag table specified by the user are presented to the user, and it is possible to favorably support the user in data analysis and decision-making based on the data analysis results.

[0096] (6) Processing Flow FIG. 7 is an example of a flowchart showing an outline of the processing executed by the data analysis device 1.

[0097] First, the data analysis device 1 acquires a tag table based on input information and the like supplied by the input device 2 (step S11). The processing of step S11 corresponds to the processing executed by the UI control unit 16. Then, the data analysis device 1 generates auxiliary information for each word present in the tag table (step S12). The processing of step S12 corresponds to the processing executed by the auxiliary information generation unit 51.

[0098] Next, the data analysis device 1 performs word clustering on the words present in the tag table (step S13). As a result, the data analysis device 1 generates word clusters. The process of step S13 corresponds to the process executed by the clustering unit 52. Then, the data analysis device 1 performs cluster filtering on the word clusters generated in step S13 (step S14). As a result, the data analysis device 1 excludes or corrects inappropriate clusters. The process of step S14 corresponds to the process executed by the cluster filtering unit 53.

[0099] Next, the data analysis device 1 determines a representative word for each cluster after the cluster filtering, and performs word filtering for each cluster based on the representative word (step S15). The process of step S15 corresponds to the process executed by the representative word determination unit 54 and the word filtering unit 55.

[0100] Next, the data analysis device 1 replaces tags in the tag table based on the representative words and corrects the replaced tags (step S16). As a result, the data analysis device 1 obtains a corrected tag table by correcting the tag table. The processing of step S16 corresponds to the processing executed by the tag replacement / correction unit 56.

[0101] Next, the data analysis device 1 determines whether re-execution is necessary (step S17). In this case, the data analysis device 1 determines whether re-execution is necessary based on the processing result of step S16 or the processing result of other steps. If it is determined that re-execution is necessary (step S17; Yes), the data analysis device 1 returns the process to one of steps S11 to S16. Furthermore, the data analysis device 1 preferably changes at least one of the parameters or algorithms (methods) used in the processing of each step from the previous processing of the same step. This allows the data analysis device 1 to obtain a modified tag table in step S16 that is different from the previous one.

[0102] On the other hand, if it is determined that re-execution is not necessary (step S17; No), the data analysis device 1 displays information about the tag aggregation result on the display device 3 (step S18). In this case, for example, the data analysis device 1 displays the corrected tag table obtained in the last executed step S16, etc., as the tag aggregation result on the display device 3.

[0103] 8 shows the configuration of a data analysis system 100A. The data analysis system 100A mainly includes a data analysis device 1A and a terminal device 5. The data analysis device 1A and the terminal device 5 perform data communication via a network 6.

[0104] The data analysis apparatus 1A is one or more devices that function as a server (including a cloud server) and performs the processing executed by the data analysis apparatus 1 in the first embodiment. In this case, the data analysis apparatus 1A receives input information from the terminal apparatus 5 via the network 6, which the data analysis apparatus 1 receives from the input device 2 in the first embodiment. The data analysis apparatus 1A also transmits display information that the data analysis apparatus 1 transmitted to the display device 3 in the first embodiment to the terminal apparatus 5 via the network 6. The data analysis apparatus 1A also includes the storage apparatus 4 of the first embodiment, or references various pieces of information stored in the storage apparatus 4 via the network 6.

[0105] The terminal device 5 is a terminal having an input function, a display function, and a communication function, and functions as the input device 2 and the display device 3 in the first embodiment. The terminal device 5 may be, for example, a personal computer, a tablet terminal, a PDA (Personal Digital Assistant), or the like. The terminal device 5 transmits input information generated based on the received user input to the data analysis device 1A via the network 6. Furthermore, when the terminal device 5 receives display information from the data analysis device 1A, it displays information based on the display information.

[0106] The data analysis device 1A according to the second embodiment can preferably perform the input process and output process that the data analysis device 1 according to the first embodiment performs on the user of the terminal device 5 .

[0107] 9 is a diagram showing the relationship between a user, data analysis device 1A, and terminal device 5. In this case, data analysis device 1A functions as a server that executes processing related to attribute estimation for a specified tag table, and terminal device 5 functions as a user terminal that accepts input related to the tag table, etc. By exchanging information with data analysis device 1A, terminal device 5 presents a display screen such as that shown in FIG. 6 to the user. This can favorably prompt the user to make a decision.

[0108] 10 is a functional block diagram of a data analysis device 1X. The data analysis device 1X mainly includes a clustering unit 52X, a representative word determination unit 54X, and a replacement unit 56X. The data analysis device 1X may be composed of multiple devices.

[0109] The clustering means 52X performs clustering of words represented by tags included in a set of combinations of data names and tags representing attributes of the data names. Examples of "sets of combinations of data names and tags representing attributes of the data names" include structured data such as tables having multiple records, and semi-structured data represented in JSON format such as {data 1: tag list 1, data 2: tag list 2, ...}. The clustering means 52X can be, for example, the clustering unit 52 in the first or second embodiment.

[0110] The representative word determination means 54X determines a representative word that represents each of the clusters generated based on the clustering. The representative word determination means 54X can be, for example, the representative word determination unit 54 in the first or second embodiment.

[0111] The replacement unit 56X replaces tags representing words belonging to clusters in the set with representative words representing the clusters to which the words belong. The replacement unit 56X can be, for example, the tag replacement / correction unit 56 in the first or second embodiment.

[0112] 11 is an example of a flowchart executed by the data analysis device 1X. The clustering means 52X clusters words represented by tags included in a set of combinations of data names and tags representing attributes of the data names (step S21). The representative word determination means 54X determines a representative word representing each of the clusters generated based on the clustering (step S22). The replacement means 56X replaces tags representing words belonging to a cluster in the set with a representative word representing the cluster to which the word belongs (step S23).

[0113] The data analysis device 1X according to the third embodiment can suitably aggregate tags included in a set of combinations of data names and tags representing attributes of the data names.

[0114] In addition, part or all of the above-described embodiments (including variations, the same applies below) may also be described as, but are not limited to, the following supplementary notes. Furthermore, not only the devices, methods, and storage media described in the supplementary notes, but also various hardware, software, various recording means for recording software, or systems may be made to depend on part or all of the configurations described in the supplementary notes, as long as they do not deviate from the above-described embodiments.

[0115] [Supplementary Note 1] A data analysis device comprising: clustering means for clustering words represented by tags included in a set of combinations of data names and tags representing attributes of the data names; representative word determination means for determining a representative word to represent each of the clusters generated based on the clustering; and replacement means for replacing, in the set, tags representing words belonging to the clusters with the representative word that represents the cluster to which the word belongs. [Supplementary Note 2] The data analysis device of Supplementary Note 1, further comprising: generation means for generating frequency information regarding the appearance frequency of each of the words in the set, wherein the clustering means filters words to be clustered based on the frequency information. [Supplementary Note 3] The data analysis device of Supplementary Note 1, further comprising: generation means for generating classification information based on the part of speech or spelling of each of the words in the set, wherein the clustering means performs the clustering for each group of the words divided based on the classification information. [Supplementary Note 4] The data analysis device of Supplementary Note 1, further comprising cluster filtering means for excluding or modifying clusters that do not satisfy predetermined conditions from the clusters generated by the clustering. [Supplementary Note 5] The data analysis device of Supplementary Note 4, wherein the cluster filtering means sets, as the predetermined condition, a condition related to the number of words in the cluster, or a condition related to the variance within the cluster of word vectors obtained by vectorizing the words. [Supplementary Note 6] The data analysis device of Supplementary Note 1, wherein the representative word determination means generates the representative word for each of the clusters based on a language model that has been machine-learned to generate an answer to a fill-in-the-blank question when given, or a language model that has been machine-learned to generate a word that represents a set of words when given. [Supplementary Note 7] The data analysis device of Supplementary Note 1, further comprising word filtering means that, based on the representative word, filters out words that are not suitable to belong to the cluster represented by the representative word.[Supplementary Note 8] The data analysis device of Supplementary Note 1, wherein the replacement means corrects the replacement based on a search result of the data name associated with the tag for which replacement has been performed and the representative word replaced by the replacement. [Supplementary Note 9] The data analysis device of Supplementary Note 1, wherein after replacement with the representative word, a prompt asking whether the representative word is appropriate as the tag representing an attribute of the data name is input to a language model that has been machine-learned to output a response to the prompt when input, and the replacement is corrected based on the response output by the language model in response to the input. [Supplementary Note 10] The data analysis device of Supplementary Note 1, wherein the clustering means performs the clustering to generate second clusters for words that do not belong to any of the clusters for which the representative word has been determined, and the replacement means replaces tags representing words belonging to the second clusters with the representative word determined for each of the second clusters. [Supplementary Note 11] The data analysis device of Supplementary Note 1, further comprising display control means for displaying information about the set after replacement with the representative word on a display device. [Supplementary Note 12] The data analysis device according to Supplementary Note 11, wherein the display control means accepts an input specifying the set and displays the information on the display device to support the decision-making of a user who makes the input. [Supplementary Note 13] A data analysis method, wherein a computer performs clustering of words represented by tags included in a set of combinations of data names and tags representing attributes of the data names, determines a representative word to represent each of the clusters generated based on the clustering, and replaces, in the set, tags representing words belonging to the clusters with the representative word to represent the cluster to which the words belong. [Supplementary Note 14] A storage medium storing a program that causes a computer to execute processes of clustering words represented by tags included in a set of combinations of data names and tags representing attributes of the data names, determines a representative word to represent each of the clusters generated based on the clustering, and replaces, in the set, tags representing words belonging to the clusters with the representative word to represent the cluster to which the words belong.[Supplementary Note 15] The data analysis method according to Supplementary Note 13, generating frequency information regarding the frequency of appearance of each of the words in the set, and filtering words to be clustered based on the frequency information. [Supplementary Note 16] The data analysis method according to Supplementary Note 13, generating classification information based on the part of speech or spelling of each of the words in the set, and performing the clustering for each group of the words separated based on the classification information. [Supplementary Note 17] The data analysis method according to Supplementary Note 13, excluding or correcting clusters that do not satisfy a predetermined condition from among the clusters generated by the clustering. [Supplementary Note 18] The data analysis method according to Supplementary Note 17, setting a condition regarding the number of words in the cluster, or a condition regarding the variance within the cluster of word vectors obtained by vectorizing the words, as the predetermined condition. [Supplementary Note 19] The data analysis method of Supplementary Note 13, wherein the representative word for each of the clusters is generated based on a language model that has been machine-learned to generate an answer to a fill-in-the-blank question when given, or a language model that has been machine-learned to generate a word that represents a set of words when given. [Supplementary Note 20] The data analysis method of Supplementary Note 13, wherein words that are not suitable to belong to the cluster represented by the representative word are filtered out based on the representative word. [Supplementary Note 21] The data analysis method of Supplementary Note 13, wherein the replacement is corrected based on a search result of the data name associated with the tag that has been replaced and the representative word that has been replaced by the replacement. [Supplementary Note 22] The data analysis method of Supplementary Note 13, wherein after the replacement with the representative word, a prompt asking whether the representative word is suitable as the tag representing an attribute of the data name is input to a language model that has been machine-learned to output an answer to the prompt when a prompt is input, and the replacement is corrected based on the answer output by the language model in response to the input.[Supplementary Note 23] The data analysis method according to Supplementary Note 13, wherein the clustering is performed to generate second clusters for words that do not belong to any of the clusters for which the representative word has been determined, and wherein tags representing words that belong to the second clusters are replaced with the representative word determined for each of the second clusters. [Supplementary Note 24] The data analysis method according to Supplementary Note 13, wherein information about the set after replacement with the representative word is displayed on a display device. [Supplementary Note 25] The data analysis method according to Supplementary Note 24, wherein an input specifying the set is received, and the information is displayed on the display device to support the decision-making of a user who made the input. [Supplementary Note 26] The storage medium according to Supplementary Note 14, wherein the program is stored that causes the computer to execute a process of generating frequency information regarding the frequency of appearance of each of the words in the set, and filtering words to be clustered based on the frequency information. [Supplementary Note 27] The storage medium according to Supplementary Note 14, wherein the program is stored that causes the computer to execute a process of generating classification information based on the part of speech or spelling of each of the words in the set, and performing the clustering for each group of the words divided based on the classification information. [Supplementary Note 28] The storage medium according to Supplementary Note 14, which stores the program causing the computer to execute a process of excluding or correcting clusters that do not satisfy a predetermined condition from among the clusters generated by the clustering. [Supplementary Note 29] The storage medium according to Supplementary Note 28, which stores the program causing the computer to execute a process of setting, as the predetermined condition, a condition related to the number of words in the cluster, or a condition related to the variability within the cluster of word vectors obtained by vectorizing the words. [Supplementary Note 30] The storage medium according to Supplementary Note 14, which stores the program causing the computer to execute a process of generating the representative word for each of the clusters, based on a language model that has been machine-learned to generate an answer to a fill-in-the-blank question when given, or a language model that has been machine-learned to generate a word that represents a set of words when given.[Supplementary Note 31] The storage medium according to Supplementary Note 14, which stores the program causing the computer to execute a process of filtering out words that are unsuitable for belonging to the cluster represented by the representative word, based on the representative word. [Supplementary Note 32] The storage medium according to Supplementary Note 14, which stores the program causing the computer to execute a process of correcting the replacement, based on search results of the data name associated with the tag that has been replaced and the representative word that has been replaced by the replacement. [Supplementary Note 33] The storage medium according to Supplementary Note 14, which stores the program causing the computer to execute a process of correcting the replacement, based on the answer output by the language model in response to the input, by inputting a prompt asking about the suitability of the representative word as the tag that represents an attribute of the data name into a language model that has been machine-learned so as to output an answer to the prompt when the prompt is input, after the replacement with the representative word. [Supplementary Note 34] The storage medium of Supplementary Note 14, storing the program that causes the computer to execute the process of performing the clustering to generate second clusters for words that do not belong to any of the clusters for which the representative words have been determined, and replacing tags representing words that belong to the second clusters with the representative words determined for each of the second clusters. [Supplementary Note 35] The storage medium of Supplementary Note 14, storing the program that causes the computer to execute the process of displaying, on a display device, information about the set after replacement with the representative word. [Supplementary Note 36] The storage medium of Supplementary Note 35, storing the program that causes the computer to execute the process of accepting an input specifying the set, and displaying the information on the display device to support the decision-making of a user who made the input.

[0116] In each of the above-described embodiments, the program can be stored using various types of non-transitory computer-readable media and supplied to a computer processor, etc. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic storage media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical storage media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, semiconductor memories (e.g., mask ROMs, programmable ROMs (PROMs), erasable PROMs (EPROMs), flash ROMs, and random access memories (RAMs). The program may also be supplied to a computer by various types of transient computer-readable media. Examples of transient computer-readable media include electric signals, optical signals, and electromagnetic waves. The transient computer-readable medium can supply the program to a computer via a wired communication path such as an electric wire or optical fiber, or via a wireless communication path.

[0117] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above embodiments. Various modifications within the scope of the present invention that would be understood by those skilled in the art can be made to the configuration and details of the present invention. In other words, the present invention naturally includes various modifications and alterations that would be possible for those skilled in the art based on the entire disclosure, including the claims, and the technical ideas. Furthermore, the disclosures of the above-cited patent and non-patent documents are incorporated herein by reference.

[0118] 1, 1A, 1X Data analysis device 2 Input device 3 Display device 4 Storage device 5 Terminal device 11 Processor 12 Memory 13 Interface 100, 100A Data analysis system

Claims

1. A data analysis apparatus comprising: clustering means for performing clustering of words represented by tags included in a set of combinations of data names and tags representing attributes of the data names; representative word determination means for determining a representative word representing each of the clusters generated based on the clustering; and replacement means for replacing, in the set, tags representing words belonging to the cluster with the representative word representing the cluster to which the word belongs.

2. The data analysis apparatus according to claim 1, further comprising generation means for generating frequency information regarding the frequency of occurrence of each of the words in the set, wherein the clustering means performs filtering of the words for which the clustering is to be performed based on the frequency information.

3. The data analysis apparatus according to claim 1, further comprising generation means for generating classification information based on the part of speech or notation of each of the words in the set, wherein the clustering means performs the clustering for each group of the words classified based on the classification information.

4. The data analysis apparatus according to claim 1, further comprising cluster filtering means for excluding or correcting clusters that do not satisfy a predetermined condition among the clusters generated by the clustering.

5. The data analysis apparatus according to claim 4, wherein the cluster filtering means sets, as the predetermined condition, a condition regarding the number of words in the cluster or a condition regarding the variation within the cluster of word vectors obtained by vectorizing the words.

6. The data analysis apparatus according to claim 1, wherein the representative word determination means generates the representative word for each of the clusters based on a language model in which machine learning is performed to generate an answer to a fill-in-the-blank problem when the fill-in-the-blank problem is given, or a language model in which machine learning is performed to generate a word representing a set of words when the set of words is given.

7. The data analysis apparatus according to claim 1, further comprising word filtering means for filtering words that are not suitable for belonging to the cluster represented by the representative word based on the representative word.

8. The data analysis apparatus according to claim 1, wherein the replacement means corrects the replacement based on a search result of the data name associated with the tag for which the replacement has been made and the representative word replaced by the replacement.

9. After replacement with the representative word, a prompt is input to a language model in which machine learning has been performed to output an answer to the prompt, and the suitability of the representative word as the tag representing the attribute of the data name is questioned. Based on the answer output by the language model in response to the input, the replacement is corrected. The data analysis apparatus according to claim 1.

10. The clustering means performs clustering to generate a second cluster for words that do not belong to any of the clusters in which the representative word has been determined. The replacement means performs replacement of tags representing words belonging to the second cluster with the representative word determined for each of the second clusters. The data analysis apparatus according to claim 1.

11. The data analysis apparatus according to claim 1, further comprising display control means for causing a display device to display information regarding the set after replacement with the representative word.

12. The display control means receives an input for designating the set and causes the display device to display the information in order to assist the decision-making of the user who makes the input. The data analysis apparatus according to claim 11.

13. A computer performs clustering of words represented by tags included in a set of combinations of data names and tags representing attributes of the data names, determines a representative word representing each of the clusters generated based on the clustering, and in the set, replaces tags representing words belonging to the cluster with the representative word representing the cluster to which the word belongs. A data analysis method.

14. A storage medium storing a program for causing a computer to perform clustering of words represented by tags included in a set of combinations of data names and tags representing attributes of the data names, determine a representative word representing each of the clusters generated based on the clustering, and in the set, replace tags representing words belonging to the cluster with the representative word representing the cluster to which the word belongs.

15. Frequency information regarding the frequency of occurrence of each of the words in the set is generated, and filtering of words for which the clustering is to be performed is performed based on the frequency information. The data analysis method according to claim 13.

16. Generate classification information based on the part of speech or notation of each word in the set, and perform the clustering for each group of words separated based on the classification information, the data analysis method according to claim 13.

17. Exclude or correct clusters that do not meet a predetermined condition among the clusters generated by the clustering, the data analysis method according to claim 13.

18. Set a condition regarding the number of words in the cluster or a condition regarding the variation within the cluster of the word vectors obtained by vectorizing the words as the predetermined condition, the data analysis method according to claim 17.

19. Generate each representative word of the clusters based on a language model in which machine learning is performed to generate an answer to the fill-in-the-blank problem when the fill-in-the-blank problem is given, or a language model in which machine learning is performed to generate a word representing the set of words when the set of words is given, the data analysis method according to claim 13.

20. Perform filtering of words that are not suitable to belong to the cluster represented by the representative word based on the representative word, the data analysis method according to claim 13.

Citation Information

Patent Citations

  • Content classification system, server, terminal device, program, and recording medium

    JP2008242689A

  • Data substitution device, data substitution method, and program

    WO2020184126A1