Data analysis device, data analysis method, and storage medium
The data analysis apparatus and method address the challenge of accurately generating tags for target data by integrating search results and language model outputs to improve the accuracy and response rate of tag generation.
Patent Information
- Application Number
- PCT/JP2023/043129
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-01
- Publication Date
- 2025-06-05
AI Technical Summary
Existing data analysis systems using large language models face challenges in accurately generating tags that represent the characteristics of target data, due to the limitations of the data used for learning the models.
A data analysis apparatus and method that utilizes a search result acquisition unit, a prompt generation unit, an answer acquisition unit, and a feature word generation unit to acquire search results, generate prompts for a language model, obtain answers, and generate feature words based on the answers, thereby improving the accuracy of tag generation.
The proposed solution enables the generation of analysis results that accurately represent the characteristics of target data, by leveraging both search results and language model outputs to enhance the response rate and accuracy of tag generation.
Smart Images

Figure JP2023043129_05062025_PF_FP_ABST
Abstract
Description
Data analysis device, data analysis method and storage medium
[0001] The present disclosure relates to the technical fields of a data analysis device, a data analysis method, and a storage medium that perform processing related to data analysis.
[0002] There are systems that perform natural language processing using large language models (LLMs). For example, Patent Literature 1 discloses a language model system that generates a prompt to be input to a large language model, adding valid sentences as reference information for an input question sentence within a set character limit.
[0003] Patent No. 7313757
[0004] To improve the efficiency of data analysis, it is possible to use a large-scale language model when generating tags that characterize the target of analysis. However, since the answers output by a large-scale language model are based on the data used to train the large-scale language model, there is room for improvement in terms of accuracy.
[0005] In view of the above-mentioned problems, one of the objects of the present disclosure is to provide a data analysis device, a data analysis method, and a storage medium that are capable of obtaining analysis results that represent the characteristics of target data.
[0006] One aspect of the data analysis device is a data analysis device having: a search result acquisition means for acquiring search results related to target data; a prompt generation means for generating, based on the search results, first text data representing a question related to the characteristics of the target data as the prompt for a language model that has undergone machine learning so as to output an answer to the prompt when the prompt is input; an answer acquisition means for acquiring, based on the first text data and the language model, second text data representing an answer to the first text data; and a feature word generation means for generating feature words related to the target data based on the second text data.
[0007] One aspect of the data analysis method is a data analysis method in which a computer obtains search results related to target data, generates first text data representing a question regarding characteristics of the target data based on the search results as the prompt for a language model that has been machine-learned to output an answer to the prompt when the prompt is input, obtains second text data representing the answer to the first text data based on the first text data and the language model, and generates characteristic words related to the target data based on the second text data.
[0008] One aspect of the storage medium is a storage medium that stores a program that causes a computer to execute the following processes: obtain search results related to target data; based on the search results, generate first text data representing a question related to characteristics of the target data as a prompt for a language model that has undergone machine learning so as to output an answer to the prompt when the prompt is input; obtain second text data representing an answer to the first text data based on the first text data and the language model; and generate characteristic words related to the target data based on the second text data.
[0009] As an example of an effect of the present disclosure, it is possible to obtain analysis results that represent the characteristics of the target data.
[0010] 1 shows the configuration of a data analysis system; 2 shows the hardware configuration of a data analysis device; 3 shows an example of functional blocks of a processor of the data analysis device; 4 shows an overview of the processing of a data analysis unit; 5 shows a functional block diagram of the data analysis unit; 6 shows an example of a display screen showing the results of attribute estimation processing; 7 is an example of a flowchart showing an overview of processing executed by a data analysis device; 8 shows the configuration of a data analysis system; 9 shows the relationship between a user, a data analysis device, and a terminal device; 10 is a functional block diagram of a data analysis device; 11 is an example of a flowchart showing the processing procedure of a data analysis device
[0011] Hereinafter, embodiments of a data analysis device, a data analysis method, and a storage medium will be described with reference to the drawings. Hereinafter, a "query" refers to a natural language inquiry (including a question or a hypothesis sentence) passed from a user to the data analysis system 100. An "answer" refers to a natural language sentence or text data representing a natural language sentence output by the data analysis system 100 in response to the query.
[0012] <First Embodiment> (1) System Configuration Fig. 1 shows the configuration of a data analysis system 100. The data analysis system 100 is a system that analyzes specified data to estimate attributes that serve as tags for the data, and mainly includes a data analysis device 1, an input device 2, a display device 3, and a storage device 4.
[0013] The data analysis device 1 performs data analysis to estimate attributes that serve as tags for specified data, and controls the display of information related to the attribute estimation results. Hereinafter, the data that is the target of attribute estimation will also be referred to as "attribute estimation target data." The attribute estimation target data may be text data representing a sentence, or table data representing a set of delimited character strings (tokens). The data analysis device 1 communicates data with the input device 2, display device 3, and storage device 4, respectively, via a communication network or by direct wireless or wired communication.
[0014] The input device 2 is an interface that accepts user input, which is external input, and corresponds to, for example, a touch panel, buttons, a keyboard, a voice input device, etc. The input device 2 supplies input information generated based on the user input to the data analysis apparatus 1.
[0015] The display device 3 is, for example, a display, a projector, or the like, and performs a predetermined display based on the display information supplied from the data analysis device 1 .
[0016] The storage device 4 is a memory that stores various information necessary for the processing performed by the data analysis device 1. The storage device 4 may store, for example, a search database used in the search described below, a search engine program that executes the search, model information (configuration information) for constructing a machine-learned large-scale language model (LLM), and model information for constructing any natural language understanding model such as BERT (Bidirectional Encoder Representations from Transformers) used in natural language processing. The model information includes, for example, various parameters of a machine-learned deep learning model, such as the layer structure, the neuron structure of each layer, the number and filter size of filters in each layer, and the weight of each element of each filter. The search database may be web information on the Internet, which is a search location for general search engines, or may be a database owned by an organization (such as a company) that manages the data analysis system 100.
[0017] The LLM is a natural language processing model trained using a large amount of text data, and receives text data representing sentences as input and outputs text data representing sentences. When text data representing a question is input to the LLM, the LLM outputs text data representing an answer. Hereinafter, the text data input to the LLM will be referred to as a "prompt," and the text data output by the LLM will be referred to as "answer text data." The LLM will also be simply referred to as a language model. In addition to the above-mentioned BERT, a specific example of a language model is a GPT (Generative Pre-Trained Transformer), which predicts a character string that is likely to follow an input character string and outputs a sentence containing the input character string. Other examples of language models include T5 (Text-to-Text Transfer Transformer), RoBERTa (Robustly optimized BERT approach), and ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately).
[0018] The storage device 4 may be a storage device such as a hard disk connected to or built into the data analysis device 1, or may be a storage medium such as a flash memory. The storage device 4 may also be a server device that performs data communication with the data analysis device 1. In this case, the storage device 4 may be composed of multiple server devices.
[0019] The configuration of the data analysis system 100 shown in FIG. 1 is an example, and various modifications may be made to the configuration. For example, the input device 2 and the display device 3 may be configured as an integrated device. In this case, the input device 2 and the display device 3 may be configured as a tablet terminal integrated with the data analysis device 1. The data analysis device 1 may be connected to or have a built-in sound output device, such as a speaker, and output information by sound. The data analysis device 1 may also be configured from multiple devices. In this case, the multiple devices that make up the data analysis device 1 exchange information necessary to execute pre-assigned processing between these multiple devices.
[0020] (2) Hardware Configuration of Data Analysis Apparatus Fig. 2 shows the hardware configuration of the data analysis apparatus 1. The data analysis apparatus 1 includes, as hardware components, a processor 11, a memory 12, and an interface 13. The processor 11, the memory 12, and the interface 13 are connected via a data bus 19.
[0021] The processor 11 executes predetermined processes by executing programs stored in the memory 12. The processor 11 is a processor such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or a TPU (Tensor Processing Unit). The processor 11 may be composed of multiple processors. The processor 11 is an example of a computer.
[0022] The memory 12 is composed of various types of volatile and non-volatile memories, such as RAM (Random Access Memory) and ROM (Read Only Memory). The memory 12 also stores programs for the data analysis apparatus 1 to execute various processes. The memory 12 is also used as a working memory, and temporarily stores information obtained from the storage device 4. The memory 12 may also function as the storage device 4. Similarly, the storage device 4 may also function as the memory 12 of the data analysis apparatus 1. The programs executed by the data analysis apparatus 1 may be stored in a storage medium other than the memory 12.
[0023] The interface 13 is an interface for electrically connecting the data analysis device 1 to other devices. These interfaces may be wireless interfaces such as network adapters for wirelessly transmitting and receiving data to and from other devices, or may be hardware interfaces for connecting to other devices via cables or the like.
[0024] The hardware configuration of the data analysis device 1 is not limited to the configuration shown in Fig. 2. For example, the data analysis device 1 may include at least one of the input device 2 and the display device 3. Furthermore, the data analysis device 1 may be connected to or have a built-in sound output device such as a speaker.
[0025] (3) Processing Overview Fig. 3 shows an example of functional blocks of the processor 11. Functionally, the processor 11 has a data analysis unit 15 and a UI (User Interface) control unit 16. Note that in Fig. 3, blocks that exchange data are connected by solid lines, but the combination of blocks that exchange data is not limited to Fig. 3. The same applies to other functional block diagrams described later.
[0026] The data analysis unit 15 refers to various information stored in the storage device 4, memory 12, etc., performs data analysis on the attribute estimation target data specified by the user, and estimates the attributes of the attribute estimation target data. The data analysis unit 15 then supplies the estimation results on the attributes of the attribute estimation target data (also referred to as "attribute estimation results") to the UI control unit 16. Details of the processing by the data analysis unit 15 will be described later.
[0027] The UI control unit 16 controls the reception of user input and the display of information to be viewed by the user. For example, the UI control unit 16 generates attribute estimation target data based on input information (i.e., external input) supplied from the input device 2 in response to a user operation, and supplies the attribute estimation target data to the data analysis unit 15. In this case, the UI control unit 16 may directly acquire text data input by the user as attribute estimation target data, or may generate attribute estimation target data corresponding to table data specified by the user or a part of the records thereof.
[0028] Furthermore, the UI control unit 16 generates display information based on the attribute estimation result generated by the data analysis unit 15, and controls the display of the display device 3 by supplying the generated display information to the display device 3. Specific processing by the UI control unit 16 will be described later with reference to display examples.
[0029] The components of the data analysis unit 15 and the UI control unit 16 described in FIG. 3 can be realized, for example, by the processor 11 executing a program. Alternatively, the necessary programs may be recorded on any non-volatile storage medium and installed as needed to realize the components. At least some of these components may not necessarily be realized by software programs, but may be realized by any combination of hardware, firmware, and software. At least some of these components may be realized using a user-programmable integrated circuit, such as an FPGA (Field-Programmable Gate Array) or a microcontroller. In this case, the integrated circuit may be used to realize a program consisting of the above components. Furthermore, at least a portion of each component may be configured by an ASSP (Application Specific Standard Product), an ASIC (Application Specific Integrated Circuit), or a quantum processor (quantum computer control chip). In this way, each component may be realized by various hardware. The same applies to other embodiments described below. Furthermore, each of these components may be realized by the cooperation of multiple computers, for example, using cloud computing technology.
[0030] (4) Attribute Estimation Processing Next, the attribute estimation processing executed by the data analysis unit 15 will be described.
[0031] (4-1) Overview FIG. 4 is a diagram illustrating an overview of the attribute estimation process performed by the data analysis unit 15. As shown in FIG. 4, when attribute estimation target data is supplied from the UI control unit 16, the data analysis unit 15 performs the attribute estimation process and estimates attributes to be used as tags representing the characteristics of the attribute estimation target data. In the example of FIG. 4, as an example of text data specifying the attribute estimation target data, text data representing the sentence "Please tell me the characteristics of customers of a restaurant named Restaurant X in Area Y." is provided. In this example, "Area Y," the restaurant name "Restaurant X," which is a proper noun, and "Customer" are the attributes to be estimated. Therefore, "Area Y," the restaurant name "Restaurant X," which is a proper noun, and "Customer" may also be referred to as attribute estimation target data.
[0032] In this case, the data analysis unit 15 generates a search query, which is a query used for a search, from the attribute estimation target data. The search query is a set of character strings. In this example, the data analysis unit 15 extracts "Store X" and "Area Y" as the search query.
[0033] Next, the data analysis unit 15 inputs the search query into the search engine, acquires search results output by the search engine, and further generates data that aggregates the acquired search results (also referred to as "aggregated search result data"). Here, as an example, the data analysis unit 15 generates text data that is a list of words (parts of speech) such as "stone oven bread, sandwich, sliced bread, and croissant" as the aggregated search result data.
[0034] Next, the data analysis unit 15 generates a prompt to be input into the LLM based on the aggregated search result data. Here, the data analysis unit 15 generates text data as the prompt, such as "Please tell me the characteristics of customers at a restaurant named 'Store X' in area Y that is known for its stone oven bread, sandwiches, white bread, and croissants." The above text data is an example of "first text data."
[0035] Next, the data analysis unit 15 inputs the generated prompt to the LLM to obtain answer text data output by the LLM. The answer text data output by the LLM in this case is an example of “second text data.”
[0036] The data analysis unit 15 then generates and modifies characteristic words, which are words that represent the characteristics of the answer sentences indicated by the answer text data, and recognizes the obtained characteristic words as tags that represent the attributes of the attribute estimation target data (more specifically, the store name "Store X," which is a proper noun included in the attribute estimation target data). Here, the data analysis unit 15 generates "woman," "bread," "sweets," "lover," "girls' night out," "casual," and the like as tags that represent the attributes of the attribute estimation target data.
[0037] In addition, in Figure 4, attribute estimation target data related to stores is shown as an example, but attribute estimation target data related to products, services (including healthcare), celebrities (including medical professionals such as doctors and nurses), diseases, patients, and any other searchable items may also be used, not limited to stores.
[0038] Here, a supplementary explanation will be given of the effect of the attribute estimation process described above.
[0039] In general, the output of an LLM is based on the knowledge contained in the LLM (i.e., the learning data used for the machine learning of the LLM). Therefore, if attribute estimation target data related to an unfamous store is directly input to the LLM as a prompt, an accurate answer may not be obtained from the LLM. In consideration of the above, the data analysis device 1 according to the present embodiment utilizes search results to suitably estimate the attributes of the attribute estimation target data, even when attribute estimation target data related to information that is not particularly contained in the learning data used for the machine learning of the LLM is specified.
[0040] Specifically, while LLM has the advantage of being able to generate answers to any query (i.e., having a high response rate), the answers generated by LLM may not be sufficiently accurate because they are inference results. On the other hand, in the case of search, the accuracy of information retrieved in response to a search query is generally high, but the response rate is not as sufficient as that of LLM, for example, when a specific search query is set, the number of hits may be zero. Furthermore, for example, in the case of a store search, basic information about the store (business type and address) may be included in the response, but reputation information such as the store's customer profile may not be included in the response. Taking the above into consideration, the data analysis device 1 in this embodiment utilizes both search and LLM in a complementary manner to estimate attributes with both a high response rate and accuracy.
[0041] 5 is a functional block diagram of the data analysis unit 15. Functionally, the data analysis unit 15 includes a search query generation unit 51, a search result aggregation unit 52, a prompt generation unit 53, an answer acquisition unit 54, a feature word generation / modification unit 55, and a re-execution unit 56.
[0042] The search query generation unit 51 generates a search query, which is a set of character strings, from attribute estimation target data supplied from the UI control unit 16 or the like. In this case, the attribute estimation target data may be data indicating sentences as shown in FIG. 4 , or may be data obtained by converting table data (structured data) into text. The search query generation unit 51 may also generate multiple sets of character strings as multiple search queries. The search query generation unit 51 then supplies the search query to the search result aggregation unit 52.
[0043] The search result aggregation unit 52 acquires search results obtained by searching a search engine using the search query supplied from the search query generation unit 51, and generates search result aggregation data representing an explanation that aggregates the search results. In this case, the search result aggregation unit 52 acquires search results from the search query using, for example, any search engine that searches web information (websites), and aggregates the acquired search results. In this case, the search engine may be executed by the data analysis device 1, or may be executed by a search system other than the data analysis device 1. In the latter case, the data analysis device 1 receives search results from an external device instead of executing the search itself. The search result aggregation unit 52 supplies the generated search result aggregation data to the prompt generation unit 53.
[0044] The prompt generation unit 53 generates a prompt based on the search result aggregation data generated by the search result aggregation unit 52. In this case, the prompt generation unit 53 may generate a prompt using the attribute estimation target data in addition to the search result aggregation data. The prompt generation unit 53 supplies the generated prompt to the answer acquisition unit 54.
[0045] The answer acquisition unit 54 acquires answer text data output by the LLM by inputting the prompt generated by the prompt generation unit 53 to the LLM. For example, if machine-learned configuration information of the LLM is stored in the storage device 4, the answer acquisition unit 54 inputs a prompt to the LLM configured by referring to the machine-learned parameters, etc. indicated in the configuration information, and acquires answer text data output by the LLM in response to the input. The answer acquisition unit 54 supplies the answer text data to the feature word generation / modification unit 55.
[0046] The device that executes the LLM may be an external device capable of data communication with the data analysis device 1. In this case, the answer acquisition unit 54 transmits an execution instruction signal for the LLM, including a prompt, to the external device via the interface 13, and receives a response signal, including answer text data, from the external device via the interface 13. The external device that receives the execution instruction signal transmits a response signal, including answer text data output by the LLM when the prompt included in the execution instruction signal is input to the LLM, to the data analysis device 1. Similarly, when another processing block uses the LLM, the data analysis device 1 may acquire the execution result of the LLM from the external device instead of executing the LLM itself.
[0047] The feature word generation and correction unit 55 generates a set of one or more feature words based on the response text data. In this case, the feature word generation and correction unit 55 also performs search and LLM on the generated set of feature words to correct (including delete) the feature words. The feature word generation and correction unit 55 supplies the generated set of feature words to the re-execution unit 56.
[0048] The re-execution unit 56 determines whether or not it is necessary to re-execute (hereinafter simply referred to as "re-execution") part or all of the series of processes for generating and correcting feature words from attribute estimation target data, based on information output by the search query generation unit 51, the search result aggregation unit 52, the prompt generation unit 53, the answer acquisition unit 54, and / or the feature word generation and correction unit 55. This determination method will be described later. If the re-execution unit 56 determines that re-execution is necessary, it instructs other functional blocks to re-execute the processes. In this case, the re-execution unit 56 may, for example, instruct the search query generation unit 51 to generate a search query so that the feature word generated by the feature word generation and correction unit 55 is included in the search query, or may instruct the prompt generation unit 53 to generate a prompt so that the feature word generated by the feature word generation and correction unit 55 is included in the prompt.
[0049] On the other hand, if the re-execution unit 56 determines that re-execution is unnecessary, it supplies the characteristic words generated most recently by the characteristic word generation / modification unit 55 to the UI control unit 16 as the attribute estimation result of the attribute estimation target data. In this case, the UI control unit 16 may display information representing the attribute estimation result supplied by the characteristic word generation / modification unit 55 on the display device 3. An example of this display will be described later.
[0050] Hereinafter, the processes executed by the search query generation unit 51, search result aggregation unit 52, prompt generation unit 53, feature word generation / modification unit 55, and re-execution unit 56 will be described in detail.
[0051] (4-2) Search Query Generation Unit The generation of a set of character strings that become search queries by the search query generation unit 51 will be described.
[0052] In the first search query generation example, the search query generation unit 51 extracts partial strings from the attribute estimation target data. In this case, if the attribute estimation target data is a sentence, the search query generation unit 51 extracts partial strings from the attribute estimation target data based on any natural language processing such as keyword extraction processing. If the attribute estimation target data is table data, the search query generation unit 51 may extract each token as a partial string. Furthermore, if the number of partial strings set as the search query is a predetermined number, the search query generation unit 51 may randomly select the predetermined number of partial strings or may select the predetermined number of partial strings based on input information generated by the input device 2. Note that the search query generation unit 51 may identify proper nouns (including store names, product names, etc.) whose attributes are to be estimated from the attribute estimation target data based on any natural language processing or user input, and select partial strings for the proper nouns so that they are included in the search query.
[0053] In the second search query generation example, the search query generation unit 51 uses the LLM as a search query generator. For example, the search query generation unit 51 extracts proper nouns from the attribute estimation target data and generates a prompt inquiring about a search query to be set for the proper noun. The search query generation unit 51 then inputs the generated prompt into the LLM and determines the search query based on the response output by the LLM. In this case, for example, template information indicating a template for generating a prompt is stored in the storage device 4 or the like, and the prompt generation unit 53 generates a prompt by referring to the template information. Note that this template information may be prepared for each category (classification) of proper nouns whose attributes are to be estimated. In this case, the prompt generation unit 53 determines the category of the proper nouns included in the attribute estimation target data using natural language processing and generates a prompt using a template corresponding to the determined category.
[0054] For example, in the example of Fig. 3, the search query generation unit 51 generates the following prompt: "When researching a restaurant named 'Store X' in area Y, please create three search query candidates for a search engine that will provide an overview of the restaurant."
[0055] In this case, the search query generation unit 51 determines a search query from three search query candidates generated by the LLM (in this case, three combinations of words that serve as search query candidates). In this case, the search query generation unit 51 may determine a search query based on a frequency count result of words included in the candidates, or may randomly select a predetermined number of words that serve as search queries from the candidate words.
[0056] (4-3) Search Result Aggregation Unit The generation of search result aggregation data by the search result aggregation unit 52 will be described.
[0057] In a first generation method of the aggregated search result data, the search result aggregator 52 performs morphological analysis on the text of the search results to break it down into words for each part of speech, assigns a score based on frequency to each word, and extracts a set of words with a score equal to or greater than a predetermined value as elements of the aggregated search result data. In this case, the aggregated search result data represents a sentence (explanatory text) that lists words with a score equal to or greater than a predetermined value.
[0058] In this case, in a first example, the search result aggregating unit 52 calculates the frequency of occurrence of each word extracted from the search results by counting the number of times each word appears as a score. The search result aggregating unit 52 may then extract the k words (k is an integer equal to or greater than 1) with the highest scores representing the frequency of occurrence as elements of the search result aggregation data, or may extract words with scores representing the frequency of occurrence equal to or greater than a predetermined threshold as elements of the search result aggregation data. The threshold is stored in, for example, the storage device 4 or the memory 12.
[0059] In a second example, the search result aggregating unit 52 uses weighted frequency as a score. In this case, the search result aggregating unit 52 calculates the weighted frequency using, for example, TF-IDF (Term Frequency-Inverse Document Frequency). In this case, the score representing the weighted frequency is a value obtained by multiplying the TF (Term Frequency) value by the IDF (Inverse Document Frequency) value as a weight. The search result aggregating unit 52 may extract the k words with the highest scores representing the weighted frequency as elements of the search result aggregation data, or may extract words with scores representing the weighted frequency equal to or greater than a predetermined threshold as elements of the search result aggregation data. The above-mentioned threshold is stored, for example, in the storage device 4 or the memory 12.
[0060] In a second method for generating search result aggregated data, the search result aggregating unit 52 uses an LLM as a generator of search result aggregated data. In this case, the search result aggregating unit 52 generates a prompt requesting that search results be aggregated, and inputs the generated prompt into the LLM, thereby using text data output by the LLM as the search result aggregated data. In this case, for example, template information indicating a template for generating the prompt is stored in the storage device 4 or the like, and the search result aggregating unit 52 generates the prompt by referring to the template information. This template information may be prepared for each category of proper nouns whose attributes are to be inferred. In this case, the search result aggregating unit 52 generates the prompt using a template corresponding to the category of proper nouns included in the attribute inference target data.
[0061] For example, in the example of FIG. 3, the search query generation unit 51 generates a prompt with the following sentence added to the beginning of the search result:
[0062] "Please use the following sentence as a reference and explain in about 20 words what store "Store X" is."
[0063] In this case, the search query generation unit 51 uses text data indicating a description of about 20 characters generated by the LLM as search result aggregation data.
[0064] The search results to be aggregated may be multiple search results output by the search engine from one search query, or multiple search results output by the search engine from multiple search queries. Furthermore, the search results to be aggregated may include web information of link destinations (URLs) included in the search results output by the search engine from the search query.
[0065] (4-4) Prompt Generator The generation of a prompt by the prompt generator 53 will now be described.
[0066] In a first prompt generation method, the prompt generation unit 53 generates text data that combines attribute estimation target data and search result aggregate data as a prompt. In this case, for example, template information indicating a template for generating a prompt using the attribute estimation target data and the search result aggregate data is stored in the storage device 4 or the like, and the prompt generation unit 53 generates a prompt by referring to the template information. Note that this template information may be prepared for each category of proper nouns whose attributes are to be estimated. In this case, the prompt generation unit 53 generates a prompt using a template corresponding to the category of proper nouns included in the attribute estimation target data.
[0067] For example, when the proper noun included in the attribute estimation target data is a shop, the prompt generation unit 53 generates a prompt using template information indicating the following template used for the shop category.
[0068] "Please tell us in about five words the characteristics of customers at {store name} at {address} that are characterized by {aggregated search results data}."
[0069] In this case, the prompt generation unit 53 substitutes the description indicated by the search result aggregation data for "{search result aggregation data}" and the address and store name ("area Y" and "store X" in the example of FIG. 4) included in the attribute estimation target data for "{address}" and "{store name}", respectively. In this case, the prompt generation unit 53 may perform a process of extracting words (the address and store name in this example) that will be used as elements to be substituted into the template from the attribute estimation target data, based on any natural language processing.
[0070] The above-described template is merely an example, and any template format may be used. For example, the prompt generator 53 may generate a prompt using the following template:
[0071] "#Question: Please describe the characteristics of customers at {Store Name} at {Address} in about five words. #Background: {Store Name} is {Search Results Aggregated Data}."
[0072] Furthermore, the prompt generating unit 53 may generate a prompt based on the attribute estimation target data and the search result aggregate data using any prompt generation method such as RAG (Retrieval Augmented Generation).
[0073] In a second prompt generation method, the prompt generation unit 53 uses an LLM as a prompt generator. In this case, the prompt generation unit 53 generates a prompt requesting the generation of a prompt for input to the LLM suitable for inquiring about features (attributes) related to attribute estimation target data using the aggregated search result data. In this case, for example, template information indicating a template for generating a prompt using the aggregated search result data and the attribute estimation target data is stored in the storage device 4, and the prompt generation unit 53 generates a prompt by referring to the template information. Note that this template information may be prepared for each category of proper nouns whose attributes are to be estimated.
[0074] (4-5) Feature Word Generation and Modification Unit First, the generation of feature words by the feature word generation and modification unit 55 will be described.
[0075] In the first method for generating feature words, the feature word generation and modification unit 55 performs morphological analysis on the answer sentences indicated by the answer text data output by the LLM to break them down into words for each part of speech, assigns a score based on frequency to each word, and extracts a set of words with scores equal to or greater than a predetermined value as feature words. In this case, the feature word generation and modification unit 55 calculates the score for each word, for example, by counting the number of times the word extracted from the search results appears as the score. In another example, the feature word generation and modification unit 55 uses weighted frequency as the score and calculates the weighted frequency using TF-IDF or the like. The feature word generation and modification unit 55 may then extract the k words with the highest scores as feature words, or may extract words whose scores representing weighted frequency are equal to or greater than a predetermined threshold. The threshold is stored, for example, in the storage device 4 or memory 12.
[0076] In the second method for generating feature words, the feature word generation and correction unit 55 uses an LLM as a generator of feature words. In this case, the feature word generation and correction unit 55 generates a prompt requesting the output of feature words (tags) appropriate for the answer sentence indicated by the answer text data, and inputs the generated prompt into the LLM to obtain the feature words output by the LLM.
[0077] In this case, for example, the feature word generation and correction unit 55 inputs a prompt into the LLM in which the following sentence is added after the answer sentence indicated by the answer text data:
[0078] "Please list about five tags that would be suitable for tagging the above sentence."
[0079] In this case, the feature word generation / modification unit 55 acquires approximately five words output by the LLM as tags as feature words.
[0080] In the third method for generating feature words, the feature word generation / modification unit 55 acquires, as feature words, in addition to the feature words obtained by the first method for generating feature words, words whose scores obtained by the first method for generating search result aggregated data are equal to or greater than a predetermined value. In this case, if the search result aggregation unit 52 has executed the first method for generating search result aggregated data, the feature word generation / modification unit 55 may acquire the execution result from the search result aggregation unit 52 and acquire, as feature words, words whose scores are equal to or greater than a predetermined value.
[0081] In the fourth method for generating feature words, the feature word generation and correction unit 55 generates feature words using an LLM based on the answer sentence indicated by the answer text data and the search results acquired by the search result aggregation unit 52. For example, template information indicating a template for generating a prompt using the search results and the answer text data is stored in the storage device 4 or the like, and the feature word generation and correction unit 55 generates the prompt by referring to the template information. This template information may be prepared for each category of proper nouns whose attributes are to be estimated.
[0082] For example, if the attribute estimation target data includes a proper noun of a store, the feature word generation / modification unit 55 inputs the following prompt to the LLM to request the generation of feature words targeting the answer sentence and search results:
[0083] "#Explanation {Answer text data} #Search results {Search results} Based on the above two pieces of information, please tell us about five words that would be suitable for describing {Store name} customers for marketing purposes."
[0084] Here, the answer sentence indicated by the answer text data is inserted into "{answer text data}", and the sentence indicating the search results acquired by the search result aggregator 52 is inserted into "{search result}".
[0085] Next, the modification of the characteristic words by the characteristic word generation and modification unit 55 will be described.
[0086] In a first feature word correction method, the feature word generation / correction unit 55 filters (deletes) feature words based on search results obtained when a search is performed using the feature words as a search query. In a first example, the feature word generation / correction unit 55 performs a search using a search engine using a pair of a proper noun ("Store X" in the example of FIG. 4) of the attribute estimation target data and each feature word as a search query, retaining feature words for which the number of hits is equal to or greater than a predetermined threshold and deleting feature words for which the number of hits is less than the threshold. In a second example, the feature word generation / correction unit 55 sequentially adds each feature word to a search query that includes at least a proper noun of the attribute estimation target data and compares the number of hits before and after the addition. The threshold is stored, for example, in the storage device 4 or memory 12. If the ratio of the number of hits after the addition to the number of hits before the addition is equal to or greater than a predetermined threshold, the feature word generation / correction unit 55 retains the added feature word in the search query. If the ratio is less than the threshold, the feature word generation / correction unit 55 removes the added feature word from the search query. Then, the characteristic word generation / modification unit 55 uses the characteristic words remaining in the search query as modified characteristic words.
[0087] In the second feature word correction method, the feature word generation / correction unit 55 filters the feature words using the LLM. In this case, the feature word generation / correction unit 55 inputs a prompt to the LLM requesting the deletion of inappropriate feature words from the generated feature words. For example, template information indicating a template for generating a prompt using the generated feature words and attribute estimation target data is stored in the storage device 4, etc., and the feature word generation / correction unit 55 generates the prompt by referring to the template information. This template information may be prepared for each category of proper nouns whose attributes are to be estimated.
[0088] For example, if the proper noun of the attribute estimation target data is a shop, the following prompt is generated as a prompt to the LLM:
[0089] "The following tags have been assigned to describe the characteristics of customers of the {business type} {store name} at {address}. {characteristic word set} Please delete the incorrect tags and leave only the correct ones."
[0090] Here, all generated feature words are inserted into {feature word set}, and {address}, {store name}, and {business type} are inserted with the address, store name, and business type of proper nouns identified from the attribute estimation target data (in the example of Figure 4, "Y area," "X store," and "food and beverage business type").
[0091] In addition, when filtering characteristic words using LLM, the characteristic word generation / modification unit 55 may generate prompts that allow not only deletion of characteristic words but also rephrasing, modification, or addition.
[0092] For example, if the proper noun of the attribute estimation target data is a shop, the following prompt is generated as a prompt to the LLM:
[0093] "The following tags have been assigned to describe the characteristics of customers of the {business type} {store name} at {address}. {characteristic word set} Please delete the incorrect tags and leave only the correct ones. You can modify the wording."
[0094] By generating such a prompt, the characteristic word generation / modification unit 55 can obtain correction results for characteristic words that are not fixed to the expression of the generated characteristic words.
[0095] (4-6) Re-execution Unit First, a specific example of determining whether or not it is necessary to re-execute part or all of the series of processes for generating and correcting feature words from attribute estimation target data will be described.
[0096] In a first example of determining whether re-execution is necessary, the re-execution unit 56 determines whether re-execution is necessary based on the number of feature words supplied from the feature word generation / correction unit 55. For example, the re-execution unit 56 determines that re-execution is necessary when the number of feature words supplied from the feature word generation / correction unit 55 is equal to or less than a predetermined threshold number (e.g., three or less), and determines that re-execution is unnecessary when the number of feature words is greater than the predetermined threshold number. The above-mentioned threshold is stored, for example, in the storage device 4 or the memory 12. In a second example of determining whether re-execution is necessary, the re-execution unit 56 determines whether re-execution is necessary based on the output result of any processing block other than the feature word generation / correction unit 55. For example, the re-execution unit 56 determines that re-execution is necessary when the number of characters in the answer sentence indicated by the answer text data acquired by the answer acquisition unit 54 is equal to or less than a threshold, and determines that re-execution is unnecessary when the number of characters is greater than the threshold. The above-mentioned threshold is stored, for example, in the storage device 4 or the memory 12. In another example, the re-execution unit 56 determines that re-execution is necessary when the number of search results (the number of hits by the search engine) obtained by the search result aggregating unit 52 is equal to or less than a threshold, and determines that re-execution is not necessary when the number of characters is greater than the threshold. The threshold is stored in, for example, the storage device 4 or the memory 12.
[0097] Next, a specific embodiment of re-execution will be described. When the re-execution unit 56 determines that re-execution is necessary, it may instruct the re-execution unit 56 to start re-execution from any process selected from the search query generation process by the search query generation unit 51, the search result aggregation data generation process by the search result aggregation unit 52, the prompt generation process by the prompt generation unit 53, the feature word generation process by the feature word generation / correction unit 55, or the feature word correction process by the feature word generation / correction unit 55. In this case, the data analysis device 1 preferably changes the method and / or parameters (such as thresholds) randomly or according to a predetermined rule and re-executes each process. In the re-execution, a series of processes is executed, from the process that starts the re-execution to the feature word correction process by the feature word generation / correction unit 55. This enables the feature word generation / correction unit 55 to generate a set of feature words that is different from the set of feature words before the re-execution.
[0098] For example, when re-execution of the process from generating a prompt to generating and correcting a characteristic word is performed, the re-execution unit 56 supplies the characteristic word most recently generated and corrected by the characteristic word generation and correction unit 55 to the prompt generation unit 53, and the prompt generation unit 53 generates a prompt that includes at least the characteristic word. In this case, the prompt generation unit 53 may generate a prompt by setting the characteristic word, instead of the search result aggregate data, as a term that represents a characteristic of the proper noun in the attribute estimation target data, or may generate a prompt by setting both the search result aggregate data and the characteristic word as a characteristic of the proper noun in the attribute estimation target data.
[0099] For example, when the proper noun included in the attribute estimation target data is a shop, the prompt generation unit 53 generates the following prompt:
[0100] "Please tell me the characteristics of customers at {store name} at {address} that are characterized by {set of characteristic words}."
[0101] All feature words supplied by the re-execution unit 56 are inserted into {feature word set}. This generates a prompt that is different from the prompt generated by the prompt generation unit 53 before re-execution. As a result, the feature word generation / modification unit 55 can generate a set of feature words that is different from the set of feature words before re-execution.
[0102] As another example, when the processes from the search query generation process to the feature word generation / modification process are re-executed, the search query generation unit 51 sets a search query that includes at least the feature words supplied from the re-execution unit 56. This generates a search query that is different from the search query generated by the search query generation unit 51 before the re-execution. As a result, the feature word generation / modification unit 55 can generate a set of feature words that is different from the set of feature words before the re-execution.
[0103] (5) Display Example Fig. 6 is an example of a display screen showing the results of the attribute estimation process. When the re-execution unit 56 determines that re-execution is unnecessary, the UI control unit 16 generates a display signal to be supplied to the display device 3 based on each processing result supplied from the data analysis unit 15, and supplies the display signal to the display device 3, thereby causing the display device 3 to display the display screen shown in Fig. 6. The UI control unit 16 provides a query specification field 61, an attribute estimation result display field 62, and an answer statement display field 63 on the display screen.
[0104] The query specification field 61 is an input GUI for a user to specify attribute estimation target data, and includes a text input field 611, an execute button 612, and a table selection button 613. The text input field 611 is an input field that accepts attribute estimation target data input (such as keyboard input or voice input) by the user using the input device 2. When the UI control unit 16 detects that the execute button 612 has been selected, it generates attribute estimation target data representing the text entered in the text input field 611. On the other hand, when it detects that the table selection button 613 has been selected, the UI control unit 16 further displays a GUI for selecting a table stored in the storage device 4, the memory 12, or the like, and accepts user input specifying the table (and record). The UI control unit 16 then generates table data corresponding to the table (and record) specified by the user input as attribute estimation target data.
[0105] The attribute estimation result display field 62 is a display field that displays the estimation result by the data analysis unit 15, and has tags 621 and an input field for adding tags 622. The tags 621 are feature words extracted by the data analysis unit 15 as the attribute estimation result, and here, tags 621 corresponding to six feature words are arranged. Each tag 621 is provided with a delete button, allowing any tag 621 to be deleted based on user input. The input field for adding tags 622 is a GUI (here, an input field and an add button) for generating tags 621 based on user input, and when the add button is selected, the character string entered in the input field is generated as the tag 621.
[0106] The answer sentence display field 63 is a field for displaying an answer sentence indicated by the answer text data generated by the answer acquisition unit 54. The UI control unit 16 acquires the answer text data generated by the answer acquisition unit 54 from the data analysis unit 15, and displays the answer sentence indicated by the answer text data in the answer sentence display field 63.
[0107] According to the display example shown in Figure 6, tags indicating the attribute estimation results for the query (data subject to attribute estimation) specified by the user are presented, thereby making it possible to effectively support the user's data analysis and decision-making based on the data analysis results.
[0108] (6) Processing Flow FIG. 7 is an example of a flowchart showing an outline of the processing executed by the data analysis device 1.
[0109] First, the data analysis device 1 identifies attribute estimation target data based on input information etc. supplied by the input device 2, and generates a search query based on the identified attribute estimation target data (step S11). The processing of step S11 corresponds to processing executed by the UI control unit 16 and the search query generation unit 51.
[0110] Next, the data analysis device 1 aggregates search results based on the search query generated in step S11 (step S12). In this case, the data analysis device 1 generates search result aggregate data that aggregates search results obtained by executing the search engine using the search query. The processing of step S12 corresponds to the processing executed by the search result aggregation unit 52.
[0111] Next, the data analysis device 1 generates a prompt based on the search result aggregate data obtained in step S12 (step S13). The process of step S13 corresponds to the process executed by the prompt generation unit 53.
[0112] Next, the data analysis device 1 inputs the prompt generated in step S13 to the LLM and acquires answer text data generated by the LLM in response to the input (step S14). The processing of step S14 corresponds to the processing executed by the answer acquisition unit 54.
[0113] Next, the data analysis device 1 generates feature words based on the response text data acquired in step S14 (step S15). The data analysis device 1 also corrects (filters) the feature words generated in step S15 (step S16). The processes of steps S15 and S16 correspond to the processes executed by the feature word generation / correction unit 55.
[0114] Next, the data analysis apparatus 1 determines whether re-execution is necessary (step S17). In this case, the data analysis apparatus 1 determines whether re-execution is necessary based on the processing result of step S16 or the processing result of other steps. If it is determined that re-execution is necessary (step S17; Yes), the data analysis apparatus 1 returns to one of steps S11 to S16. In this case, the data analysis apparatus 1 may use the feature words obtained in the immediately preceding step S16. Furthermore, the data analysis apparatus 1 preferably changes at least one of the parameters or algorithms (methods) used in the processing of each step from the previous processing of the same step. This enables the data analysis apparatus 1 to obtain feature words in step S16 that are different from those used previously.
[0115] On the other hand, if it is determined that re-execution is not necessary (step S17; No), the data analysis device 1 displays information representing the attribute estimation results for the attribute estimation target data based on the feature words on the display device 3 (step S18). In this case, for example, the data analysis device 1 displays the set of corrected feature words obtained in the last execution of step S16 as the attribute estimation results for the attribute estimation target data on the display device 3.
[0116] 8 shows the configuration of a data analysis system 100A. The data analysis system 100A mainly includes a data analysis device 1A and a terminal device 5. The data analysis device 1A and the terminal device 5 perform data communication via a network 6.
[0117] The data analysis apparatus 1A is one or more devices that function as a server (including a cloud server) and performs the processing executed by the data analysis apparatus 1 in the first embodiment. In this case, the data analysis apparatus 1A receives input information from the terminal apparatus 5 via the network 6, which the data analysis apparatus 1 receives from the input device 2 in the first embodiment. The data analysis apparatus 1A also transmits display information that the data analysis apparatus 1 transmitted to the display device 3 in the first embodiment to the terminal apparatus 5 via the network 6. The data analysis apparatus 1A also includes the storage apparatus 4 of the first embodiment, or references various pieces of information stored in the storage apparatus 4 via the network 6.
[0118] The terminal device 5 is a terminal having an input function, a display function, and a communication function, and functions as the input device 2 and the display device 3 in the first embodiment. The terminal device 5 may be, for example, a personal computer, a tablet terminal, a PDA (Personal Digital Assistant), or the like. The terminal device 5 transmits input information generated based on the received user input to the data analysis device 1A via the network 6. Furthermore, when the terminal device 5 receives display information from the data analysis device 1A, it displays information based on the display information.
[0119] The data analysis device 1A according to the second embodiment can preferably perform the input process and output process that the data analysis device 1 according to the first embodiment performs on the user of the terminal device 5 .
[0120] 9 is a diagram showing the relationship between a user, a data analysis device 1A, and a terminal device 5. In this case, the data analysis device 1A functions as a server that executes processing related to attribute estimation for specified attribute estimation target data, and the terminal device 5 functions as a user terminal that accepts input related to the attribute estimation target data, etc. The terminal device 5 exchanges information with the data analysis device 1A to present a display screen such as that shown in FIG. 6 to the user. This can favorably prompt the user to make a decision.
[0121] 10 is a functional block diagram of a data analysis device 1X. The data analysis device 1X mainly includes a search result acquisition unit 52X, a prompt generation unit 53X, an answer acquisition unit 54X, and a feature word generation unit 55X. The data analysis device 1X may be composed of multiple devices.
[0122] The search result acquisition unit 52X acquires search results related to the target data. The attribute estimation target data in the first or second embodiment is an example of "target data." The search result acquisition unit 52X may be the search result aggregation unit 52 in the first or second embodiment.
[0123] The prompt generation means 53X generates first text data representing a question regarding the characteristics of the target data as a prompt for the language model based on the search results. Here, the language model is machine-trained to output an answer to the prompt when the prompt is input. The LLM in the first or second embodiment is an example of a "language model." The prompt generation means 53X can be the prompt generation unit 53 in the first or second embodiment.
[0124] The answer acquisition unit 54X acquires second text data representing an answer to the first text data based on the first text data and the language model. The answer acquisition unit 54X is an example of the answer acquisition unit 54 in the first or second embodiment.
[0125] The feature word generating means 55X generates feature words related to the target data based on the second text data. The feature word generating means 55X is an example of the feature word generating and correcting unit 55 in the first or second embodiment.
[0126] 11 is an example of a flowchart executed by the data analysis apparatus 1X. The search result acquisition means 52X acquires search results related to the target data (step S21). The prompt generation means 53X generates, based on the search results, first text data representing a question related to the characteristics of the target data as a prompt for a language model that has undergone machine learning so as to output an answer to the prompt when the prompt is input (step S22). The answer acquisition means 54X acquires, based on the first text data and the language model, second text data representing an answer to the first text data (step S23). The feature word generation means 55X generates feature words related to the target data based on the second text data (step S24).
[0127] The data analysis device 1X according to the third embodiment can generate characteristic words that suitably represent target data.
[0128] In addition, part or all of the above-described embodiments (including variations, the same applies below) may also be described as, but are not limited to, the following supplementary notes. Furthermore, not only the devices, methods, and storage media described in the supplementary notes, but also various hardware, software, various recording means for recording software, or systems may be made to depend on part or all of the configurations described in the supplementary notes, as long as they do not deviate from the above-described embodiments.
[0129] [Supplementary Note 1] A data analysis device comprising: a search result acquisition means for acquiring search results related to target data; a prompt generation means for generating, based on the search results, first text data representing a question related to characteristics of the target data as the prompt for a language model that has been machine-learned so as to output an answer to the prompt when the prompt is input; an answer acquisition means for acquiring, based on the first text data and the language model, second text data representing an answer to the first text data; and a feature word generation means for generating feature words related to the target data based on the second text data. [Supplementary Note 2] The data analysis device according to Supplementary Note 1, further comprising: a feature word correction means for correcting the feature words based on search results related to the feature words. [Supplementary Note 3] The data analysis device according to Supplementary Note 1, further comprising: a feature word correction means for generating, as the prompt, third text data for requesting correction of the feature words, and for obtaining, based on the third text data and the language model, the feature words resulting from the correction of the feature words. [Supplementary Note 4] The data analysis device of Supplementary Note 1, wherein the search result acquisition means generates search result aggregate data representing an explanatory sentence that aggregates the search results, and the prompt generation means generates the first text data based on the search result aggregate data. [Supplementary Note 5] The data analysis device of Supplementary Note 4, wherein the search result acquisition means generates the search result aggregate data representing the explanatory sentence using words extracted from words included in the search results based on the frequency of appearance of the words. [Supplementary Note 6] The data analysis device of Supplementary Note 4, wherein the search result acquisition means generates the search result aggregate data based on the search results and the language model. [Supplementary Note 7] The data analysis device of Supplementary Note 1, wherein the feature word generation means selects the feature words from the words based on the frequency of appearance of the words included in the second text data. [Supplementary Note 8] The data analysis device of Supplementary Note 1, wherein the feature word generation means generates the feature words based on the second text data and the language model.[Supplementary Note 9] The data analysis device according to Supplementary Note 1, further comprising a re-execution means for instructing the execution of at least one of a process of re-acquiring the search results based on the feature words and generating the feature words based on the re-acquired search results, or a process of re-generating the first text data based on the feature words and generating the feature words based on the re-generated first text data. [Supplementary Note 10] The data analysis device according to Supplementary Note 9, wherein the re-execution means determines whether or not the execution is necessary based on the number of feature words. [Supplementary Note 11] The data analysis device according to Supplementary Note 1, further comprising a display control means for displaying, on a display device, information representing an estimation result of an attribute related to the target data based on the feature words. [Supplementary Note 12] The data analysis device according to Supplementary Note 11, wherein the display control means receives an input specifying the target data and displays the information on the display device to support decision-making by a user who inputs the input. [Supplementary Note 13] A data analysis method in which a computer acquires search results for target data, generates first text data representing a question regarding characteristics of the target data based on the search results, as the prompt for a language model that has been machine-learned to output an answer to the prompt when a prompt is input, acquires second text data representing an answer to the first text data based on the first text data and the language model, and generates feature words for the target data based on the second text data. [Supplementary Note 14] A storage medium storing a program that causes a computer to execute the following processes: acquire search results for target data, generates first text data representing a question regarding characteristics of the target data based on the search results, as the prompt for a language model that has been machine-learned to output an answer to the prompt when a prompt is input, acquires second text data representing the answer to the first text data based on the first text data and the language model, and generates feature words for the target data based on the second text data. [Supplementary Note 15] A data analysis method in which the computer modifies the feature words based on the search results for the feature words.[Supplementary Note 16] The data analysis method of Supplementary Note 13, which generates third text data as the prompt, requesting correction of the feature word, and obtains a feature word obtained by correcting the feature word based on the third text data and the language model. [Supplementary Note 17] The data analysis method of Supplementary Note 13, which generates search result aggregated data representing an explanatory sentence summarizing the search results, and generates the first text data based on the search result aggregated data. [Supplementary Note 18] The data analysis method of Supplementary Note 17, which generates search result aggregated data representing the explanatory sentence using words extracted from words included in the search results based on the frequency of appearance of the words. [Supplementary Note 19] The data analysis method of Supplementary Note 17, which generates the search result aggregated data based on the search results and the language model. [Supplementary Note 20] The data analysis method of Supplementary Note 13, which selects the feature word from the words based on the frequency of appearance of words included in the second text data. [Supplementary Note 21] The data analysis method of Supplementary Note 13, which generates the feature word based on the second text data and the language model. [Supplementary Note 22] The data analysis method of Supplementary Note 13, which instructs the execution of at least one of a process of re-acquiring the search results based on the characteristic words and generating the characteristic words based on the re-acquired search results, or a process of re-generating the first text data based on the characteristic words and generating the characteristic words based on the re-generated first text data. [Supplementary Note 23] The data analysis method of Supplementary Note 22, which determines whether or not the execution is necessary based on the number of characteristic words. [Supplementary Note 24] The data analysis method of Supplementary Note 13, which causes a display device to display information representing an estimation result of attributes related to the target data based on the characteristic words. [Supplementary Note 25] The data analysis method of Supplementary Note 24, which accepts input specifying the target data and displays the information on the display device to support decision-making by a user who makes the input. [Supplementary Note 26] The storage medium of Supplementary Note 14, having stored therein the program for modifying the characteristic words based on search results related to the characteristic words.[Appendix 27] The storage medium of Appendix 14, storing the program, generates third text data as the prompt, requesting correction of the feature word, and obtains a feature word obtained by correcting the feature word based on the third text data and the language model. [Appendix 28] The storage medium of Appendix 14, storing the program, generates search result aggregation data representing an explanatory sentence that aggregates the search results, and generates the first text data based on the search result aggregation data. [Appendix 29] The storage medium of Appendix 28, storing the program, generates search result aggregation data representing the explanatory sentence using words extracted from words included in the search results based on the frequency of appearance of the words. [Appendix 30] The storage medium of Appendix 28, storing the program, generates the search result aggregation data based on the search results and the language model. [Appendix 31] The storage medium of Appendix 14, storing the program, selects the feature word from the words based on the frequency of appearance of words included in the second text data. [Supplementary Note 32] The storage medium of Supplementary Note 14, in which the program for generating the feature words based on the second text data and the language model is stored. [Supplementary Note 33] The storage medium of Supplementary Note 14, in which the program for instructing the execution of at least one of a process for re-acquiring the search results based on the feature words and generating the feature words based on the re-acquired search results, or a process for re-generating the first text data based on the feature words and generating the feature words based on the re-generated first text data, is stored. [Supplementary Note 34] The storage medium of Supplementary Note 33, in which the program for determining whether the execution is necessary based on the number of feature words is stored. [Supplementary Note 35] The storage medium of Supplementary Note 14, in which the program for displaying information representing the estimation results of attributes related to the target data based on the feature words on a display device is stored. [Supplementary Note 36] The storage medium of Supplementary Note 35, in which the program for receiving an input specifying the target data and displaying the information on the display device to support the decision-making of a user who made the input is stored.
[0130] In each of the above-described embodiments, the program can be stored using various types of non-transitory computer-readable media and supplied to a computer processor, etc. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic storage media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical storage media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, semiconductor memories (e.g., mask ROMs, programmable ROMs (PROMs), erasable PROMs (EPROMs), flash ROMs, and random access memories (RAMs). The program may also be supplied to a computer by various types of transient computer-readable media. Examples of transient computer-readable media include electric signals, optical signals, and electromagnetic waves. The transient computer-readable medium can supply the program to a computer via a wired communication path such as an electric wire or optical fiber, or via a wireless communication path.
[0131] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above embodiments. Various modifications within the scope of the present invention that would be understood by those skilled in the art can be made to the configuration and details of the present invention. In other words, the present invention naturally includes various modifications and alterations that would be possible for those skilled in the art based on the entire disclosure, including the claims, and the technical ideas. Furthermore, the disclosures of the above-cited patent and non-patent documents are incorporated herein by reference.
[0132] 1, 1A, 1X Data analysis device 2 Input device 3 Display device 4 Storage device 5 Terminal device 11 Processor 12 Memory 13 Interface 100, 100A Data analysis system
Claims
1. A data analysis apparatus comprising: a search result acquisition means for acquiring a search result regarding target data; a prompt generation means for generating, as the prompt to a language model in which machine learning has been performed to output an answer to the prompt when a prompt is input, first text data representing a question sentence regarding the characteristics of the target data based on the search result; an answer acquisition means for acquiring second text data representing an answer to the first text data based on the first text data and the language model; and a characteristic word generation means for generating a characteristic word regarding the target data based on the second text data.
2. The data analysis apparatus according to claim 1, further comprising a characteristic word correction means for correcting the characteristic word based on a search result regarding the characteristic word.
3. The data analysis apparatus according to claim 1, wherein the characteristic word correction means generates the third text data for requesting correction of the characteristic word as the prompt, and acquires a corrected characteristic word based on the third text data and the language model.
4. The data analysis apparatus according to claim 1, wherein the search result acquisition means generates search result aggregation data representing an explanatory text obtained by aggregating the search results, and the prompt generation means generates the first text data based on the search result aggregation data.
5. The data analysis apparatus according to claim 4, wherein the search result acquisition means generates the search result aggregation data representing the explanatory text using words extracted from the words based on the frequency of appearance of the words included in the search results.
6. The data analysis apparatus according to claim 4, wherein the search result acquisition means generates the search result aggregation data based on the search results and the language model.
7. The data analysis apparatus according to claim 1, wherein the characteristic word generation means selects the characteristic word from the words based on the frequency of appearance of the words included in the second text data.
8. The data analysis apparatus according to claim 1, wherein the characteristic word generation means generates the characteristic word based on the second text data and the language model.
9. Further comprising re-execution means for instructing execution of at least one of: processing of re-acquiring the search results based on the characteristic words and generating characteristic words based on the re-acquired search results; or processing of re-generating the first text data based on the characteristic words and generating characteristic words based on the re-generated first text data. The data analysis apparatus according to claim 1.
10. The re-execution means determines the necessity of the execution based on the number of the characteristic words. The data analysis apparatus according to claim 9.
11. Further comprising display control means for causing a display device to display information representing an estimation result of an attribute regarding the target data based on the characteristic words. The data analysis apparatus according to claim 1.
12. The display control means receives an input for designating the target data and causes the display device to display the information in order to assist the decision-making of the user who makes the input. The data analysis apparatus according to claim 11.
13. A data analysis method, wherein a computer: acquires a search result regarding target data; generates, based on the search result, first text data representing a question sentence regarding the characteristics of the target data as a prompt to a language model in which machine learning has been performed to output an answer to the prompt when the prompt is input; acquires second text data representing an answer to the first text data based on the first text data and the language model; and generates characteristic words regarding the target data based on the second text data.
14. A storage medium storing a program for causing a computer to execute: acquiring a search result regarding target data; generating, based on the search result, first text data representing a question sentence regarding the characteristics of the target data as a prompt to a language model in which machine learning has been performed to output an answer to the prompt when the prompt is input; acquiring second text data representing an answer to the first text data based on the first text data and the language model; and generating characteristic words regarding the target data based on the second text data.
15. Modifying the characteristic words based on a search result regarding the characteristic words. The data analysis method according to claim 13.
16. The data analysis method according to claim 13, wherein third text data for requesting correction of the characteristic word is generated as the prompt, and a characteristic word obtained by correcting the characteristic word is acquired based on the third text data and the language model.
17. The data analysis method according to claim 13, wherein search result aggregation data representing an explanatory text obtained by aggregating the search results is generated, and the first text data is generated based on the search result aggregation data.
18. The data analysis method according to claim 17, wherein the search result aggregation data representing the explanatory text using words extracted from the words based on the frequency of occurrence of the words included in the search results is generated.
19. The data analysis method according to claim 17, wherein the search result aggregation data is generated based on the search results and the language model.
20. The data analysis method according to claim 13, wherein the characteristic word is selected from the words based on the frequency of occurrence of the words included in the second text data.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A
Text generation device and text generation method
JP7313757B1